header

Welcome


bdrp


Welcome to the Computer Vision Group at RWTH Aachen University!

The Chair for Computer Vision, headed by Prof. Bastian Leibe, draws upon many years of research expertise in the fields of computer vision and machine learning, particularly for applications in mobile robotics and intelligent vehicles. Our research topics span all areas of 3D scene understanding, including object and person detection, multi-object tracking, image and video segmentation, and 3D reconstruction. The goal of our research—which has been awarded multiple ERC grants—is to use computer vision methods and modern deep learning approaches to achieve a visual scene understanding that enables situated AI systems, such as robots or autonomous vehicles, to perform complex tasks in realistic application scenarios.

We offer lectures and seminars about computer vision and machine learning.

You can browse through our publications and the projects we are working on.

News

•

ECCV'26

We have two papers accepted at the European Conference on Computer Vision (ECCV) 2026!

Sept. 1, 2026
•

RLC'26

Our paper Biased Dreams: Limitations to Epistemic Uncertainty Quantification in Latent Dynamics Models was accepted at the 2026 Reinforcement Learning Conference (RLC)!

Aug. 12, 2026
•

SPAICE'26

Our paper "AstroPIE-53: Towards Understanding Failure Modes of Human Pose Estimation in Space" will be presented at the European Space Agency (AI for Space Applications Conference (SPAICE)) in Noordwijk, Netherlands!

June 24, 2026
•

WACV'26

Our paper "We Still See Broken Limbs: Towards Anatomical Realism in GenAI via Human Preference Learning" was accepted at the 5th Workshop on Image/Video/Audio Quality Assessment in Computer Vision, VLM and Diffusion Models at IEEE/CVF Winter Conference on Applications of Computer Vision 2026. See you in Tucson, Arizona!

Jan. 23, 2026
•

ICCV'25

Our paper DONUT: A Decoder-Only Model for Trajectory Prediction was accepted at the 2025 International Conference on Computer Vision (ICCV)!

Our project: Sa2VA-i: Improving Sa2VA Results with Consistent Training and Inference achieves 3rd Place of LSVOS Workshop, MeViS Track.

Oct. 1, 2025
•

RO-MAN'25

Our paper How do Foundation Models Compare to Skeleton-Based Approaches for Gesture Recognition in Human-Robot Interaction? has been accepted!

June 12, 2025

Recent Publications

pubimg
LVMT: Video Mask Transformer for Long-term Video Segmentation

arXiv

Existing online video segmentation methods struggle to track objects in long, complex videos with long-term occlu- sions. We hypothesize that this limitation is caused by (i) the inability of their temporal propagation mechanism to adaptively select the object information that is propagated across time, and (ii) their inability to be trained on long videos due to memory requirements and vanishing gradi- ents. To address the first limitation, we propose to use a lightweight GRU-based temporal propagation module that can learn to select which information it keeps in memory and propagates across time. Second, to allow training on long videos, we leverage Truncated Query Propagation (TQP), a training strategy in which the model processes a video in chunks of frames, where information about tracked objects is propagated between chunks but backpropagation is only conducted in individual chunks, enabling longer temporal supervision without out-of-memory issues, inference over- head, or vanishing gradients. The resulting model is called the Long-term Video Mask Transformer (LVMT). Extensive experiments on six benchmarks show that LVMT sets a new state of the art across a range of video segmentation tasks, while retaining the speed of the highly efficient model it is based on, making it 10×faster than the prior state of the art.

fadeout
 
pubimg
Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding

European Conference on Computer Vision (ECCV)

Transformers have become a common foundation across deep learning, yet 3D scene understanding still relies on specialized backbones with strong domain priors. This isolates the field from the broader Transformer ecosystem, limiting the transfer of research advances from other domains and the benefits of increasingly optimized software and hardware stacks. To bridge this gap, we propose the Volume Transformer (Volt), which adapts the vanilla Transformer encoder to 3D scenes with minimal modifications. Specifically, Volt partitions 3D scenes into volumetric patch tokens, processes them with full global self-attention, and injects positional information via 3D rotary positional embeddings (RoPE). Our initial experiments reveal that naively training Volt on standard 3D benchmarks leads to poor generalization, highlighting the limited scale of current 3D supervision. To overcome this, we introduce a data-efficient training recipe based on strong 3D augmentations, regularization, and distillation from a convolutional teacher, making Volt competitive with state-of-the-art methods. We then scale supervision through joint training on multiple datasets and show that Volt benefits more from increased scale than domain-specific 3D backbones, achieving state-of-the-art results on several indoor and outdoor semantic segmentation benchmarks. Finally, as a drop-in backbone in a standard 3D instance segmentation pipeline, Volt also sets a new state of the art, highlighting its potential as a simple, scalable, and general-purpose backbone for 3D scene understanding.

fadeout
 
pubimg
Towards Metric-Agnostic Trajectory Forecasting

European Conference on Computer Vision (ECCV) 2026

Accurate trajectory forecasting of surrounding traffic participants is a core capability for autonomous driving, enabling vehicles to anticipate behavior and plan safe maneuvers. We observe that current state-of-the-art forecasting models on Argoverse 2 and the Waymo Open Motion Dataset tailor their training objectives to the different benchmark metrics. Because these metrics encourage conflicting behavior, we propose a paradigm change for trajectory forecasting: training models with metric-agnostic probabilistic objectives and treating metric optimization as a downstream task applied to the predictive distribution. Concretely, we introduce Trajectory Distribution Evaluation (TraDiE) policies, metric-specific policies that map a predictive distribution to the set of K trajectories and confidences required by trajectory forecasting metrics. We evaluate this framework by introducing DONUT-NLL, which adapts the training objective of the state-of-the-art trajectory forecasting model DONUT to directly optimize the predictive distribution. Using our policies, DONUT-NLL achieves state-of-the-art results on all metrics of the Waymo motion prediction benchmark.

fadeout
Datenschutzerklärung/Privacy Policy Home Visual Computing institute RWTH Aachen University