RayPose: Ray Bundling Diffusion for Template Views in Unseen 6D Object Pose Estimation

Junwen Huang; Shishir Reddy Vutukur; Peter KT Yu; Nassir Navab; Slobodan Ilic; Benjamin Busam

2025 ICCV ICCV 2025

RayPose: Ray Bundling Diffusion for Template Views in Unseen 6D Object Pose Estimation

Abstract

Typical template-based object pose pipelines estimate the pose by retrieving the closest matching template and aligning it with the observed image. However, failure to retrieve the correct template often leads to inaccurate pose predictions. To address this, we reformulate template-based object pose estimation as a ray alignment problem, where the viewing directions from multiple posed template images are learned to align with a non-posed query image. Inspired by recent progress in diffusion-based camera pose estimation, we embed this formulation into a diffusion transformer architecture that aligns a query image with a set of posed templates. We reparameterize object rotation using object-centered camera rays and model object translation by extending scale-invariant translation estimation to dense translation offsets. Our model leverages geometric priors from the templates to guide accurate query pose inference. A coarse-to-fine training strategy based on narrowed template sampling improves performance without modifying the network architecture. Extensive experiments across multiple benchmark datasets show competitive results of our method compared to state-of-the-art approaches in unseen object pose estimation.

🌉 Interdisciplinary Bridge — Computer Vision and Deep Learning and Robotics

🧭 Keyword Pioneer — ray alignment

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Healthcare & Medicine, Interdisciplinary, Knowledge & Reasoning, Machine Learning, Mathematics & Optimization, Natural Language Processing, Reinforcement Learning, Robotics, Security & Privacy, Speech & Audio

Authors

Junwen Huang , Shishir Reddy Vutukur , Peter KT Yu , Nassir Navab , Slobodan Ilic , Benjamin Busam

Topics

Deep Learning > Models > Diffusion Models Computer Vision > Analysis > 3D Vision Robotics > Capabilities > Perception

Keywords

template matching geometric prior diffusion model camera pose estimation diffusion transformer coarse-to-fine training 6d object pose estimation ray alignment six-dimensional object pose estimation

Download PDF

Related papers

MA-CIR: A Multimodal Arithmetic Benchmark for Composed Image Retrieval 2025

SimMLM: A Simple Framework for Multi-modal Learning with Missing Modality 2025

MonSTeR: a Unified Model for Motion, Scene, Text Retrieval 2025

ASGS: Single-Domain Generalizable Open-Set Object Detection via Adaptive Subgraph Searching 2025

Robust Dataset Condensation using Supervised Contrastive Learning 2025