Geometry Matters: 3D Foundation Priors for Learning Semantic Correspondence

Artur Jesslen¹, Olaf Dünkel², and Adam Kortylewski³

¹University of Freiburg, Germany ²Max Planck Institute for Informatics, Saarland Informatics Campus, Germany ³CISPA Helmholtz Center for Information Security, Germany
Arxiv 2026

Abstract

Foundation features from self-supervised vision models and text-to-image diffusion models have proven effective for semantic correspondence estimation. However, because these features are learned primarily from 2D image objectives, they lack explicit 3D awareness and often confuse symmetric object sides, repeated parts, and visually similar structures that are distinct in 3D. We introduce a 3D-aware post-training framework that goes beyond available 2D foundation features by incorporating priors from 3D foundation models. Given an image, our method uses SAM3D to estimate object geometry and pose, and refines the pose through render-and-compare optimization. Subsequently, we render PartField descriptors from the reconstructed geometry into the image plane based on the estimated object pose. The resulting geometry-aware feature maps complement DINO and Stable Diffusion features, while geodesic distances on the reconstructed shapes enable reliable filtering of candidate correspondences. We use the filtered matches as supervision to train a lightweight adapter on top of DINO and Stable Diffusion for semantic correspondence. In contrast to prior post-training approaches that require pose annotations and rely on coarse spherical geometry, our method automatically obtains instance-specific 3D structure and uses it to guide correspondence learning. Experiments show that our approach improves semantic correspondence over prior methods while reducing manual geometric supervision.

Method

Canonicalized 3D object reconstruction pipeline. Given an image, we obtain an instance mask and a mesh from foundation models. We refine the mesh pose via a two-phase render-and-compare optimization (distance-transform + soft-IoU), then resolve the four-fold yaw ambiguity using OrientAnything V2 with majority voting.

Pseudo-label correspondences pipeline. DINO, SD, and PartField features (rasterized from the reconstructed meshes) are fused and candidate matches are proposed via nearest-neighbor search with relaxed cyclic consistency. Each candidate is then verified by lifting matched pixels onto the meshes and computing the bicyclic geodesic error; candidates exceeding the threshold are rejected.

@inproceedings{jesslen2026geometry, author = {Artur Jesslen and Olaf D{\"u}nkel and Adam Kortylewski}, title = {Geometry Matters: 3D Foundation Priors for Learning Semantic Correspondence}, booktitle = {Arxiv}, year = {2026} }

Geometry Matters: 3D Foundation Priors for Learning Semantic Correspondence

Abstract

(a) SD+DINO: existing zero-shot pipelines suffer from left–right and repeated-part confusion, producing many incorrect matches.

(b) SD+DINO + Geodesic Filtering: adding our geodesic filter removes wrong matches but is bottlenecked by feature quality, often leaving few surviving correspondences.

(c) SD+DINO+PartField + Geodesic Filtering: adding PartField features yields dense and accurate correspondences even with large pose changes.

Qualitative pseudo-annotations. 3D-SC produces denser and more geometrically consistent pseudo-annotations than DIY-SC, free from left–right ambiguities.

Method

BibTeX