We present CalibBEV, a novel Bird’s Eye View (BEV) alignment approach for LiDAR-camera calibration. Our method unifies LiDAR and camera data into a shared 3D spatial representation, enabling accurate and robust cross- modal calibration. CalibBEV extracts sensor-wise BEV fea- tures from each modality using domain-specific architec- tures and estimates the calibration matrix through a two- step alignment process. First, we perform an implicit align- ment by regressing a coarse calibration matrix directly from the BEV features. To ease this alignment, we enforce seman- tic consistency between BEV representations across modal- ities using a contrastive loss inspired by CLIP, guiding both networks toward a unified feature space. In the second step, we leverage our BEV formulation to explicitly align the features of one modality with the other, refining the ini- tial coarse estimate into a final, more accurate calibration matrix. CalibBEV significantly outperforms prior point-to- pixel matching methods, achieving state-of-the-art calibra- tion accuracy. On the KITTI and nuScenes benchmarks, our method reduces the Relative Rotation Error (RRE) by 51% and 68%, and the Relative Translation Error (RTE) by 80% and 91%, respectively, compared to previous methods.
D'Addeo, F., Cipelli, L., Cardace, A., Ghelfi, E., Zinelli, A., Bertozzi, M. (2026). CalibBEV: LiDAR-Camera Calibration via BEV Alignment. IEEE [10.1109/WACV61042.2026.00423].
CalibBEV: LiDAR-Camera Calibration via BEV Alignment
D'Addeo, Filippo;Cardace, Adriano;Bertozzi, Massimo
2026
Abstract
We present CalibBEV, a novel Bird’s Eye View (BEV) alignment approach for LiDAR-camera calibration. Our method unifies LiDAR and camera data into a shared 3D spatial representation, enabling accurate and robust cross- modal calibration. CalibBEV extracts sensor-wise BEV fea- tures from each modality using domain-specific architec- tures and estimates the calibration matrix through a two- step alignment process. First, we perform an implicit align- ment by regressing a coarse calibration matrix directly from the BEV features. To ease this alignment, we enforce seman- tic consistency between BEV representations across modal- ities using a contrastive loss inspired by CLIP, guiding both networks toward a unified feature space. In the second step, we leverage our BEV formulation to explicitly align the features of one modality with the other, refining the ini- tial coarse estimate into a final, more accurate calibration matrix. CalibBEV significantly outperforms prior point-to- pixel matching methods, achieving state-of-the-art calibra- tion accuracy. On the KITTI and nuScenes benchmarks, our method reduces the Relative Rotation Error (RRE) by 51% and 68%, and the Relative Translation Error (RTE) by 80% and 91%, respectively, compared to previous methods.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



