We introduce MADNet-$\varepsilon$ , a lightweight architecture for event-based stereo matching. Unlike state-of-the-art models designed for this task that can barely achieve 30 FPS, the MADNet-$\varepsilon$ design enables disparity inference at more than 140 FPS without compromising accuracy, thus dramatically reducing the gap with the microseconds resolution offered by event cameras. This is achieved by stacking the event streams into histograms over small time frames (i.e., 10 ms) and by designing a recurrent feature extractor capable of maintaining temporal information beyond the single histogram using spiking convolutional long short-term memory cells. To train MADNet-$\varepsilon$ , we generate a new synthetic dataset using the CARLA simulator, $\varepsilon$-CARLA , which is used for pre-training before fine-tuning on real data. Our experiments on both $\varepsilon$-CARLA and the popular DSEC dataset demonstrate that MADNet-$\varepsilon$ achieves a favourable trade-off between accuracy and speed, representing a significant step towards event-based stereo depth estimation at millisecond resolution.

Mannocci, E., Bartolomei, L., Tosi, F., Poggi, M., Mattoccia, S. (2026). Towards Event‐Based Stereo Depth Estimation at Millisecond Resolution. IET IMAGE PROCESSING, 20(1), 1-11 [10.1049/ipr2.70472].

Towards Event‐Based Stereo Depth Estimation at Millisecond Resolution

Mannocci, Enrico;Bartolomei, Luca;Tosi, Fabio;Poggi, Matteo;Mattoccia, Stefano
2026

Abstract

We introduce MADNet-$\varepsilon$ , a lightweight architecture for event-based stereo matching. Unlike state-of-the-art models designed for this task that can barely achieve 30 FPS, the MADNet-$\varepsilon$ design enables disparity inference at more than 140 FPS without compromising accuracy, thus dramatically reducing the gap with the microseconds resolution offered by event cameras. This is achieved by stacking the event streams into histograms over small time frames (i.e., 10 ms) and by designing a recurrent feature extractor capable of maintaining temporal information beyond the single histogram using spiking convolutional long short-term memory cells. To train MADNet-$\varepsilon$ , we generate a new synthetic dataset using the CARLA simulator, $\varepsilon$-CARLA , which is used for pre-training before fine-tuning on real data. Our experiments on both $\varepsilon$-CARLA and the popular DSEC dataset demonstrate that MADNet-$\varepsilon$ achieves a favourable trade-off between accuracy and speed, representing a significant step towards event-based stereo depth estimation at millisecond resolution.
2026
Mannocci, E., Bartolomei, L., Tosi, F., Poggi, M., Mattoccia, S. (2026). Towards Event‐Based Stereo Depth Estimation at Millisecond Resolution. IET IMAGE PROCESSING, 20(1), 1-11 [10.1049/ipr2.70472].
Mannocci, Enrico; Bartolomei, Luca; Tosi, Fabio; Poggi, Matteo; Mattoccia, Stefano
File in questo prodotto:
Eventuali allegati, non sono esposti

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11585/1081390
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact