We introduce MADNet-$\varepsilon$ , a lightweight architecture for event-based stereo matching. Unlike state-of-the-art models designed for this task that can barely achieve 30 FPS, the MADNet-$\varepsilon$ design enables disparity inference at more than 140 FPS without compromising accuracy, thus dramatically reducing the gap with the microseconds resolution offered by event cameras. This is achieved by stacking the event streams into histograms over small time frames (i.e., 10 ms) and by designing a recurrent feature extractor capable of maintaining temporal information beyond the single histogram using spiking convolutional long short-term memory cells. To train MADNet-$\varepsilon$ , we generate a new synthetic dataset using the CARLA simulator, $\varepsilon$-CARLA , which is used for pre-training before fine-tuning on real data. Our experiments on both $\varepsilon$-CARLA and the popular DSEC dataset demonstrate that MADNet-$\varepsilon$ achieves a favourable trade-off between accuracy and speed, representing a significant step towards event-based stereo depth estimation at millisecond resolution.
Mannocci, E., Bartolomei, L., Tosi, F., Poggi, M., Mattoccia, S. (2026). Towards Event‐Based Stereo Depth Estimation at Millisecond Resolution. IET IMAGE PROCESSING, 20(1), 1-11 [10.1049/ipr2.70472].
Towards Event‐Based Stereo Depth Estimation at Millisecond Resolution
Mannocci, Enrico;Bartolomei, Luca;Tosi, Fabio;Poggi, Matteo;Mattoccia, Stefano
2026
Abstract
We introduce MADNet-$\varepsilon$ , a lightweight architecture for event-based stereo matching. Unlike state-of-the-art models designed for this task that can barely achieve 30 FPS, the MADNet-$\varepsilon$ design enables disparity inference at more than 140 FPS without compromising accuracy, thus dramatically reducing the gap with the microseconds resolution offered by event cameras. This is achieved by stacking the event streams into histograms over small time frames (i.e., 10 ms) and by designing a recurrent feature extractor capable of maintaining temporal information beyond the single histogram using spiking convolutional long short-term memory cells. To train MADNet-$\varepsilon$ , we generate a new synthetic dataset using the CARLA simulator, $\varepsilon$-CARLA , which is used for pre-training before fine-tuning on real data. Our experiments on both $\varepsilon$-CARLA and the popular DSEC dataset demonstrate that MADNet-$\varepsilon$ achieves a favourable trade-off between accuracy and speed, representing a significant step towards event-based stereo depth estimation at millisecond resolution.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



