Distillation-free Scaling of Large State-Space Models for Images and Videos
State-space models ({SSMs}), exemplified by S4, have introduced a novel context modeling method by integrating state-space techniques into deep learning. Despite their effectiveness, {SSMs} struggle with global context modeling due to data-independent matrices. The Mamba model addresses this with data-dependent variants enabled by the S6 selective-scan algorithm, enhancing context modeling, especially for long sequences. However, Mamba-based architectures face significant parameter scalability challenges, limiting their utility in vision applications. This paper tackles the scalability issue of large {SSMs} for image classification and action recognition without relying on additional techniques like knowledge distillation. We analyze the distinct characteristics of Mamba-based and Attention-based models, proposing a Mamba-Attention interleaved architecture that enhances scalability, robustness, and performance. We demonstrate that the stable and efficient interleaved architecture resolves the scalability issue of Mamba-based architectures and increases robustness to common corruption artifacts. Our thorough evaluation on the {ImageNet}-1K, Kinetics-400, and Something-Something-v2 benchmarks demonstrates that our approach improves the accuracy of state-of-the-art Mamba-based architectures by up to \$\$+1.7\$\$\%.
- Published in:
International Journal of Computer Vision - Type:
Article - Authors:
- Year:
2026 - Source:
https://doi.org/10.1007/s11263-026-02824-0
Citation information
: Distillation-free Scaling of Large State-Space Models for Images and Videos, International Journal of Computer Vision, 2026, 134, 5, 232, April, https://doi.org/10.1007/s11263-026-02824-0, Suleman.etal.2026a,
@Article{Suleman.etal.2026a,
author={Suleman, Hamid; Wasim, Syed Talal; Naseer, Muzammal; Gall, Juergen},
title={Distillation-free Scaling of Large State-Space Models for Images and Videos},
journal={International Journal of Computer Vision},
volume={134},
number={5},
pages={232},
month={April},
url={https://doi.org/10.1007/s11263-026-02824-0},
year={2026},
abstract={State-space models ({SSMs}), exemplified by S4, have introduced a novel context modeling method by integrating state-space techniques into deep learning. Despite their effectiveness, {SSMs} struggle with global context modeling due to data-independent matrices. The Mamba model addresses this with data-dependent variants enabled by the S6 selective-scan algorithm, enhancing context modeling,...}}