Multimodal spiking-mixer with ODE neurons

The first vision-language multimodal SNN, with robustness from ODE stability analysis

A vision-language multimodal spiking neural network for image captioning, with adversarial robustness derived from a neural-ODE view of spiking dynamics.

  • The first vision-language multimodal SNN for image-caption applications.
  • Viewing the SNN as a neural ODE, and analyzing the stability of that ODE to gain adversarial robustness.
  • Extending multimodal vision-language adversarial attacks to the SNN domain, demonstrating effectiveness compared with naive unimodal implementations.