Multimodal spiking-mixer with ODE neurons
The first vision-language multimodal SNN, with robustness from ODE stability analysis
A vision-language multimodal spiking neural network for image captioning, with adversarial robustness derived from a neural-ODE view of spiking dynamics.
- The first vision-language multimodal SNN for image-caption applications.
- Viewing the SNN as a neural ODE, and analyzing the stability of that ODE to gain adversarial robustness.
- Extending multimodal vision-language adversarial attacks to the SNN domain, demonstrating effectiveness compared with naive unimodal implementations.