PROTEUS is an 18 mm2 programmable general-purpose digital compute-in-memory (GP-DCIM) accelerator integrating 4 Mb RRAM and 2.6 Mb tensor SRAM with a 32-bit hierarchical DCIM instruction set architecture, fabricated in 40 nm ultra-low-power CMOS with foundry RRAM. Fine-grained 1-D matrix tiling and a reconfigurable DCIM datapath/pipeline achieve near-100% memory utilization across INT8/INT16/FP8/FP16 computations. PROTEUS unifies SRAM/RRAM dataflows and embeds nonvolatile micro-programs in RRAM, enabling rapid switching among pre-stored kernels without off-chip instruction feeds or RRAM rewrites. The chip delivers 702 GOPS throughput, 6.4 TOPS/W energy efficiency, and 0.039 TOPS/mm2 compute density, validated on ResNet-20, BERT-Tiny, MobileViT, GraphSAGE, and Vision Mamba.
@article{zheng2026proteus,title={PROTEUS: A 40 nm Programmable General-Purpose Digital Compute-In-Memory Accelerator With eNVM and Hierarchical ISA for Versatile Edge AI},author={Zheng, Luqi and Bavani, Amir Massah and Chen, Mufeng and Du, Shuting and Hsin, Te-Yu and Khwa, Win-San and Wu, Ping-Sheng and Lele, Ashwin Sanjay and Zhang, Bo and Crafton, Brian and Chuang, Harry and Chih, Yu-Der and Chang, Meng-Fan and Li, Haitong},journal={IEEE Journal of Solid-State Circuits},year={2026},note={Early Access},doi={10.1109/JSSC.2026.3694285},}
ISQED
Heterogeneous Compute-in-Memory Fabrics for Efficient, Scalable Edge Inference and Learning
Luqi Zheng, Zeshu Wang, Shuting Du, and 3 more authors
In 2026 27th International Symposium on Quality Electronic Design (ISQED), 2026
Compute-in-memory fabrics built from a single memory technology force a hard trade-off among density, endurance, retention, and compute efficiency. This work examines heterogeneous CIM fabrics that combine complementary memory technologies within one architecture, and analyzes how the resulting design space maps onto efficient and scalable edge inference and on-device learning workloads.
@inproceedings{zheng2026heterogeneous,title={Heterogeneous Compute-in-Memory Fabrics for Efficient, Scalable Edge Inference and Learning},author={Zheng, Luqi and Wang, Zeshu and Du, Shuting and Chen, Mufeng and Bavani, Amir Massah and Li, Haitong},booktitle={2026 27th International Symposium on Quality Electronic Design (ISQED)},year={2026},}
2025
A-SSCC
CENTAUR: A 38.5-TFLOPS/W 600MHz Floating-Point Digital Compute-In-Memory Engine with 40nm Fusion RRAM-eDRAM Macros Featuring 3D-MAC Operation
Luqi Zheng, Amir Massah Bavani, Shuting Du, and 8 more authors
In 2025 IEEE Asian Solid-State Circuits Conference (A-SSCC), 2025
Existing NVM-based floating-point CIM macros face significant area and energy overhead from integer/FP conversion or large pre-alignment logic, accuracy degradation from row-wise pre-alignment of weights, and limited operating frequency constrained by slow NVM sensing. CENTAUR is a floating-point CIM engine featuring RRAM-eDRAM fusion macros and a novel FP 3D-MAC dataflow that eliminates alignment-induced accuracy loss, reduces area overhead, and enables high-speed, energy-efficient FP computation. Fabricated in 40 nm CMOS with foundry RRAM, CENTAUR achieves 600 MHz operation and 38.5 TFLOPS/W, running Tiny-ViT on CIFAR-10 with only 1.75% accuracy degradation versus the software baseline.
@inproceedings{zheng2025centaur,title={CENTAUR: A 38.5-TFLOPS/W 600MHz Floating-Point Digital Compute-In-Memory Engine with 40nm Fusion RRAM-eDRAM Macros Featuring 3D-MAC Operation},author={Zheng, Luqi and Bavani, Amir Massah and Du, Shuting and Hsin, Te-Yu and Chen, Mufeng and Khwa, Win-San and Lele, Ashwin and Chuang, Harry and Chih, Yu-Der and Chang, Meng-Fan and Li, Haitong},booktitle={2025 IEEE Asian Solid-State Circuits Conference (A-SSCC)},year={2025},}
IMW
Analog Multilevel eDRAM-RRAM CIM for Zeroth-Order Fine-tuning of LLMs
Mufeng Chen, Luqi Zheng, Jheng-Yu Lin, and 2 more authors
In 2025 IEEE International Memory Workshop (IMW), 2025
Zeroth-order fine-tuning eliminates explicit back-propagation and reduces memory overhead for large language models, making it a promising approach for on-device fine-tuning. However, existing memory-centric accelerators fail to fully leverage these benefits due to inefficiencies in balancing bit density, compute-in-memory capability, and the endurance-retention trade-off. We present a reliability-aware, analog multi-level-cell eDRAM-RRAM compute-in-memory solution co-designed with zeroth-order optimization for language model fine-tuning. An RRAM-assisted eDRAM MLC programming scheme is developed, along with a PVT-robust, large-sensing-window time-to-digital converter. The MLC-eDRAM integrating two-finger MOM provides 12x improvement in bit density over state-of-the-art MLC designs, with another 5x density and 2x retention benefit from BEOL In2O3 FETs.
@inproceedings{chen2025analog,title={Analog Multilevel eDRAM-RRAM CIM for Zeroth-Order Fine-tuning of LLMs},author={Chen, Mufeng and Zheng, Luqi and Lin, Jheng-Yu and Ye, Peide D. and Li, Haitong},booktitle={2025 IEEE International Memory Workshop (IMW)},pages={1--4},year={2025},doi={10.1109/IMW61990.2025.11026966},}
2023
EM Sci.
An Electromagnetic Perspective of Artificial Intelligence Neuromorphic Chips
Er-Ping Li, Hanzhi Ma, Manareldeen Ahmed, and 6 more authors
Traditional general-purpose chips based on von Neumann architecture face the memory wall problem when applied to artificial intelligence. Inspired by the efficiency of the human brain, many neuromorphic chips have been proposed to emulate its working mechanism and neuron-synapse structure. Here we review neuromorphic circuit design, algorithms and applications, with a focus on signal integrity issues, modeling, and optimization, and on the heterogeneous integration of neuromorphic circuits with memory arrays and sensors.
@article{li2023electromagnetic,title={An Electromagnetic Perspective of Artificial Intelligence Neuromorphic Chips},author={Li, Er-Ping and Ma, Hanzhi and Ahmed, Manareldeen and Tao, Tianhao and Gu, Zheming and Chen, Mufeng and Chen, Quankun and Li, Da and Chen, Wenchao},journal={Electromagnetic Science},volume={1},number={3},pages={0030151},year={2023},doi={10.23919/emsci.2023.0015},}
2022
Front. Neurosci.
MAP-SNN: Mapping Spike Activities with Multiplicity, Adaptability, and Plasticity into Bio-Plausible Spiking Neural Networks
Chengting Yu, Yangkai Du, Mufeng Chen, and 3 more authors
Spiking Neural Networks are considered more biologically realistic and power-efficient as they imitate the fundamental mechanism of the human brain. Toward bio-plausible backpropagation-based SNNs, we consider three properties in modeling spike activities: Multiplicity, Adaptability, and Plasticity (MAP). We propose a Multiple-Spike Pattern with multiple spike transmission to strengthen model robustness in discrete time iteration, adopt Spike Frequency Adaption to decrease spike activities for improved efficiency, and propose a trainable convolutional synapse that models spike response current to enhance the diversity of spiking neurons for temporal feature extraction.
@article{yu2022mapsnn,title={MAP-SNN: Mapping Spike Activities with Multiplicity, Adaptability, and Plasticity into Bio-Plausible Spiking Neural Networks},author={Yu, Chengting and Du, Yangkai and Chen, Mufeng and Wang, Aili and Wang, Gaoang and Li, Erping},journal={Frontiers in Neuroscience},volume={16},pages={945037},year={2022},doi={10.3389/fnins.2022.945037},}