
System-level architecture of the proposed neural signal compression framework powered by the RAMAN tiny AI accelerator
Imagine a future where brain-computer interfaces (BCIs) enable people to seamlessly interact with machines using thought alone, restoring communication, enhancing sensory perception, and enabling intuitive control of assistive devices, prosthetics, and digital systems. While these systems are rapidly becoming a reality, one major challenge is that they generate enormous amounts of neural data that must be transmitted wirelessly without overheating surrounding brain tissue or consuming excessive power.
Researchers at the Department of Electronic Systems Engineering (DESE), IISc, led by Chetan Singh Thakur, have developed a neural signal compression system powered by a tiny AI accelerator called RAMAN, which compresses neural data directly within the implantable head unit before wireless transmission. The system achieves up to 150× compression on neural recordings, enabling efficient wireless transmission while still allowing the original brain signals to be reconstructed with high fidelity at the receiver for subsequent analysis and decoding.
Unlike the large AI chips used in cloud servers and chatbots, RAMAN is purpose-built for extreme low-power edge environments. The accelerator leverages ‘sparsity,’ a property in which many neural network computations are redundant and can therefore be skipped. By avoiding redundant computations and compressing stored model parameters, the accelerator dramatically reduces energy consumption and memory usage.” The team also introduced a novel stochastic pruning scheme that enables the accelerator to reconstruct pruned neural network connections without relying on large indexing tables in memory. This hardware-software co-design reduces storage overhead by up to 32% while maintaining performance.
To validate the design, the researchers implemented RAMAN on FPGA hardware and further evaluated a custom ASIC design through post-layout simulations in a TSMC 65-nm CMOS process. The post-layout results indicate that the RAMAN encoder occupies only 0.0187 mm² per channel and consumes just 15.1 µW per channel, highlighting its suitability for ultra-low-power wearable and implantable neural systems.
Overall, this work demonstrates how custom-built edge AI hardware can fundamentally reshape the design of future brain-computer interfaces. By integrating intelligent compression directly into the implantable head unit, RAMAN reduces the need for high-bandwidth wireless transmission while preserving the ability to reconstruct neural signals with high fidelity. This transition from raw data streaming to on-device compression addresses one of the key barriers to scalable, long-term neural implants: power and thermal constraints. As these systems continue to evolve toward practical clinical and assistive applications, such efficient hardware–algorithm co-designs will be critical in enabling safer, smaller, and more deployable neural interfaces.
The key team includes first author Adithya Krishna, who is jointly supervised by Chetan Singh Thakur, Andre van Schaik (Director of ICNS, University of Manchester), and Mahesh Mehendale (TI Fellow).
REFERENCES:
Krishna A, Debnath S, Srivatsav M, van Schaik A, Mehendale M, Thakur CS, Neural Signal Compression using RAMAN tinyML Accelerator for BCI Applications, IEEE Transactions on Circuits and Systems for Artificial Intelligence (2026).
https://doi.org/10.1109/TCASAI.2026.3674434
Krishna A et al, RAMAN: A Reconfigurable and Sparse tinyML Accelerator for Inference on Edge. IEEE Internet of Things Journal (2024).
https://doi.org/10.1109/JIOT.2024.3386832
LAB WEBSITE:
https://labs.dese.iisc.ac.in/neuronics/