Publications
1. Efficient AI Computing Platforms
I develop memory-centric, heterogeneous, and system-aware computing platforms that reduce data movement, conversion overhead, and integration bottlenecks in AI workloads.
- R2S-CIM: Random-Reference Stochastic Interface for ADC-Less Analog Compute-in-Memory (arXiv)
TL;DR. A stochastic interface that removes high-cost ADCs from analog CIM while retaining flexible data representation, reducing conversion overhead for efficient AI execution.
- NOVA-CIM: Noise- and Correlation-Tolerant Stochastic Interfaces for Analog Compute-in-Memory (ASP-DAC 2027, accepted)
TL;DR. A stochastic CIM interface that mitigates analog noise and correlation, improving the practical efficiency and reliability of in-memory AI computation.
- Scale-CIM: A Stochastic-Computing Interface for Configurable Precision in Analog CIM Systems (arXiv)
TL;DR. A configurable-precision stochastic interface that allows analog CIM systems to adapt computation cost to workload accuracy requirements.
- TrainCIM: Design Space Exploration of Heterogeneous Multi-Core Compute-in-Memory Architectures for AI Training (arXiv)
TL;DR. A design-space exploration framework for heterogeneous multi-core CIM architectures that exposes system-level tradeoffs in efficient AI training.
- PIMScope: Package-Aware Analytical Modeling and Design-Space Exploration of Heterogeneous In-Memory Computing Systems for LLM Training (arXiv)
TL;DR. A package-aware analytical modeling framework for evaluating how compute, memory, communication, and packaging choices shape the efficiency of LLM training systems.
2. Hardware-Aware Model Representation and Adaptation
I redesign model precision, numerical representation, and adaptation mechanisms so that models can better exploit the capabilities of resource-constrained hardware.
- Binary Weight Multi-Bit Activation Quantization for Compute-in-Memory CNN Accelerators (IEEE TCAD)
TL;DR. A hardware-efficient quantization scheme that combines binary weights with multi-bit activations to reduce the cost of CIM-based CNN inference.
- NANQ: Noise-Aware Mixed-Precision Non-Uniform Quantization for Neural Networks on Analog Compute-in-Memory (ASP-DAC 2027, accepted)
TL;DR. A noise-aware mixed-precision quantization method that assigns precision non-uniformly across a model, jointly improving efficiency and robustness on analog CIM hardware.
- Exploring Layer-Wise Information Effectiveness for Post-Training Quantization in Small Language Models (ACL 2027, accepted)
TL;DR. A layer-wise analysis of information effectiveness that guides post-training quantization decisions for small language models.
- ASIQ: Adaptive-Scale Integer Quantization with 1FeFET–1RRAM Cells for Efficient LLM Inference on Analog Compute-in-Memory Systems (arXiv)
TL;DR. An adaptive-scale integer quantization approach co-designed with hybrid 1FeFET–1RRAM hardware to improve the efficiency of LLM inference.
- HaLoRA: Hardware-Aware Low-Rank Adaptation for Large Language Models Based on Hybrid Compute-in-Memory Architecture (ACM TODAES)
TL;DR. A hardware-aware LoRA framework that redesigns low-rank adaptation around the resource and architectural constraints of hybrid CIM systems.
3. Efficient Foundation-Model Execution and Generation
I reduce the computational, memory, and sequential bottlenecks of foundation-model inference by rethinking attention, caching, verification, and generation strategies.
- Recall Before You Rank: Similarity-Guided Top-K Reuse for Efficient Long-Context Attention (arXiv)
TL;DR. A similarity-guided top-k reuse strategy that avoids redundant work in long-context attention while preserving retrieval quality.
- Approximate Speculative Decoding (arXiv)
TL;DR. A speculative decoding framework that strategically relaxes exact verification to reduce the sequential cost of autoregressive language-model generation.
- RECAP: Recent-Context-Aware KV-Cache Pruning for Efficient Long-Context Speculative Decoding (arXiv)
TL;DR. A recent-context-aware KV-cache pruning method that reduces memory and computation overhead in long-context speculative decoding.
- CoCommit: Coordinating Parallel Token Commitment in Few-Step Diffusion Large Language Models (arXiv)
TL;DR. A coordinated token-commitment mechanism that improves parallel generation efficiency in few-step diffusion language models.
4. Reliable AI under Approximation and Hardware Imperfections
I study how approximation and physical nonidealities propagate through AI workloads, and develop cross-layer methods that preserve robust model behavior.
- ABNAT: Attention-Based Noise-Aware Training for Robust Transformers on Analog Compute-in-Memory Systems (ASP-DAC 2027, accepted)
TL;DR. An attention-based noise-aware training method that improves Transformer robustness to analog CIM nonidealities.
- ASSERT: Adaptive Stochastic Sampling for Robust Diffusion Models on Analog Compute-in-Memory Hardware (ASP-DAC 2027, accepted)
TL;DR. An adaptive stochastic sampling strategy that maintains diffusion-model generation quality under noisy analog CIM computation.
- Guard-of-Sink: Selective KV-Cache Protection for Noise-Resilient LLM Inference on Analog Compute-in-Memory Systems (arXiv)
TL;DR. A selective protection mechanism that identifies and safeguards critical KV-cache states, improving LLM reliability under analog CIM noise.
- ROMER: Expert Replacement and Router Calibration for Robust MoE LLMs on Analog Compute-in-Memory Systems (arXiv)
TL;DR. A cross-layer robustness framework that combines expert replacement and router calibration to stabilize MoE LLM inference under CIM nonidealities.
- Beyond Autoregression: Diffusion Language Models for Robust Analog In-Memory Generation (arXiv)
TL;DR. An investigation of diffusion language models as a generation paradigm that can be more resilient than autoregressive decoding under analog in-memory computation errors.
- When Guidance Goes Off-Scale: Recalibrating Diffusion Transformers under Analog Compute-in-Memory Nonidealities (arXiv)
TL;DR. A recalibration method that restores effective guidance in diffusion Transformers affected by analog CIM errors.
- Can We Trust LLMs on Memristors? Diving into Reasoning Ability under Non-Ideality (COLM 2027, accepted)
TL;DR. A systematic study of how memristor nonidealities affect LLM reasoning behavior, revealing reliability challenges beyond conventional task-level accuracy.
