Publications

1. Efficient AI Computing Platforms

I develop memory-centric, heterogeneous, and system-aware computing platforms that reduce data movement, conversion overhead, and integration bottlenecks in AI workloads.

  • R2S-CIM: Random-Reference Stochastic Interface for ADC-Less Analog Compute-in-Memory (arXiv)

    TL;DR. A stochastic interface that removes high-cost ADCs from analog CIM while retaining flexible data representation, reducing conversion overhead for efficient AI execution.

  • NOVA-CIM: Noise- and Correlation-Tolerant Stochastic Interfaces for Analog Compute-in-Memory (ASP-DAC 2027, accepted)

    TL;DR. A stochastic CIM interface that mitigates analog noise and correlation, improving the practical efficiency and reliability of in-memory AI computation.

  • Scale-CIM: A Stochastic-Computing Interface for Configurable Precision in Analog CIM Systems (arXiv)

    TL;DR. A configurable-precision stochastic interface that allows analog CIM systems to adapt computation cost to workload accuracy requirements.

  • TrainCIM: Design Space Exploration of Heterogeneous Multi-Core Compute-in-Memory Architectures for AI Training (arXiv)

    TL;DR. A design-space exploration framework for heterogeneous multi-core CIM architectures that exposes system-level tradeoffs in efficient AI training.

  • PIMScope: Package-Aware Analytical Modeling and Design-Space Exploration of Heterogeneous In-Memory Computing Systems for LLM Training (arXiv)

    TL;DR. A package-aware analytical modeling framework for evaluating how compute, memory, communication, and packaging choices shape the efficiency of LLM training systems.

2. Hardware-Aware Model Representation and Adaptation

I redesign model precision, numerical representation, and adaptation mechanisms so that models can better exploit the capabilities of resource-constrained hardware.

  • Binary Weight Multi-Bit Activation Quantization for Compute-in-Memory CNN Accelerators (IEEE TCAD)

    TL;DR. A hardware-efficient quantization scheme that combines binary weights with multi-bit activations to reduce the cost of CIM-based CNN inference.

  • NANQ: Noise-Aware Mixed-Precision Non-Uniform Quantization for Neural Networks on Analog Compute-in-Memory (ASP-DAC 2027, accepted)

    TL;DR. A noise-aware mixed-precision quantization method that assigns precision non-uniformly across a model, jointly improving efficiency and robustness on analog CIM hardware.

  • Exploring Layer-Wise Information Effectiveness for Post-Training Quantization in Small Language Models (ACL 2027, accepted)

    TL;DR. A layer-wise analysis of information effectiveness that guides post-training quantization decisions for small language models.

  • ASIQ: Adaptive-Scale Integer Quantization with 1FeFET–1RRAM Cells for Efficient LLM Inference on Analog Compute-in-Memory Systems (arXiv)

    TL;DR. An adaptive-scale integer quantization approach co-designed with hybrid 1FeFET–1RRAM hardware to improve the efficiency of LLM inference.

  • HaLoRA: Hardware-Aware Low-Rank Adaptation for Large Language Models Based on Hybrid Compute-in-Memory Architecture (ACM TODAES)

    TL;DR. A hardware-aware LoRA framework that redesigns low-rank adaptation around the resource and architectural constraints of hybrid CIM systems.

3. Efficient Foundation-Model Execution and Generation

I reduce the computational, memory, and sequential bottlenecks of foundation-model inference by rethinking attention, caching, verification, and generation strategies.

  • Recall Before You Rank: Similarity-Guided Top-K Reuse for Efficient Long-Context Attention (arXiv)

    TL;DR. A similarity-guided top-k reuse strategy that avoids redundant work in long-context attention while preserving retrieval quality.

  • Approximate Speculative Decoding (arXiv)

    TL;DR. A speculative decoding framework that strategically relaxes exact verification to reduce the sequential cost of autoregressive language-model generation.

  • RECAP: Recent-Context-Aware KV-Cache Pruning for Efficient Long-Context Speculative Decoding (arXiv)

    TL;DR. A recent-context-aware KV-cache pruning method that reduces memory and computation overhead in long-context speculative decoding.

  • CoCommit: Coordinating Parallel Token Commitment in Few-Step Diffusion Large Language Models (arXiv)

    TL;DR. A coordinated token-commitment mechanism that improves parallel generation efficiency in few-step diffusion language models.

4. Reliable AI under Approximation and Hardware Imperfections

I study how approximation and physical nonidealities propagate through AI workloads, and develop cross-layer methods that preserve robust model behavior.

  • ABNAT: Attention-Based Noise-Aware Training for Robust Transformers on Analog Compute-in-Memory Systems (ASP-DAC 2027, accepted)

    TL;DR. An attention-based noise-aware training method that improves Transformer robustness to analog CIM nonidealities.

  • ASSERT: Adaptive Stochastic Sampling for Robust Diffusion Models on Analog Compute-in-Memory Hardware (ASP-DAC 2027, accepted)

    TL;DR. An adaptive stochastic sampling strategy that maintains diffusion-model generation quality under noisy analog CIM computation.

  • Guard-of-Sink: Selective KV-Cache Protection for Noise-Resilient LLM Inference on Analog Compute-in-Memory Systems (arXiv)

    TL;DR. A selective protection mechanism that identifies and safeguards critical KV-cache states, improving LLM reliability under analog CIM noise.

  • ROMER: Expert Replacement and Router Calibration for Robust MoE LLMs on Analog Compute-in-Memory Systems (arXiv)

    TL;DR. A cross-layer robustness framework that combines expert replacement and router calibration to stabilize MoE LLM inference under CIM nonidealities.

  • Beyond Autoregression: Diffusion Language Models for Robust Analog In-Memory Generation (arXiv)

    TL;DR. An investigation of diffusion language models as a generation paradigm that can be more resilient than autoregressive decoding under analog in-memory computation errors.

  • When Guidance Goes Off-Scale: Recalibrating Diffusion Transformers under Analog Compute-in-Memory Nonidealities (arXiv)

    TL;DR. A recalibration method that restores effective guidance in diffusion Transformers affected by analog CIM errors.

  • Can We Trust LLMs on Memristors? Diving into Reasoning Ability under Non-Ideality (COLM 2027, accepted)

    TL;DR. A systematic study of how memristor nonidealities affect LLM reasoning behavior, revealing reliability challenges beyond conventional task-level accuracy.