HKUST Computer Architecture Group
HKUST Computer Architecture Group
News
People
Events
Publications
Contact
F. Tu
Latest
[Forthcoming] Symposium on VLSI Technology and Circuits (VLSI) 2026
A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding
D2CIM: A 28nm 53.3 TFLOPS/W Decoding Digital CIM Macro for Efficient FlashMLA-based LLM Inference
Denim: Heterogeneous Compute-in-Memory Accelerator Exploiting Denoising-similarity for Diffusion Models
ETCIM: Error-Tolerant Digital CIM Processor with Redundancy-Free Hard Error Repair and Run-Time Soft Error Correction
HR-DCIM: High-Reliability Floating-Point Digital CIM Architecture with Unified Low-Cost Iterative Error Correction
MRCIM: A Many-Core Reconfigurable Computing-in-Memory Processor Combining CPU and Tensor Modes for NN Acceleration
VAR-Turbo: Unlocking the Potential of Visual Autoregressive Models through Dual Redundancy
A 28nm 0.22uJ/Token Memory-Compute-Intensity-Aware CNN-Transformer Accelerator with Hybrid-Attention-Based Layer-Fusion and Cascaded Pruning for Semantic-Segmentation
CELLA: A 28nm Compute-Memory Co-Optimized Real-Time Digital CIM-based Edge LLM Accelerator with 1.78ms-response in Prefill and 31.32 Token/s in Decoding
CoXplorer: Multi-staged Co-exploration Framework for AI Model Compression and Accelerator Design
CV-CIM: A Hybrid Domain XOR-derived Similarity-aware Computation-in-memory Supporting Cost Volume Construction
ER-DCIM: Error-Resilient Digital CIM with Run-Time MAC-Cell Error Correction
Exploiting the Memory-Compute-Coupling Feature for CIM Accelerator Design Optimization
LLM-CIM: A 28nm 126.7 TOPS/W Input-LUT-based Digital CIM Macro with Reconfigurable Matrix Multiplication and Nonlinear Operation Modes for LLMs
PipeDCIM: 28nm 115.11 TOPS/mm2xTOPS/W@1.24GHz Pipeline Digital CIM Macro with an Auto-Design Tool for Diverse High-Performance AI Scenarios
Rethinking Control Flow in Spatial Architectures: Insights into Control Flow Plane Design
Robotic Computing System and Embodied AI Evolution: An Algorithm-Hardware Co-Design Perspective
TensorCIM: Digital Computing-In-Memory Tensor Processor with Multi-Chip-Module-based Architecture for Beyond-NN Acceleration
Cite
×