Research and Materials on Hardware implementation of Transformer Model
-
Updated
Feb 28, 2025 - Jupyter Notebook
Research and Materials on Hardware implementation of Transformer Model
PIMeval simulator and PIMbench suite
AIM, A Framework for High-throughput Sequence Alignment using Real Processing-in-Memory Systems, Bioinformatics, btad155, https://doi.org/10.1093/bioinformatics/btad155
A PIM instrumentation, compilation, execution, simulation, and evaluation repository for BLIMP-style architectures.
Source code & scripts for experimental characterization and demonstration of performing NOT and up to 16-input AND, NAND, OR, and NOR operations in real DDR4 DRAM chips. Described in our HPCA'24 paper by Yuksel et al. at https://arxiv.org/abs/2402.18736
UpPipe is an RNA abundance quantification design on a real processing-near-memory system (UPMEM DPU); the paper of this project is published in Design Automation Conference (DAC) 2023
Running state-of-the-art RNA-seq abundance quantification software "kallisto" on UPMEM DPU system
Aletheia is a memory-centric experimentation framework for exploring Processing near Memory (PNM) concepts on commodity hardware.
Processing-In-Memory acceleration of Breadth-First Search on DPUs.
Code for distributed inference of WNN on the UPMEM PiM System
Deyuan's fork with -print_tech_params, -print_decoder_breakdown, and macOS compatibility
Implementation of WNNs on the UPMEM PiM system. This implementation is optimizing for MRAM utilization.
A reference implementation of Matrix Multiplication algorithms for ML on UPMEM PIM - a processing-in-memory platform
A Basic Processing in Memory (PIM) System (Project of "Modern VLSI Design" Course)
Distributed Training on the UPMEM PiM system
Paper critiques & labs on emerging memory and storage systems — from NAND Flash/SSD design and phase-change memory to shingled disks, key-value stores and processing-in-memory. NCKU coursework.
Analytical model: video VLM serving is limited by KV cache capacity, not bandwidth. 113x fewer concurrent users than text on the same node.
Master's thesis research artifact for exact tensor-network quantum-circuit simulation on UPMEM Processing-in-Memory hardware.
Multicycle RISC-V CPU and 32-bit PIM accelerator in SystemVerilog, with verified MMIO, shared-memory integration, and matrix-vector benchmarks.
We present an analog in-memory computing (AIMC) evaluation framework, providing SW/HW performance for LLM inference.
To associate your repository with the processing-in-memory topic, visit your repo's landing page and select "manage topics."