![[Paper Reading Club] Stop Wasting GPU VRAM — PagedAttention & vLLM Architecture [Paper Reading Club] Stop Wasting GPU VRAM — PagedAttention & vLLM Architecture](https://json.commudle.com/rails/active_storage/blobs/proxy/eyJfcmFpbHMiOnsibWVzc2FnZSI6IkJBaHBBMFdRQnc9PSIsImV4cCI6bnVsbCwicHVyIjoiYmxvYl9pZCJ9fQ==--6357857fd28692fbe5d9f61fd8ad6bff22a768fb/com_7dc86c4cf581c684_20260807125248.png)
[Paper Reading Club] Stop Wasting GPU VRAM — PagedAttention & vLLM Architecture
Attendee Registration
Welcome to Session #01 of the ML Indore Paper Reading Club!
In this online session, we are opening up and reading through: 📄 Efficient Memory Management for Large Language Model Serving with PagedAttention Authors: Woosuk Kwon et al. (UC Berkeley SkyLab) — Read on arXiv
📖 How We Read Papers Together
You don't need to be a C++ kernel developer or math wizard to join. We will have the PDF on screen, walking through the core architectural diagrams, section by section:
-
The Context & Bottleneck (Sections 1–2): Why static KV cache pre-allocation wastes 60–80% of GPU VRAM.
-
The Core Intuition (Section 3):
View More
