Details
Welcome to Session #01 of the ML Indore Paper Reading Club!
In this online session, we are opening up and reading through: π Efficient Memory Management for Large Language Model Serving with PagedAttention Authors: Woosuk Kwon et al. (UC Berkeley SkyLab) β Read on arXiv
π How We Read Papers Together
You don't need to be a C++ kernel developer or math wizard to join. We will have the PDF on screen, walking through the core architectural diagrams, section by section:
-
The Context & Bottleneck (Sections 1β2): Why static KV cache pre-allocation wastes 60β80% of GPU VRAM.
-
The Core Intuition (Section 3): How the paper adapts OS Virtual Memory Paging to GPU tensors.
-
Architecture & Implementation (Section 4): Looking at logical-to-physical block tables and dynamic allocation.
-
Group Discussion & Critique: Where does this paper break down? What are the hardware trade-offs?
π€ Session Facilitator
Karan Mittal (CTO & Co-Founder at Dextar Intelligence) will lead the walkthrough, breaking down key diagrams and facilitating group discussion.
-
π Portfolio & Background: karan-s-mittal.github.io/about
π‘ How to Prepare
-
π Skim the PDF ahead of time: Don't worry about understanding every equationβjust glance at Figure 1 and Figure 2.
-
β Bring 1 question or takeaway: Note down one thing that made you curious or confused.
-
π» Have the PDF open: We encourage everyone to read along during the session.
π Session Details
-
π Date: Tuesday, August 11, 2026
-
β° Time: 10:00 PM β 11:00 PM IST
-
π Venue: Online (Virtual Meet link sent upon registration)

![[Paper Reading Club] Stop Wasting GPU VRAM β PagedAttention & vLLM Architecture [Paper Reading Club] Stop Wasting GPU VRAM β PagedAttention & vLLM Architecture](https://json.commudle.com/rails/active_storage/blobs/proxy/eyJfcmFpbHMiOnsibWVzc2FnZSI6IkJBaHBBMFdRQnc9PSIsImV4cCI6bnVsbCwicHVyIjoiYmxvYl9pZCJ9fQ==--6357857fd28692fbe5d9f61fd8ad6bff22a768fb/com_7dc86c4cf581c684_20260807125248.png)



