[Paper Reading Club] Stop Wasting GPU VRAM — PagedAttention & vLLM Architecture

[Paper Reading Club] Stop Wasting GPU VRAM — PagedAttention & vLLM Architecture

Attendee Registration

Welcome to Session #01 of the ML Indore Paper Reading Club!

In this online session, we are opening up and reading through: 📄 Efficient Memory Management for Large Language Model Serving with PagedAttention Authors: Woosuk Kwon et al. (UC Berkeley SkyLab) — Read on arXiv

📖 How We Read Papers Together

You don't need to be a C++ kernel developer or math wizard to join. We will have the PDF on screen, walking through the core architectural diagrams, section by section:

  1. The Context & Bottleneck (Sections 1–2): Why static KV cache pre-allocation wastes 60–80% of GPU VRAM.

  2. The Core Intuition (Section 3):

View More

Never enter any personal or sensitive information which can be misused (including but not limited to passwords) in on Commudle. If you find something inappropriately asked, please report it to more@commudle.com immediately.