Inference Engine from Scratch

A Qwen2.5-7B inference engine with paged KV cache and continuous batching.

A compact, single-GPU inference engine implementing the core ideas behind vLLM for Qwen2.5-7B: paged KV-cache allocation, a block allocator, FCFS scheduling, and continuous batching.

  • Paged KV cache: one shared GPU pool with fixed-size blocks and a per-sequence block table.
  • Block allocator: ref-counted physical blocks returned to the pool when a sequence finishes.
  • Continuous batching: an FCFS scheduler slots new requests into the running batch between decode steps.
  • Batched execution: prefill and decode backfill available execution slots from a waiting queue.

Benchmarked on H100: 2.6× higher decode throughput than static batching and 3.5× more sequences under an equal token budget than fixed maximum-length reservation.

GitHub · Write-up

Paged versus naive KV memory capacity
Paged versus naive KV memory capacity.