Inference Engine from Scratch
A Qwen2.5-7B inference engine with paged KV cache and continuous batching.
A compact, single-GPU inference engine implementing the core ideas behind vLLM for Qwen2.5-7B: paged KV-cache allocation, a block allocator, FCFS scheduling, and continuous batching.
- Paged KV cache: one shared GPU pool with fixed-size blocks and a per-sequence block table.
- Block allocator: ref-counted physical blocks returned to the pool when a sequence finishes.
- Continuous batching: an FCFS scheduler slots new requests into the running batch between decode steps.
- Batched execution: prefill and decode backfill available execution slots from a waiting queue.
Benchmarked on H100: 2.6× higher decode throughput than static batching and 3.5× more sequences under an equal token budget than fixed maximum-length reservation.