- By:
- Gao, Shouwei ; Dong, Wenqian; Wang, Feiyi ; Yin, Junqi
- Page Number:
- 17-29
- Book Title:
- ICS '26: Proceedings of the 40th ACM International Conference on Supercomputing
- Publication Date:
- July 27, 2026
- Conference Name:
- 2026 International Conference on Supercomputing (ICS '26)
- Conference Location:
- Belfast, Ireland
- Conference Sponsor:
- ACM
- View DOI Listing:
- https://doi.org/10.1145/3797905.3800525
Abstract
Production LLM serving must simultaneously deliver high throughput, low latency, and sufficient context capacity under non-stationary traffic and mixed request requirements. Data parallelism (DP) maximizes throughput by running independent replicas, while tensor parallelism (TP) reduces per-request latency and pools memory for long-context inference. However, existing serving stacks typically commit to a static parallelism configuration at deployment; adapting to bursts, priorities, or long-context requests is often disruptive and slow. We present Flying Serving, a vLLM-based system that enables online DP-TP switching without restarting engine workers. Flying Serving makes reconfiguration practical by virtualizing the state that would otherwise force data movement: (i) a zero-copy Model Weights Manager that exposes TP shard views on demand, (ii) a KV Cache Adaptor that preserves request KV state across DP/TP layouts, (iii) an eagerly initialized Communicator Pool to amortize collective setup, and (iv) a deadlock-free scheduler that coordinates safe transitions under execution skew. Across three popular LLMs and realistic serving scenarios, Flying Serving improves performance by up to 4.79 × under high load and 3.47 × under low load while supporting latency- and memory-driven requests.