Skip to main content

Showing 1–4 of 4 results for author: Patke, A

.
  1. arXiv:2407.00047  [pdf, other

    cs.DC cs.CL cs.LG

    One Queue Is All You Need: Resolving Head-of-Line Blocking in Large Language Model Serving

    Authors: Archit Patke, Dhemath Reddy, Saurabh Jha, Haoran Qiu, Christian Pinto, Shengkun Cui, Chandra Narayanaswami, Zbigniew Kalbarczyk, Ravishankar Iyer

    Abstract: $ $Large language models (LLMs) have become an increasingly important workload for cloud providers catering to both enterprise and consumer applications. LLM inference requests from these applications have end-to-end latency SLOs that must be adhered to in production settings. However, existing LLM serving systems focus on optimization objectives such as request serving throughput or request execu… ▽ More

    Submitted 5 June, 2024; originally announced July 2024.

  2. arXiv:2404.08509  [pdf, other

    cs.DC cs.CL cs.LG

    Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction

    Authors: Haoran Qiu, Weichao Mao, Archit Patke, Shengkun Cui, Saurabh Jha, Chen Wang, Hubertus Franke, Zbigniew T. Kalbarczyk, Tamer Başar, Ravishankar K. Iyer

    Abstract: Large language models (LLMs) have been driving a new wave of interactive AI applications across numerous domains. However, efficiently serving LLM inference requests is challenging due to their unpredictable execution times originating from the autoregressive nature of generative models. Existing LLM serving systems exploit first-come-first-serve (FCFS) scheduling, suffering from head-of-line bloc… ▽ More

    Submitted 12 April, 2024; originally announced April 2024.

    Comments: Accepted at AIOps'24

  3. arXiv:2012.07755  [pdf, other

    cs.DC cs.NI

    Application-aware Congestion Mitigation for High-Performance Computing Systems

    Authors: Archit Patke, Saurabh Jha, Haoran Qiu, Jim Brandt, Ann Gentile, Joe Greenseid, Zbigniew Kalbarczyk, Ravishankar Iyer

    Abstract: High-performance computing (HPC) systems frequently experience congestion leading to significant application performance variation. However, the impact of congestion on application runtime differs from application to application depending on their network characteristics (such as bandwidth and latency requirements). We leverage this insight to develop Netscope, an automated ML-driven framework tha… ▽ More

    Submitted 3 February, 2021; v1 submitted 14 December, 2020; originally announced December 2020.

  4. arXiv:1907.05312  [pdf, other

    cs.DC cs.NI

    A Study of Network Congestion in Two Supercomputing High-Speed Interconnects

    Authors: Saurabh Jha, Archit Patke, Jim Brandt, Ann Gentile, Mike Showerman, Eric Roman, Zbigniew T. Kalbarczyk, William T. Kramer, Ravishankar K. Iyer

    Abstract: Network congestion in high-speed interconnects is a major source of application run time performance variation. Recent years have witnessed a surge of interest from both academia and industry in the development of novel approaches for congestion control at the network level and in application placement, map**, and scheduling at the system-level. However, these studies are based on proxy applicat… ▽ More

    Submitted 11 July, 2019; originally announced July 2019.

    Comments: Accepted for HOTI2019