Show HN: Proxima serves 4x more requests with no hardware change on vLLM https://ift.tt/LD9F3Ml

Show HN: Proxima serves 4x more requests with no hardware change on vLLM hey everyone, i decided to make a vLLM plugin that implements the Star-KV paper. the results are quiet promising with a decode kernel thats faster than FA2 in higher batch sizes. would love any thoughts and recommendations https://ift.tt/vuUfbpx August 11, 2026 at 08:34PM

Comments

Popular posts from this blog