<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Vllm on SREKubeCraft | Nick Nikolakakis</title><link>https://srekubecraft.io/tags/vllm/</link><description>Recent content in Vllm on SREKubeCraft | Nick Nikolakakis</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sat, 15 Aug 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://srekubecraft.io/tags/vllm/index.xml" rel="self" type="application/rss+xml"/><item><title>llm-d - Kubernetes-Native Distributed LLM Inference at Scale</title><link>https://srekubecraft.io/posts/llm-d-distributed-inference/</link><pubDate>Sat, 15 Aug 2026 00:00:00 +0000</pubDate><guid>https://srekubecraft.io/posts/llm-d-distributed-inference/</guid><description>&lt;p&gt;Three months ago I &lt;a href="https://srekubecraft.io/posts/kserve/"&gt;wrote about KServe&lt;/a&gt; and how the &lt;code&gt;InferenceService&lt;/code&gt; CRD had become the closest thing cloud-native has to a standard for putting a trained model behind an API. That post ended on a deliberate cliffhanger: KServe gives you a great single-model serving primitive, but it does not solve GPU sharing, fractional scheduling, or how you load-balance inference traffic across many replicas of a large model. I pointed at Volcano and Kueue and moved on.&lt;/p&gt;</description></item></channel></rss>