<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>LLM Serving on Hellokitty&#39;s Blog</title>
    <link>https://yangyang233333.github.io/tags/llm-serving/</link>
    <description>Recent content in LLM Serving on Hellokitty&#39;s Blog</description>
    <generator>Hugo</generator>
    <language>zh-CN</language>
    <lastBuildDate>Wed, 26 Aug 2026 13:50:00 +0800</lastBuildDate>
    <atom:link href="https://yangyang233333.github.io/tags/llm-serving/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>NVIDIA Dynamo 源码阅读（一）：数据中心级推理栈的整体架构</title>
      <link>https://yangyang233333.github.io/posts/nvidia-dynamo-architecture-overview/</link>
      <pubDate>Wed, 26 Aug 2026 13:50:00 +0800</pubDate>
      <guid>https://yangyang233333.github.io/posts/nvidia-dynamo-architecture-overview/</guid>
      <description>从请求路径、控制面、数据面和后端边界出发，理解 NVIDIA Dynamo 为什么是推理引擎之上的分布式编排层。</description>
    </item>
    <item>
      <title>Mini-SGLang 源码阅读（一）：一次 LLM 请求如何穿过推理引擎</title>
      <link>https://yangyang233333.github.io/posts/mini-sglang-source-reading-request-lifecycle/</link>
      <pubDate>Tue, 25 Aug 2026 20:30:00 +0800</pubDate>
      <guid>https://yangyang233333.github.io/posts/mini-sglang-source-reading-request-lifecycle/</guid>
      <description>从 Prefill、Decode、Continuous Batching 和多进程流水线出发，梳理 Mini-SGLang 一次请求的完整执行流程。</description>
    </item>
  </channel>
</rss>
