<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>AIEII</title><link>https://aieii.com/tags/%E6%8E%A8%E7%90%86%E6%80%A7%E8%83%BD/</link><description>来自 AIEII 传播部门的最新通讯</description><generator>Hugo -- gohugo.io</generator><language>zh-cn</language><managingEditor>newsletter@aieii.com (Giorgio)</managingEditor><lastBuildDate>Sun, 16 Aug 2026 19:17:51 +0800</lastBuildDate><atom:link href="https://aieii.com/tags/%E6%8E%A8%E7%90%86%E6%80%A7%E8%83%BD/index.xml" rel="self" type="application/rss+xml"/><item><title>vLLM 和 Ollama 怎么选：高并发场景实测对比</title><link>https://aieii.com/posts/2026-08-21-vllm-vs-ollama-concurrent-inference/</link><pubDate>Fri, 21 Aug 2026 09:00:00 +0800</pubDate><guid>https://aieii.com/posts/2026-08-21-vllm-vs-ollama-concurrent-inference/</guid><description>&lt;p&gt;先说结论：一个人用、或者团队里三五个人零星调用，Ollama 完全够用，装好就跑，不用纠结。但如果你要拿本地模型接一个对外服务，哪怕只是给公司内部几十号人用的工具，并发请求一多，Ollama 的排队延迟会立刻暴露，这时候该换 vLLM 了。&lt;/p&gt;</description></item></channel></rss>