<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Inovello</title>
    <link>https://inovello.dev/</link>
    <description>Benchmark writeups and failure notes on running large mixture-of-experts models on two RTX 3090s and 192 GB of DDR4.</description>
    <language>en</language><lastBuildDate>Sat, 05 Sep 2026 00:00:00 &#43;0000</lastBuildDate>
    <atom:link href="https://inovello.dev/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>A silent -ot downgrade under mmap, and a 9-line loader fix</title>
      <link>https://inovello.dev/writeups/pinned-host-experts-under-mmap/</link>
      <pubDate>Sat, 05 Sep 2026 00:00:00 &#43;0000</pubDate>
      <guid>https://inovello.dev/writeups/pinned-host-experts-under-mmap/</guid>
      <description>Why -ot ...=CUDA_Host did nothing on master, what my PR #28223 changes, and where the 8 minutes of load time went.</description>
    </item>
    <item>
      <title>Qwen3.8-Flash-Next on 2x3090 &#43; DDR4, part 2: 25-29 to 37-41 t/s with UD-Q4_K_XL, the expert cache, and MTP</title>
      <link>https://inovello.dev/writeups/qwen3-flash-next-2x3090-q4-mtp/</link>
      <pubDate>Fri, 04 Sep 2026 00:00:00 &#43;0000</pubDate>
      <guid>https://inovello.dev/writeups/qwen3-flash-next-2x3090-q4-mtp/</guid>
      <description>A quant swap, MTP on top of the cache, a 4x faster load, a bug in the cache PR, and the discovery that my RAM had been thermal throttling the whole time.</description>
    </item>
    <item>
      <title>Qwen3.8-Flash-Next on 2x3090 &#43; DDR4: 17 to 25-29 t/s decode with the expert cache PR</title>
      <link>https://inovello.dev/writeups/qwen3-flash-next-2x3090-expert-cache/</link>
      <pubDate>Thu, 03 Sep 2026 00:00:00 &#43;0000</pubDate>
      <guid>https://inovello.dev/writeups/qwen3-flash-next-2x3090-expert-cache/</guid>
      <description>All 48 expert layers in host RAM, an LRU cache of hot experts in VRAM, and the VRAM budget that made it pay.</description>
    </item>
  </channel>
</rss>
