A friend asked me to host Qwen 3.8 Flash Next for him: a 125-billion-parameter model, on a machine with one RTX 4070 (12 GB). So I tried Strata, the free engine people are calling "6× faster than llama.cpp". It works. Writing ran at about 35 tokens a second against llama.cpp's 15, and a 100,000-token codebase was read in 2 min 22 s instead of 24 minutes. Then the logs showed something odd: the graphics card was barely trying. This video explains why. The same reason makes this setup brilliant for one person and no faster than llama.cpp once a second person joins. CHAPTERS 0:00 A 125B model on one 12 GB card 1:16 The request 2:18 How Strata splits the model 4:07 The machine 5:28 One person, one chat 6:39 The long read 8:41 The card that was barely trying 9:56 Two's a crowd 11:05 Verdict THE RESULTS (same model file, same card, both engines) • Short chats: Strata 32.5–38.1 tok/s vs llama.cpp 15.3–16.5 tok/s (stock llama.cpp, no MTP draft) • Reading a 100,000-token C++ codebase: Strata 2 min 22 s vs llama.cpp 24 min (~71 tok/s) • Writing after that 100K read: Strata 47 tok/s (373 of 404 drafted tokens accepted on code) vs llama.cpp 15.5 • 235,000-token document: read in 5 min 5 s; hidden passphrase found. Follow-up question: 11.5 s (229,376 tokens reused from cache, 6,054 new) • Writing as the chat grows: 30.0 tok/s (~1K) → 25.7 (120K) → 24.6 (235K) • GPU power: average 93 W of a 200 W limit (peak 117 W), clocks at full speed (~2,865 MHz) • Where the needed experts ran (single user): 58% already on the card, 4% copied to it, 38% on the CPUs • Two users (parallel on): ~18 tok/s combined vs llama.cpp ~19. Four users: 24.2 vs 23.2 • Two users each pasting ~90,000-token documents at once: 23 min 45 s • Four short chats: 66.5 s in parallel, 49.8 s one after another THE SETUP • GPU: NVIDIA RTX 4070, 12 GB • CPU: 2× Intel Xeon E5-2680 v4 (2016) • RAM: ~170 GB available (the model's experts need ~50 GB; 64 GB total is the practical minimum for this size) • Model: OrcaRouter's Qwen 3.8 Flash Next Uncensored, IQ3_XXS GGUF (85 GB); n-gram table (~29 GB) read from the SSD • Engines: Strata 0.1.40 (262K context, MTP draft from the original Qwen model) vs llama.cpp (stock) • All numbers come from our own logged runs on 7 Oct 2026. WHAT IT COSTS Around 7 Oct 2026, a used RTX 4070 was about $600 and a 64 GB DDR4 kit about $490. RAM prices are moving fast, so check current listings. RAM BY MODEL SIZE (from Strata's docs) 32 GB: Coder version · 48 GB: 2-bit versions · 64 GB: 3-bit versions (used here) · 96 GB+: ~4-bit #LocalAI #Qwen #Strata #LocalLLM #llamacpp #RTX4070 #AI #LLM #NVIDIA #GPU #SelfHosted #HomeLab #OpenSourceAI #MoE #AIServer #strata #strataai
The information provided is not trading advice. kdj.com does not assume any responsibility for any investments made based on the information provided in this article. Cryptocurrencies are highly volatile and it is highly recommended that you invest with caution after thorough research!
If you believe that the content used on this website infringes your copyright, please contact us immediately (info@kdj.com) and we will delete it promptly.