Release: 2026/10/08 12:52 Reading: 0
Original author:Weights & Wonders
Original source:https://www.youtube.com/embed/8kzSGQzGigc
A friend asked me to host Qwen 3.8 Flash Next for him: a 125-billion-parameter model, on a machine with one RTX 4070 (12 GB). So I tried Strata, the free engine people are calling "6× faster than llama.cpp". It works. Writing ran at about 35 tokens a second against llama.cpp's 15, and a 100,000-token codebase was read in 2 min 22 s instead of 24 minutes. Then the logs showed something odd: the graphics card was barely trying. This video explains why. The same reason makes this setup brilliant for one person and no faster than llama.cpp once a second person joins. CHAPTERS 0:00 A 125B model on one 12 GB card 1:16 The request 2:18 How Strata splits the model 4:07 The machine 5:28 One person, one chat 6:39 The long read 8:41 The card that was barely trying 9:56 Two's a crowd 11:05 Verdict THE RESULTS (same model file, same card, both engines) • Short chats: Strata 32.5–38.1 tok/s vs llama.cpp 15.3–16.5 tok/s (stock llama.cpp, no MTP draft) • Reading a 100,000-token C++ codebase: Strata 2 min 22 s vs llama.cpp 24 min (~71 tok/s) • Writing after that 100K read: Strata 47 tok/s (373 of 404 drafted tokens accepted on code) vs llama.cpp 15.5 • 235,000-token document: read in 5 min 5 s; hidden passphrase found. Follow-up question: 11.5 s (229,376 tokens reused from cache, 6,054 new) • Writing as the chat grows: 30.0 tok/s (~1K) → 25.7 (120K) → 24.6 (235K) • GPU power: average 93 W of a 200 W limit (peak 117 W), clocks at full speed (~2,865 MHz) • Where the needed experts ran (single user): 58% already on the card, 4% copied to it, 38% on the CPUs • Two users (parallel on): ~18 tok/s combined vs llama.cpp ~19. Four users: 24.2 vs 23.2 • Two users each pasting ~90,000-token documents at once: 23 min 45 s • Four short chats: 66.5 s in parallel, 49.8 s one after another THE SETUP • GPU: NVIDIA RTX 4070, 12 GB • CPU: 2× Intel Xeon E5-2680 v4 (2016) • RAM: ~170 GB available (the model's experts need ~50 GB; 64 GB total is the practical minimum for this size) • Model: OrcaRouter's Qwen 3.8 Flash Next Uncensored, IQ3_XXS GGUF (85 GB); n-gram table (~29 GB) read from the SSD • Engines: Strata 0.1.40 (262K context, MTP draft from the original Qwen model) vs llama.cpp (stock) • All numbers come from our own logged runs on 7 Oct 2026. WHAT IT COSTS Around 7 Oct 2026, a used RTX 4070 was about $600 and a 64 GB DDR4 kit about $490. RAM prices are moving fast, so check current listings. RAM BY MODEL SIZE (from Strata's docs) 32 GB: Coder version · 48 GB: 2-bit versions · 64 GB: 3-bit versions (used here) · 96 GB+: ~4-bit #LocalAI #Qwen #Strata #LocalLLM #llamacpp #RTX4070 #AI #LLM #NVIDIA #GPU #SelfHosted #HomeLab #OpenSourceAI #MoE #AIServer #strata #strataai
Crypto Angler
2026-10-08 15:48
Weights & Wonders
2026-10-08 15:48
Pasteur Bernard Fwamba Mirec
2026-10-08 15:48
3.0 TV
2026-10-08 15:48
GemBoy Crypto
2026-10-08 15:48
区块拿铁
2026-10-08 14:35
Fuma Boy'Z
2026-10-08 14:15
The Crypto Report
2026-10-08 13:56
Crypto Việt Nam
2026-10-08 13:19
Select Currency
US Dollar
USD
Chinese Yuan
CNY
Japanese Yen
JPY
South Korean Won
KRW
New Taiwan Dollar
TWD
Canadian Dollar
CAD
Euro
EUR
Pound Sterling
GBP
Danish Krone
DKK
Hong Kong Dollar
HKD
Australian Dollar
AUD
Brazilian Real
BRL
Swiss Franc
CHF
Chilean Peso
CLP
Czech Koruna KČ
CZK
Singapore Dollar
SGD
Indian Rupee
INR
Saudi Riyal
SAR
Vietnamese Dong
VND
Thai Baht
THB
Select Currency
US Dollar
USD-$
Chinese Yuan
CNY-¥
Japanese Yen
JPY-¥
South Korean Won
KRW -₩
New Taiwan Dollar
TWD-NT$
Canadian Dollar
CAD-$
Euro
EUR - €
Pound Sterling
GBP-£
Danish Krone
DKK-KR
Hong Kong Dollar
HKD- $
Australian Dollar
AUD-$
Brazilian Real
BRL -R$
Swiss Franc
CHF -FR
Chilean Peso
CLP-$
Czech Koruna KČ
CZK -KČ
Singapore Dollar
SGD-S$
Indian Rupee
INR -₹
Saudi Riyal
SAR -SAR
Vietnamese Dong
VND-₫
Thai Baht
THB -฿