Hacker NewsWednesday · September 2, 2026FREE

Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s

llmmacmemory-optimization

A Hacker News post titled "Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s" presents a project called SlotStream, available on GitHub at https://github.com/carloslfu/slotstream. The post claims that the project enables running a 104GB Qwen3.8-Flash-Next model on a Mac with 48GB of memory, achieving a speed of approximately 12 tokens per second. The exact model name "Qwen3.8-Flash-Next" and the speed "~12 tok/s" are explicitly stated in the title. The project appears to be a tool or technique for running large language models on hardware with less memory than the model size, likely through some form of memory optimization or streaming. The source text is limited to the title and URL, so no further details about the implementation, performance benchmarks, or user impact are provided.

// why it matters

This demonstrates a potential method for running large models on limited hardware, which could expand local AI capabilities for developers.

Sources

Primary · Hacker News
▸ Read original at github.com

Like this? Get the next digest.