Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
A Hacker News item published on 2026-10-04 points to the GitHub repository Niko1221/Strata. The submission title states that Qwen 3.8 Flash Next, described there as a 125B model, can run on consumer hardware, naming the RTX 4090, at 100T/s. No repository documentation, benchmark configuration, quantization scheme, or measurement methodology appears in the supplied source text, which instead consists of a GitHub feature-flag payload listing internal flags. Because the excerpt contains no README, release notes, or performance data, the throughput and hardware claims cannot be verified from the source itself; they are reproduced here only as they appear in the submission title. The digest is limited to what the source states: a repository link, a model name and parameter count as given in the title, a named consumer GPU, and a tokens-per-second figure. No information is available in the source about memory requirements, context length, licensing, or how the reported speed was obtained.
The submission claims a 125B model runs on a single consumer GPU, but the source provides no benchmark details to verify it.