I'm (mostly) picking models on speed now, not intelligence
Martin Alderson, in a blog post, describes a shift in his model selection criteria: he now prioritizes speed over raw intelligence. He states that models around the 'Opus 4.6 level' are 'smart enough' for most daily tasks, including coding, research, and analytical work. He mentions that while he was initially excited about Fable, the US Gov shutdown gave time to get used to Opus again, and when Fable returned with additional guardrails, it felt slow, leading him to switch back to Opus. He argues that speed is crucial for user experience, citing that even a basic product feels good if fast. He suggests that 100-200 tokens per second (tok/s) feels fast, while below 50 tok/s feels slow, and notes that going past 200 tok/s can feel unnerving. He highlights that open-weights models like GLM5.2 and DeepSeek V4 Flash GA are clearing the intelligence bar and offer a range of speeds, with GLM5.2 on OpenRouter varying from less than 30 tok/s to 129 tok/s. He also discusses limitations: as models get faster, bottlenecks shift to tool calls and human oversight, and he gives a rough example where a 5x speedup on the model only yields a 2x speedup on the turn. He predicts a price war, noting that OpenAI reduced the cost of its Luna variant by 80% before the DeepSeek V4 Flash GA release, and that GLM5.2 pricing on OpenRouter has dropped to $0.42/$1.32 per MTok, which is 5% of Opus's price.
Developers may increasingly choose faster models over smarter ones, and open-weights models are driving price competition.