Give Frontier Models Frontier Problems
If today’s strongest AI systems are becoming extraordinary engineers and researchers as claimed, their best demonstrations should tackle problems worth solving and expand what we know how to build. The models impressed me more than the demonstrations did Last week (i.e., early September 2026) might feel like a festival to many AI enthusiasts. Anthropic released Fable 5.1 and OpenAI began rolling out GPT-6 Astra. Both releases were framed around difficult work: coding, research, computer use, and long-running professional tasks. The release materials from Anthropic and OpenAI made a large capability claim, and the public response supplied a familiar kind of evidence almost immediately. The releases themselves also point toward more consequential work: Anthropic reports a higher-resolution map of part of Venus and GPU-kernel optimizations for biology models, while OpenAI emphasizes software engineering, science, and cybersecurity evaluations. Those claims deserve artifact-level scrutiny, but they are closer to the kind of demonstration I want to see. ...