Why Two Leaderboards Rank The Same Model Differently
One board has a model first, another has it fifth, and both are honest. The four reasons scores diverge, and which board to trust for writing work.
Read article →Model releases arrive faster than anyone can evaluate them, and most announcements describe capabilities that matter enormously for building software agents and very little for writing a difficult email. These articles separate the two: what each release actually changes, which specifications are marketing, and how to keep a working setup without chasing every launch.
29 articles
One board has a model first, another has it fifth, and both are honest. The four reasons scores diverge, and which board to trust for writing work.
Read article →GPT-5.6, Gemini 3.6, Claude Sonnet 5, GLM-5.2. The numbers are marketing, not measurement. How to work out whether a release actually matters to you.
Read article →The internet decided one punctuation mark exposes machine writing. The measurements say otherwise, and the real tells are harder to spot and harder to fake.
Read article →A month of model launches, price changes, and a regulatory deadline. What genuinely changes if your job involves writing emails and documents, and what is noise.
Read article →A major model provider logged a multi-hour outage in early August 2026. What that means if your writing depends on a hosted model, and what a real fallback looks like.
Read article →A mid-tier model gets 50 percent more expensive overnight when a promotion expires. What that says about planning AI spend, and how to build a setup that survives it.
Read article →Meta shipped Muse Spark 1.2 in August 2026 behind an API and has not released a new open Llama in over a year. Who carries open weights now, and what it changes for you.
Read article →An 80 percent cut on the cheapest tier and 20 percent on the middle one. What falling token prices actually change for work writing, and what they do not.
Read article →Alibaba's largest model yet lands in August 2026 with open weights promised. Here is what the specifications mean and why your rewrite still runs on 1.7 billion parameters.
Read article →OpenAI shipped an enterprise agent in July 2026 that runs for hours and produces finished documents. What that changes for professional writing, and what it does not.
Read article →Three new Gemini models shipped in July 2026, all in the cheap and fast tier, none a new Pro. Why the workhorse tier is where the useful work now happens.
Read article →Providers now market doing more with fewer tokens rather than higher scores. Why that shift happened, and why it is better news for writing than for anything else.
Read article →Enterprises say they have adopted AI agents. Very few are running them in production. What the gap tells you about where AI actually delivers value today.
Read article →Agents that read your screen, your tabs, and your clipboard transmit far more than you intend. What that means for confidential work, and how to draw a boundary you can hold.
Read article →Detectors are still unreliable, watermarking is partial, and the professional norms have shifted. A current picture of what can and cannot be proven about AI-assisted text.
Read article →The most dangerous rewrite is the fluent one where a figure moved. How to build a checking habit that survives model upgrades, deprecations, and silent version swaps.
Read article →Models accept a million tokens and attend well to far fewer. Where quality drops, why the middle of a long input is the danger zone, and how to work around it.
Read article →Two pricing mechanisms that cut cloud AI costs substantially, when each applies, and why neither helps the interactive rewrite you are waiting on.
Read article →Data residency, jurisdictional control, and operational access are three different things. What sovereignty claims mean, and the configuration that settles all three.
Read article →Prices fell roughly 80 percent in a year and open models sit within a few points of proprietary ones. When the model is a commodity, the durable decisions are about everything else.
Read article →Anthropic shipped four Claude 5 models in under two months. AI has moved from blockbuster launches to continuous iteration, and that changes how you should build workflows on top of it.
Read article →DeepSeek, Qwen, Kimi, GLM, MiniMax, and Hunyuan now define the open-weights frontier and undercut Western pricing dramatically. The capability story, and the procurement questions it raises.
Read article →Anthropic shipped Claude Opus 5 in July 2026 with a million-token context, adaptive thinking, and effort levels. Which of those features matter for work writing, and which are for other jobs.
Read article →The best model in the world and a 1.7B model on your laptop, given the same email. What actually differs, what does not, and why the expensive one sometimes loses.
Read article →Million-token windows are now standard on frontier models. What that capacity is actually for, why retrieval quality degrades long before the limit, and what professional writing really needs.
Read article →OpenAI's GPT-5.6 ships as three tiers with different capability and cost. Which one you actually want for professional writing, and why the answer is usually not the flagship.
Read article →Four frontier models in two months and a new open-weights release most weeks. A filtering system for staying current on AI without turning it into a second job.
Read article →Both labs ship excellent models. For rewriting work messages the differences that matter are not the ones the benchmarks measure. A practical comparison, including where neither is the answer.
Read article →Every hosted model you build a habit on will eventually be retired. How to notice early, migrate without breaking your prompts, and keep one option that cannot be taken away.
Read article →