Why Two Leaderboards Rank The Same Model Differently
One board has a model first, another has it fifth, and both are honest. The four reasons scores diverge, and which board to trust for writing work.
Read article →Topic
41 articles tagged AI Models. For the wider topic, see AI Models & News.
One board has a model first, another has it fifth, and both are honest. The four reasons scores diverge, and which board to trust for writing work.
Read article →GPT-5.6, Gemini 3.6, Claude Sonnet 5, GLM-5.2. The numbers are marketing, not measurement. How to work out whether a release actually matters to you.
Read article →A month of model launches, price changes, and a regulatory deadline. What genuinely changes if your job involves writing emails and documents, and what is noise.
Read article →A mid-tier model gets 50 percent more expensive overnight when a promotion expires. What that says about planning AI spend, and how to build a setup that survives it.
Read article →Meta shipped Muse Spark 1.2 in August 2026 behind an API and has not released a new open Llama in over a year. Who carries open weights now, and what it changes for you.
Read article →An 80 percent cut on the cheapest tier and 20 percent on the middle one. What falling token prices actually change for work writing, and what they do not.
Read article →Alibaba's largest model yet lands in August 2026 with open weights promised. Here is what the specifications mean and why your rewrite still runs on 1.7 billion parameters.
Read article →Three new Gemini models shipped in July 2026, all in the cheap and fast tier, none a new Pro. Why the workhorse tier is where the useful work now happens.
Read article →Providers now market doing more with fewer tokens rather than higher scores. Why that shift happened, and why it is better news for writing than for anything else.
Read article →DeepSeek shipped an MIT-licensed 284B mixture-of-experts model on 31 July 2026. What the license actually gives you, and why cheap hosting is the real story.
Read article →Microsoft announced an on-device small model family for Windows, naming rewriting as a use case. What that concession signals about where writing tools are heading.
Read article →The most dangerous rewrite is the fluent one where a figure moved. How to build a checking habit that survives model upgrades, deprecations, and silent version swaps.
Read article →Models accept a million tokens and attend well to far fewer. Where quality drops, why the middle of a long input is the danger zone, and how to work around it.
Read article →Nearly every large model in 2026 is a mixture of experts. What that means, why the two parameter counts matter differently, and the planning mistake it causes.
Read article →Small models handle some languages far better than others. How to test coverage on your own text, and which families are worth trying first.
Read article →Two pricing mechanisms that cut cloud AI costs substantially, when each applies, and why neither helps the interactive rewrite you are waiting on.
Read article →Research and economics both point the same way: most agent steps do not need a frontier model. What that means for anyone choosing tools in 2026.
Read article →Reasoning models deliberate before answering. For a tone change that is pure overhead. How to tell which mode you are in, and how to turn it off.
Read article →Prices fell roughly 80 percent in a year and open models sit within a few points of proprietary ones. When the model is a commodity, the durable decisions are about everything else.
Read article →Anthropic shipped four Claude 5 models in under two months. AI has moved from blockbuster launches to continuous iteration, and that changes how you should build workflows on top of it.
Read article →A practical shortlist of open-weights models that actually help with professional writing, sorted by what hardware you have rather than by benchmark score.
Read article →DeepSeek, Qwen, Kimi, GLM, MiniMax, and Hunyuan now define the open-weights frontier and undercut Western pricing dramatically. The capability story, and the procurement questions it raises.
Read article →Anthropic shipped Claude Opus 5 in July 2026 with a million-token context, adaptive thinking, and effort levels. Which of those features matter for work writing, and which are for other jobs.
Read article →DeepSeek V4 competes on inference efficiency rather than raw benchmark scores. Why that is the more interesting strategy, and what it means for anyone paying for AI by the token.
Read article →The best model in the world and a 1.7B model on your laptop, given the same email. What actually differs, what does not, and why the expensive one sometimes loses.
Read article →Million-token windows are now standard on frontier models. What that capacity is actually for, why retrieval quality degrades long before the limit, and what professional writing really needs.
Read article →Gemma 4 moved to Apache 2.0 and ships sizes built for laptops rather than datacenters. What it is good at, where it fits against Qwen, and why on-device is the interesting category.
Read article →Zhipu's GLM line built its reputation on function calling and structured output rather than benchmark scores. Why that reliability is harder to achieve than raw capability, and where it matters.
Read article →OpenAI's GPT-5.6 ships as three tiers with different capability and cost. Which one you actually want for professional writing, and why the answer is usually not the flagship.
Read article →OpenAI ships gpt-oss under Apache 2.0 in 120B and 20B sizes. Where they fit, why a reasoning-oriented model is a mixed blessing for rewriting, and how they compare to the small Qwen and Gemma models.
Read article →Four frontier models in two months and a new open-weights release most weeks. A filtering system for staying current on AI without turning it into a second job.
Read article →Model cards and system cards are the closest thing to a datasheet AI has. What to look for, what the omissions tell you, and why this is becoming a compliance document.
Read article →Thinking Machines Lab shipped its first open-weights model in July 2026. Why a US entry matters in a field that had tilted heavily toward Chinese labs.
Read article →Moonshot AI released Kimi K3 as a 2.8-trillion-parameter open-weights model in July 2026. What that actually means, who can run it, and why it changes the field even if you never touch it.
Read article →MiniMax M3 handles a million-token context at a fraction of standard transformer compute. How sparse attention works in plain terms, and why long context is cheaper but not better.
Read article →NVIDIA keeps shipping compressed open-weight models derived from larger ones. How pruning and distillation work, and why compressed models are the ones that reach your laptop.
Read article →SWE-bench, MMLU-Pro, and AIME scores predict almost nothing about whether a model will rewrite your email well. What to measure instead.
Read article →Both labs ship excellent models. For rewriting work messages the differences that matter are not the ones the benchmarks measure. A practical comparison, including where neither is the answer.
Read article →Alibaba ships Qwen models faster than anyone can track. A guide to which sizes matter, what the naming means, and why the small ones are the important ones.
Read article →Hunyuan 3.0 shipped open weights in July 2026 with three selectable inference modes. Why letting the user choose how hard the model thinks is the most practical feature of 2026.
Read article →Every hosted model you build a habit on will eventually be retired. How to notice early, migrate without breaking your prompts, and keep one option that cannot be taken away.
Read article →