The composite benchmark gap between a mid-tier LLM and the most expensive frontier model right now is about five points on quality indices — 0.75 versus 0.
DataFusion just pushed version 54.0.
Qualcomm is shipping 80 TOPS in the Snapdragon X2 Elite. AMD hit 60 with Ryzen AI 400.
Moonshot's latest model just took the #1 spot on LMArena's Frontend Code benchmark, beating Claude Fable 5 in blind developer testing. It clocked 93.
The largest open-weight model in history just shipped, and the number that matters most isn't 2.8 trillion.
Grok 4.5 dropped last week at 2 per million input tokens and 6 output.
Flagship phone benchmarks love to report peak throughput. Forty tokens per second on an iPhone 16 Pro running Qwen 2.
Tencent dropped Hy3 on July 6 with a claim that sounded like a typo: a 295B MoE model with 21B active parameters beating DeepSeek V4 Pro — a 1.
A two-word change in your prompt drops accuracy by fifteen points. You've seen the claims.
SpaceXAI's newest model landed on Tuesday, and the discourse immediately split into two camps that aren't talking to each other.
OpenAI flipped the switch on GPT-5.6 yesterday.
OpenAI's GPT-5.6 trio — Sol, Terra, and Luna — went public yesterday after two weeks of government-gated preview.
Transformers.
Semgrep published their IDOR detection benchmark results last week, and the headline number stopped a few people mid-scroll: GLM-5.
Open-weight models have been trailing proprietary ones on real-world coding tasks for over a year now.
Two days ago at Build, Microsoft did something it's never done before: announced a full family of homegrown AI models that compete directly with the...
Two days ago MiniMax dropped a launch announcement for M3 that reads like a greatest-hits compilation of everything the open-weight community has been asking...
You're building a content moderation pipeline. The stakes are high — wrong calls mean either letting harmful content through or censoring legitimate speech.
A Miami startup with 11 researchers just dropped claims that, if true, would reshape how every LLM handles long sequences.
Endor Labs put GPT-5.5 through two agent harnesses last month — OpenAI's own Codex scaffold and Cursor's.