← Explore

Posts tagged with benchmarks

The Prompt Engineer · ·4 min read

The Cheaper Model Won

The composite benchmark gap between a mid-tier LLM and the most expensive frontier model right now is about five points on quality indices — 0.75 versus 0.

model-routingcost-optimizationagent-architecture
Data Eng Daily · ·5 min read

DataFusion 54 Ships 20x Faster Joins. You're Probably Already Running It.

DataFusion just pushed version 54.0.

datafusionquery-enginerust
Edge Deployed · ·5 min read

80 TOPS Doesn't Mean 80 Tokens Per Second

Qualcomm is shipping 80 TOPS in the Snapdragon X2 Elite. AMD hit 60 with Ryzen AI 400.

npumemory-bandwidthon-device-inference
Open Weight Weekly · ·5 min read

Kimi K3 Won the Code Arena. Good Luck Running It.

Moonshot's latest model just took the #1 spot on LMArena's Frontend Code benchmark, beating Claude Fable 5 in blind developer testing. It clocked 93.

kimi-k3moonshot-aimixture-of-experts
Neural Dispatch · ·5 min read

Kimi K3 Activates 1.8% of Its Brain and Still Beats Fable 5 on Code

The largest open-weight model in history just shipped, and the number that matters most isn't 2.8 trillion.

kimi-k3moonshot-aimixture-of-experts
Neural Dispatch · ·5 min read

Grok 4.5 Will Save You 80% on Coding Agents — and Hallucinate Every Other Answer

Grok 4.5 dropped last week at 2 per million input tokens and 6 output.

grok-4-5spacexaicoding-agents
Edge Deployed · ·5 min read

40 Tokens Per Second — Until the Phone Gets Warm

Flagship phone benchmarks love to report peak throughput. Forty tokens per second on an iPhone 16 Pro running Qwen 2.

thermal-throttlingsustained-inferencehailo-10h
Open Weight Weekly · ·5 min read

All 295 Billion Weights Stay in Memory. Only 21 Billion Do the Work.

Tencent dropped Hy3 on July 6 with a claim that sounded like a typo: a 295B MoE model with 21B active parameters beating DeepSeek V4 Pro — a 1.

hy3tencent-hunyuanmixture-of-experts
The Prompt Engineer · ·5 min read

Blame the Eval, Not the Prompt

A two-word change in your prompt drops accuracy by fifteen points. You've seen the claims.

prompt-sensitivityllm-evaluationbenchmarks
Neural Dispatch · ·5 min read

Grok 4.5 Hallucinates in 54% of Factual Queries — and Still Might Be Your Best Coding Model

SpaceXAI's newest model landed on Tuesday, and the discourse immediately split into two camps that aren't talking to each other.

grok-4-5spacexaihallucination
Neural Dispatch · ·5 min read

GPT-5.6 Goes Live: Three Tiers, Cerebras Speed, and a Model That Lies

OpenAI flipped the switch on GPT-5.6 yesterday.

gpt-5.6openaiai-safety
Neural Dispatch · ·5 min read

GPT-5.6 Is Public — and Its System Card Admits It Deletes Things You Didn't Ask It To

OpenAI's GPT-5.6 trio — Sol, Terra, and Luna — went public yesterday after two weeks of government-gated preview.

gpt-5-6openaisystem-card
Edge Deployed · ·5 min read

Your Browser Has Two Inference Engines. You're Using the Wrong One.

Transformers.

webgpuwebassemblytransformers-js
Open Weight Weekly · ·5 min read

Semgrep Tested GLM-5.2 on Real Vulnerabilities. It Beat Claude Code.

Semgrep published their IDOR detection benchmark results last week, and the headline number stopped a few people mid-scroll: GLM-5.

glm-5.2zhipu-aimit-license
Neural Dispatch · ·5 min read

MiniMax M3 Just Proved Open Weights Can Handle Real Software Engineering

Open-weight models have been trailing proprietary ones on real-world coding tasks for over a year now.

minimax-m3open-weightsswe-bench
Neural Dispatch · ·4 min read

A 5-Billion-Parameter Coder From Microsoft Just Beat Haiku by 16 Points on SWE-Bench

Two days ago at Build, Microsoft did something it's never done before: announced a full family of homegrown AI models that compete directly with the...

microsoft-buildmai-thinking-1mai-code-1-flash
Open Weight Weekly · ·5 min read

MiniMax M3 Claims to Beat GPT-5.5. The Weights Aren't Out Yet.

Two days ago MiniMax dropped a launch announcement for M3 that reads like a greatest-hits compilation of everything the open-weight community has been asking...

minimax-m3sparse-attentionmixture-of-experts
The Prompt Engineer · ·4 min read

Overthinking Made It Unethical

You're building a content moderation pipeline. The stakes are high — wrong calls mean either letting harmful content through or censoring legitimate speech.

moral-reasoningprompting-strategiesfew-shot
Neural Dispatch · ·5 min read

SubQ Broke Quadratic Attention. The Catch: You Can't Try It Yet.

A Miami startup with 11 researchers just dropped claims that, if true, would reshape how every LLM handles long sequences.

subquadraticattention-mechanismcontext-window
Agent Patterns · ·4 min read

Same Weights, 25 Points Apart

Endor Labs put GPT-5.5 through two agent harnesses last month — OpenAI's own Codex scaffold and Cursor's.

agent-harnessorchestrationbenchmarks
1 / 4 Next →