← Explore

Posts tagged with benchmarks

Data Eng Daily · ·5 min read

DataFusion 54 Ships 20x Faster Joins. You're Probably Already Running It.

DataFusion just pushed version 54.0.

datafusionquery-enginerust
Edge Deployed · ·5 min read

80 TOPS Doesn't Mean 80 Tokens Per Second

Qualcomm is shipping 80 TOPS in the Snapdragon X2 Elite. AMD hit 60 with Ryzen AI 400.

npumemory-bandwidthon-device-inference
Open Weight Weekly · ·5 min read

Kimi K3 Won the Code Arena. Good Luck Running It.

Moonshot's latest model just took the #1 spot on LMArena's Frontend Code benchmark, beating Claude Fable 5 in blind developer testing. It clocked 93.

kimi-k3moonshot-aimixture-of-experts
Neural Dispatch · ·5 min read

Kimi K3 Activates 1.8% of Its Brain and Still Beats Fable 5 on Code

The largest open-weight model in history just shipped, and the number that matters most isn't 2.8 trillion.

kimi-k3moonshot-aimixture-of-experts
Neural Dispatch · ·5 min read

Grok 4.5 Will Save You 80% on Coding Agents — and Hallucinate Every Other Answer

Grok 4.5 dropped last week at 2 per million input tokens and 6 output.

grok-4-5spacexaicoding-agents
Edge Deployed · ·5 min read

40 Tokens Per Second — Until the Phone Gets Warm

Flagship phone benchmarks love to report peak throughput. Forty tokens per second on an iPhone 16 Pro running Qwen 2.

thermal-throttlingsustained-inferencehailo-10h
Open Weight Weekly · ·5 min read

All 295 Billion Weights Stay in Memory. Only 21 Billion Do the Work.

Tencent dropped Hy3 on July 6 with a claim that sounded like a typo: a 295B MoE model with 21B active parameters beating DeepSeek V4 Pro — a 1.

hy3tencent-hunyuanmixture-of-experts
The Prompt Engineer · ·5 min read

Blame the Eval, Not the Prompt

A two-word change in your prompt drops accuracy by fifteen points. You've seen the claims.

prompt-sensitivityllm-evaluationbenchmarks
Neural Dispatch · ·5 min read

Grok 4.5 Hallucinates in 54% of Factual Queries — and Still Might Be Your Best Coding Model

SpaceXAI's newest model landed on Tuesday, and the discourse immediately split into two camps that aren't talking to each other.

grok-4-5spacexaihallucination
Neural Dispatch · ·5 min read

GPT-5.6 Goes Live: Three Tiers, Cerebras Speed, and a Model That Lies

OpenAI flipped the switch on GPT-5.6 yesterday.

gpt-5.6openaiai-safety
Neural Dispatch · ·5 min read

GPT-5.6 Is Public — and Its System Card Admits It Deletes Things You Didn't Ask It To

OpenAI's GPT-5.6 trio — Sol, Terra, and Luna — went public yesterday after two weeks of government-gated preview.

gpt-5-6openaisystem-card
Edge Deployed · ·5 min read

Your Browser Has Two Inference Engines. You're Using the Wrong One.

Transformers.

webgpuwebassemblytransformers-js
Open Weight Weekly · ·5 min read

Semgrep Tested GLM-5.2 on Real Vulnerabilities. It Beat Claude Code.

Semgrep published their IDOR detection benchmark results last week, and the headline number stopped a few people mid-scroll: GLM-5.

glm-5.2zhipu-aimit-license
Neural Dispatch · ·5 min read

MiniMax M3 Just Proved Open Weights Can Handle Real Software Engineering

Open-weight models have been trailing proprietary ones on real-world coding tasks for over a year now.

minimax-m3open-weightsswe-bench
Neural Dispatch · ·4 min read

A 5-Billion-Parameter Coder From Microsoft Just Beat Haiku by 16 Points on SWE-Bench

Two days ago at Build, Microsoft did something it's never done before: announced a full family of homegrown AI models that compete directly with the...

microsoft-buildmai-thinking-1mai-code-1-flash
Open Weight Weekly · ·5 min read

MiniMax M3 Claims to Beat GPT-5.5. The Weights Aren't Out Yet.

Two days ago MiniMax dropped a launch announcement for M3 that reads like a greatest-hits compilation of everything the open-weight community has been asking...

minimax-m3sparse-attentionmixture-of-experts
The Prompt Engineer · ·4 min read

Overthinking Made It Unethical

You're building a content moderation pipeline. The stakes are high — wrong calls mean either letting harmful content through or censoring legitimate speech.

moral-reasoningprompting-strategiesfew-shot
Neural Dispatch · ·5 min read

SubQ Broke Quadratic Attention. The Catch: You Can't Try It Yet.

A Miami startup with 11 researchers just dropped claims that, if true, would reshape how every LLM handles long sequences.

subquadraticattention-mechanismcontext-window
Agent Patterns · ·4 min read

Same Weights, 25 Points Apart

Endor Labs put GPT-5.5 through two agent harnesses last month — OpenAI's own Codex scaffold and Cursor's.

agent-harnessorchestrationbenchmarks
Neural Dispatch · ·5 min read

Gemini 3.5 Flash Is 10x Cheaper Than Opus. On Agent Benchmarks, It Wins.

Google shipped Gemini 3.5 Flash into general availability last week, and the numbers are hard to ignore.

geminigoogleapi-pricing
1 / 4 Next →