← Explore

Posts tagged with mixture-of-experts

Neural Dispatch · ·5 min read

Inkling Doesn't Want to Be the Smartest Model. That's the Point.

Mira Murati's Thinking Machines Lab dropped Inkling last week with a line you almost never hear from a model release: "Inkling is not the strongest...

inklingthinking-machinesmira-murati
Open Weight Weekly · ·5 min read

Mistral Built a Theorem Prover. It Found Bugs Instead.

Mistral didn't set out to build a bug finder. They built Leanstral 1.

leanstralmistral-aiformal-verification
Neural Dispatch · ·5 min read

LongCat-2.0: The Stealth Model That Proved You Don't Need Nvidia to Go Frontier

For two months, a model called "Owl Alpha" sat near the top of OpenRouter's leaderboards. First place on Hermes Agent by call volume.

longcat-2meituanchinese-chips
Open Weight Weekly · ·5 min read

Kimi K3 Won the Code Arena. Good Luck Running It.

Moonshot's latest model just took the #1 spot on LMArena's Frontend Code benchmark, beating Claude Fable 5 in blind developer testing. It clocked 93.

kimi-k3moonshot-aimixture-of-experts
Neural Dispatch · ·5 min read

Kimi K3 Activates 1.8% of Its Brain and Still Beats Fable 5 on Code

The largest open-weight model in history just shipped, and the number that matters most isn't 2.8 trillion.

kimi-k3moonshot-aimixture-of-experts
Open Weight Weekly · ·5 min read

Murati's Lab Just Shipped 975 Billion Weights. The Pitch Isn't "We're Smarter."

Mira Murati left OpenAI, raised $2 billion at seed, and the first thing her lab ships is a model that openly admits it isn't the best at everything.

inklingthinking-machinesmixture-of-experts
Neural Dispatch · ·5 min read

Copilot's First Open-Weight Model Is a Trillion-Parameter MoE Built in Beijing

Open your VS Code model picker today and you'll see a new name sitting next to Claude Sonnet 5 and GPT-5.6 Terra: Kimi K2.

github-copilotkimi-k2-7moonshot-ai
Open Weight Weekly · ·5 min read

All 295 Billion Weights Stay in Memory. Only 21 Billion Do the Work.

Tencent dropped Hy3 on July 6 with a claim that sounded like a typo: a 295B MoE model with 21B active parameters beating DeepSeek V4 Pro — a 1.

hy3tencent-hunyuanmixture-of-experts
Open Weight Weekly · ·5 min read

K2.7 Code Scores Near GPT-5.5 by Thinking Less, Not More

Moonshot AI dropped K2.7 Code on June 12 and it landed with a strange flex: on Kimi Code Bench v2, it scores 62.

kimi-k2.7-codemoonshot-aimixture-of-experts
Neural Dispatch · ·5 min read

Grok 4.5 Hallucinates in 54% of Factual Queries — and Still Might Be Your Best Coding Model

SpaceXAI's newest model landed on Tuesday, and the discourse immediately split into two camps that aren't talking to each other.

grok-4-5spacexaihallucination
Open Weight Weekly · ·5 min read

Qwen3-Coder-Next Runs Like a 3B Model and Codes Like a 70B

Alibaba's Qwen team has a running habit of shipping models that make you question your assumptions about parameter counts.

qwen3-coder-nextalibabamixture-of-experts
Neural Dispatch · ·4 min read

The $4.40 Model That Codes Better Than GPT-5.5

Z.ai dropped GLM-5.

glm-5-2z-aiopen-weights
Neural Dispatch · ·5 min read

Meituan's LongCat-2.0 Was Secretly Winning OpenRouter — on Zero Nvidia Hardware

For two months, a model called "Owl Alpha" quietly climbed the OpenRouter charts. No launch event, no press tour, no Twitter hype cycle.

longcat-2meituanopen-source
Neural Dispatch · ·4 min read

A 5-Billion-Parameter Coder From Microsoft Just Beat Haiku by 16 Points on SWE-Bench

Two days ago at Build, Microsoft did something it's never done before: announced a full family of homegrown AI models that compete directly with the...

microsoft-buildmai-thinking-1mai-code-1-flash
Open Weight Weekly · ·5 min read

MiniMax M3 Claims to Beat GPT-5.5. The Weights Aren't Out Yet.

Two days ago MiniMax dropped a launch announcement for M3 that reads like a greatest-hits compilation of everything the open-weight community has been asking...

minimax-m3sparse-attentionmixture-of-experts
Neural Dispatch · ·5 min read

Project Polaris: GitHub Built Its Own Coding Model to Replace OpenAI Inside Copilot

The biggest announcement from Microsoft Build 2026 wasn't a new Azure service or another Windows agent feature. It was a quiet decoupling.

project-polarisgithub-copilotmicrosoft-build
Open Weight Weekly · ·5 min read

Poolside Spent $626M Training in Secret. Their Open Model Runs on a MacBook.

Poolside raised $626 million, hired a team of ex-DeepMind and ex-Meta researchers, and then went quiet for nearly three years.

laguna-xs2poolsidemixture-of-experts
Neural Dispatch · ·5 min read

How 760 Million Active Parameters Outscored Claude on HMMT — Without a Single NVIDIA GPU

Zyphra just dropped a model that makes me question everything I assumed about parameter efficiency. ZAYA1-8B is an 8.

open-sourcemixture-of-expertsamd
Open Weight Weekly · ·5 min read

LLaDA 2.0 Doesn't Predict the Next Token. It's Faster Than Models That Do.

Every large language model you've touched — GPT, Claude, Llama, Qwen, Gemma — generates text the same fundamental way. Predict the next token.

llada-2discrete-diffusionant-group
Neural Dispatch · ·4 min read

DeepSeek V4 Pro Costs 13x Less Than GPT-5.5. NIST Measured the Gap.

DeepSeek just made a move that'll force some uncomfortable spreadsheet conversations. The 75% promotional discount on V4 Pro?

deepseekopen-weightsnist
1 / 3 Next →