< Back to Blog

Qwen Flash-Next vs Laguna at 200K: 47s vs 455s

Qwen3.8-Flash-Next vs Laguna-S-2.1 on a Mac Studio M5 Ultra: same-day 8K to 200K speed ladder plus a hard quality eval. Speed was a sweep, quality a draw.

Qwen Flash-Next vs Laguna at 200K: 47s vs 455s

My first run said 200,000 tokens prefill in 21 seconds. That number was a lie my own harness told me, and catching it was the most useful part of the day.

Back in June, Claude Code and I benchmarked five ways to serve Qwen on a Mac Studio, and the wrong first answer turned into the best finding. This is the sequel. Same shape: a careful test, a wrong first answer, and a correction that mattered more than the raw numbers. This time the question was a head-to-head. Qwen3.8-Flash-Next against Laguna-S-2.1, same machine, same day, same prompts, 8K all the way to 200K.

Test rig: agent harness on Mac Mini driving oMLX and LM Studio lanes on a Mac Studio M5 Ultra

The setup, and who actually built it

The machine is a Mac Studio M5 Ultra with 256GB of unified memory. Qwen3.8-Flash-Next (182GB, oQ8e quant with multi-token prediction on) served through oMLX. Laguna-S-2.1 GGUF (128GB, full GPU offload) served through LM Studio. Both capped at 262,144 tokens of context. That cap was my call. Laguna can address a million tokens and Qwen cannot, so an uncapped test would compare context windows instead of models.

Straightforward about authorship, same as last time. The agent wrote the benchmark harness and drove every run over SSH. I set direction, picked the tests, and overruled the agent three times: the 262K cap, testing Laguna through LM Studio instead of the oMLX build it first reached for, and one rule that shaped everything below. No trading quality for speed. We capture speed only where quality stays fixed.

Test 1: the prefill ladder

Five prompt sizes, 8K to 200K tokens, temp 0, 160-token generations, unique salted content per run. The salting matters. My first attempt reused a shared prefix across sizes, the KV cache carried over between runs, and 200K "prefilled" in 21 seconds. The server's own cached_tokens counter exposed it. Round two used unique content per run with zero cached tokens verified on every request. Those are the numbers below.

Test TTFT Qwen / Laguna Decode Qwen / Laguna End to end
8K 2.0s / 10.3s 65 / 68 tok/s 4.6s / 12.6s
32K 7.4s / 30.4s 74 / 56 tok/s 9.6s / 33.3s
64K 14.7s / 70.8s 74 / 53 tok/s 17.0s / 73.7s
131K 30.1s / 217.2s 59 / 42 tok/s 33.1s / 221.0s
200K 47.1s / 455.4s 65 / 34 tok/s 49.9s / 460.1s

Time to first token from 8K to 200K prompts, log scale: Qwen holds near linear while Laguna climbs superlinear to 455 seconds

Qwen prefills at a near constant 4,200 tokens per second across all five sizes. Laguna degrades superlinearly, which is exactly what quadratic attention predicts and exactly what my last test showed at 64K and up (70.8s vs 68.3s, 217.2s vs 212.6s, 455.4s vs 455.6s, reproduced almost to the second). At 200K, prefill is 94 percent of total time on both models. Past 100K, prefill is the only benchmark that matters, and it is not close.

Memory favored Qwen too. Its engine footprint held flat at 144GB while system wide usage ran 172 to 184GB across the ladder. Laguna at the 262K cap sat near 202GB system wide, most of it preallocated KV cache.

Test 2: decode and the MTP receipt

Qwen shipped with multi-token prediction enabled at depth 3, so I checked the server logs instead of trusting the headline rate. The logs keep receipts:

oMLX server log lines showing MTP acceptance between 70 and 76 percent and about 2x tokens per cycle

Acceptance between 70 and 76 percent, roughly 2 tokens per cycle. That is a genuine doubling of decode, and the scheduler is adaptive about it. It parks speculation when acceptance drops and re-probes later, and nearly all accepts land at depth 1, so raising the depth buys nothing. My earlier Qwen test ran without MTP, which explains most of the decode gap between then and now.

Decode rates across prompt sizes: Qwen holds 59 to 74 tok/s while Laguna fades from 68 to 34

Laguna has no speculation here, deliberately. Its DFlash draft path lost to plain decoding on this exact hardware last year, a closed decision I am not reopening. Plain decode fades from 68 to 34 tok/s as context grows. Nothing wrong with it. Just physics.

Test 3: the optimizations that failed

With quality held fixed, I tried the two remaining speed knobs: speculative prefill and aggressive burst decode. Protocol was strict. Five fixed probes, greedy decoding, full output hashes before and after. Both changes came back bit identical across all five probes. Bit exact, and zero measurable speed change on non streamed requests. I reverted both.

That is a finding, not a failure. A toggle that changes nothing and risks confusion should not stay on. The MTP depth knob got the same treatment from the logs instead of a test. Already optimal, left alone. Net after the whole sweep: this model is at its quality neutral ceiling for single request latency. Everything left either costs quality or only helps multi client throughput.

Test 4: the quality draw

Speed means nothing if the answers differ, so both models faced four problems with script verified ground truth: exact combinatorics (a 9 digit integer), a novel interval coverage spec with 12 hidden tests, a fresh knights and knaves puzzle with a unique solution, and an asyncio ordering trap. Thinking stayed on for both. Both scored 4 out of 4.

The interesting part is how differently they got there. Laguna answers fast and commits. Three of its four landed in 5 to 10 seconds with a few hundred tokens. Qwen deliberates exhaustively. Its asyncio answer took 119 seconds and 11,000 tokens to say what Laguna said in 6 seconds and 417. Both models also showed the same failure mode once each: burning an entire 8,000 token budget on hidden reasoning with zero visible answer, then converting cleanly on a 16,000 token retry. And the day's best moment belonged to Qwen. My own puzzle verifier disagreed with its answer. Rechecking proved the model right and my verifier buggy.

The decision

Qwen is now the default resident model at the 262K cap. Faster prefill at every size by 2 to 10x, faster decode, 20 to 30GB lighter, equal top end quality. Laguna stays staged on disk for three jobs: fallback if oMLX ever has a bad day, latency sensitive short task lanes where its decisiveness still wins, and anything that ever needs to go past 262K.

I believe most local LLM benchmarks are really benchmarking whose defaults nobody checked. Cache state, thinking budgets, context caps, speculation settings. The numbers are easy. The methodology calls are the whole game, and this time they were mine.