← All articles
02 · raceSep 9, 202611 min read5,146 views on X

The fastest decode does not mean the best - Ornith 1.5 35B vs BigBang-v1 - Final Test

Same prompt, same M2 Max, same harness. BigBang finished in 54 minutes, Ornith took 119. Then I read the code.

Ornith AI vs Endless FrontierOrnith-1.5-35B-A3B MLX 4-bitBigBang-v1 oQ4e-mtp
TL;DR
  • I gave the exact same prompt to two local models on the same M2 Max. The task: plan an Apple-style nested todo app, then build it.
  • BigBang-v1 finished in 54 minutes. Ornith-1.5 took 119. BigBang generated a third of the tokens and a seventh of the reasoning.
  • Interesting fact: both apps contain almost exactly the same amount of application code, about 1,100 lines each. Ornith's entire surplus went into tests (27 vs 8 unit) and a hand-rolled CSS design system (565 lines vs Tailwind).
  • Same browser-verification rule for both. Ornith substituted jsdom tests and disclosed the limit. BigBang installed Playwright and Chromium by itself and clicked through the app for real.
  • My verdict after voxel test and to do app: BigBang wins speed, efficiency, and verification fidelity. Ornith wins deliverable completeness, solid code, and process discipline. Both apps work. But Ornith made it feel like an actual app.

Some time ago I decided to start getting into local AI which I can actually run. Since that time I am downloading models, testing them with popular voxel type prompts as well as in actual app building scenarios.

My mission is to popularize small local LLMs by doing tests on the hardware that I’ve got - my M2 Max, 64 GB.

The two runners:

  • Ornith-1.5-35B-A3B-MLX-4bit. A 35B-total, 3B-active MoE text model, 19 GB in 4-bit. Plain autoregressive decoding.
  • BigBang-v1-oQ4e-mtp. A ~20.5 GB vision-language model. The -mtp suffix means it runs with MTP speculative decoding, which drafts several tokens per cycle and accepts about 86% of them.

The harness is Grok Build, the grok CLI agent, in the same config for both: sandbox off, yolo mode, a stacked set of built-in policies. Grok ships a <browser_verification> rule in its system prompt: you must verify UI work in a browser whenever browser tools are available. And it ships a default user rule with the fallback clause: if no browser tools exist, verify through the closest substitute and say what you could not verify. Both are compiled into grok itself and injected into every session, I did not write either of them. No browser tool was connected in either session.

Identical client sampling for both: temp 0.6, top_p 0.95, 32,768 max tokens, 262,144 context window. The prompt was identical:

I want to build a web app which will be a todo app, you can add items to the todo list and then delete them, each item should have the option to add subitems, no login so far needed, we need an option to save the todo lists and remove them as well, this needs to be visually looking like an apple product, nice rich and polished ui. prepare a full implementation plan, you can check internet for similar examples if needed. put the plan.md in the project root - plan first.

The only fairness caveat I need to flag up front: part of BigBang's speed advantage is its speculative decoding variant, not the base model. Ornith decoded the plain way.

Figure · interactiveWall clock, replayed

The numbers

All from the session store and the oMLX server logs, checked 2026-09-09. One methodological note before you read the tokens row: Ornith's reasoning is invisible to the server meter, and I estimate it at around 55K characters-worth of extra generation on top of the 133,819 metered tokens. BigBang's reasoning is much shorter and sits inside its metered numbers.

MetricOrnithBigBang
Wall clock119 min54 min
Model calls170123
Tool calls166124
Generated tokens (metered)133,81943,087
Reasoning (est.)~55K tok, unmetered~7.9K tok
Cumulative input18.36M, 97% cache-read7.05M, 95% cache-read
Peak context193,143 tok (73% of window)86,378 tok (32%)
Decode speed, mean / max43.9 / 78.9 tok/s51.2 / 108.6 tok/s
Effective end-to-end rate20.3 tok/s19.0 tok/s
TTFT, average27.2 s11.8 s
Tool failures30
User interactions1 question form + 1 typed message1 approve click

Two things worth mentioning here.

First, the effective end-to-end rate is nearly identical, and Ornith is actually slightly ahead. BigBang's raw decode is faster, but its replies are short, so each request is dominated by prefill, re-reading the whole prompt before it can answer. Ornith writes long, which amortizes prefill better.

Second, the context. Ornith let its context grow to 193K of 262K, and its decode speed decayed from about 79 tok/s at 16K of prompt down to about 29 tok/s near 190K. That is a 2.5x tax purely from context size, same weights, same machine. BigBang sat at 86K and kept flatter decay.

Figure · interactiveThe context tax

How the two runs actually went

Ornith, the careful engineer.

Turn 1, about 10 minutes. It explored the empty project, ran two real web searches, one on Things 3 and Apple's design language, one on glassmorphism CSS recipes, and the results visibly shaped its plan. Then, instead of guessing, it asked me a structured 3-question form: stack, storage, nesting model. I answered React + TS + Vite, localStorage, unlimited nesting. It wrote a 12.9 KB plan with actual UX decisions in it, like "checking a parent does NOT auto-check children, matches Reminders." One neat detail: plan mode only lets it write the session's own plan file, so it copied the plan to the project root with a terminal command to satisfy my literal instruction.

Turn 2, about 107 minutes. It wrote 22 files, and then spent roughly the last 40% of the session verifying and repairing. Six rounds of TypeScript errors with real error batches, unclosed JSX tags, type mismatches. About 13 test cycles. One debugging move I want to quote because it is a real engineer's move, I would say, didn’t expect that: when a reorderItem function signature would not reconcile, it backed up the file, injected a type-probe line to make the compiler print the actual signature, patched the call site, and restored the backup.

It also caught a UX bug by reading its own code, not by a test failing. The empty list state had no way to add the first item. It found that in its reasoning, fixed it, and moved on.

Final state: TypeScript clean, lint clean, 27 out of 27 tests green, production build ok, dev server launched and curl-checked. And because no browser tool existed, it disclosed exactly that, in almost the exact words of grok's own fallback rule.

BigBang, the sprinter.

Plan mode was quick and it actually didn’t use the opportunity to ask me any questions. I decided to accept the plan as is to see how it handles the task

Its first move was a fumble: 4 calls to an MCP tool-discovery tool, used as if it were a web search box. All empty. It self-corrected immediately and did two real web searches after. Its scaffold also failed twice, the interactive npm create vite prompt hanging in a non-terminal, so it just did npm init -y and installed the deps manually.

Then it wrote, a lot: 40 writes against only 3 re-reads of its own files. Ornith re-read its code 25 times. BigBang writes code and just trusts it.

It tackled testing in a different way. Its reasoning says it plainly: "I need to verify the app using browser tools... use a headless browser to verify." No browser tool was connected. So it installed its own. npx playwright install chromium, the Playwright test package, a real e2e suite, dev server in the background. The first e2e run failed on a stale process holding port 3000. It found the process with lsof, killed it, re-ran, and finished 6 out of 6 e2e green, including a mobile viewport test at 375x667 and a dark mode test. Then it killed its own background tasks on the way out.

Two models read the same browser rule and produced two opposite philosophies. One complied by substitute. One complied by acquiring the capability. I did not expect a 20 GB local model to solve a missing-tool problem by building the tool, and that is the single most impressive behavior in this whole comparison.

What the screens show

I had both apps screenshotted in the same four states: light mode, a second list, dark mode, and a 375x667 mobile viewport. Same states for both apps, so you can compare like with like.

Ornith app, light mode, as built. Two-pane shell: sidebar with list management and per-list progress counts, content pane with nested tasks, progress ring, export and import icons.
Ornith app, light mode, as built. Two-pane shell: sidebar with list management and per-list progress counts, content pane with nested tasks, progress ring, export and import icons.

Ornith went for a two-pane app shell, a sidebar for lists plus a content pane. The screenshots show per-list progress counts, a circular progress ring, export and import icons in the header, nested tasks with expand arrows, and an "Add item" affordance at every level. It is the only one of the two with a real app-shell layout, and it looks actually closest to what I would expect from a modern to do app.

BigBang app, light mode, as built. Single centered column, emoji list headers with item counts, inline Add input, per-task add and delete buttons.
BigBang app, light mode, as built. Single centered column, emoji list headers with item counts, inline Add input, per-task add and delete buttons.

BigBang went for a single centered column. Its screenshots show emoji-prefixed lists with item counts, a progress percentage, an inline "Add a task..." input with a blue Add button, and per-task buttons for adding subtasks and deleting. It has features Ornith skipped: per-list color and icon choices, a sample-data loader, and a README. One small grammar slip is visible in the UI: the list footer says "1 items". However the UI is ‘like a student’s project from 2010’.

Ornith dark mode. Gradient backdrop so the frosted glass has something to blur, with reduced-transparency fallbacks.
Ornith dark mode. Gradient backdrop so the frosted glass has something to blur, with reduced-transparency fallbacks.

Dark mode is also shows the design gaps. Ornith's dark mode is real frosted glass, and the reason it works is a code detail: it paints a pastel gradient backdrop specifically so the blur has something to blur, and it hand-encoded an iOS spring curve for motion, with an SVG checkmark that draws itself when you tick a task. It also has reduced-motion and reduced-transparency fallbacks.

BigBang dark mode. Clean single accent color, consistent and readable, with per-task progress percentages.
BigBang dark mode. Clean single accent color, consistent and readable, with per-task progress percentages.

BigBang's dark mode is clean and consistent, one blue accent on near-black, no complaints about readability. But its glassmorphism does not actually work, and I can point at why in its CSS:

backdrop-filter

applied over a flat background, so there is nothing to blur, and a shadow variable referenced in a rule but never defined anywhere. Generic ease curves, checkbox drawn with a border trick. To my eye Ornith's app looks noticeably better, and the reasons are in the CSS files. BigBang's is the more conventional of the two.

Ornith mobile at 375x667. Two-pane shell stacks cleanly and keeps every control reachable.
Ornith mobile at 375x667. Two-pane shell stacks cleanly and keeps every control reachable.

Ornith mobile at 375x667. Two-pane shell stacks cleanly and keeps every control reachable.

BigBang mobile at 375x667. readable, but the paired icon-only buttons sit near minimum touch-target size, poor design.
BigBang mobile at 375x667. readable, but the paired icon-only buttons sit near minimum touch-target size, poor design.

BigBang's mobile view works and is readable, but it is spartan. The paired plus and delete buttons sit close to minimum comfortable touch size, the icon-only controls are unlabeled, and its "2 remaining" footer counts only top-level tasks, not all open items. A small logic quirk, not a bug you would catch from the code alone, which is exactly the kind of thing a real-browser test at deeper coverage would have caught.

The verification split, stated fairly

So to be fair.

BigBang's browser tests are with highlighting, cause I’ve seen 4bit models trying to do and either fail or get stuck in the loop. But they are smoke-level. The mobile test only asserts that the page title is visible at 375px. The dark mode test asserts a CSS class exists on the html element. That is more of a "the app opens" test, not "the app works".

Ornith's jsdom suite is the opposite trade: no real browser, but far deeper coverage. Persistence round-trips, corrupt-data normalization, immutability checks, nesting at any depth, reorder, export and import round-trips, and a full UI flow from list creation through nesting, completing, deleting, switching, and re-persisting. Plus to BigBang for handling the browser but Ornith won verification depth.

On the gaps, because it matters to me. BigBang's package.json runs vitest with no config excluding the Playwright spec file, so the npm run test command fails today, and the model saw that exact failure during the session. Its final summary reported "8 unit tests passing" without flagging the broken command. Ornith's test command is green as delivered, and Ornith, when it could not do something, said so out loud.

Code and features, side by side

Ornith appBigBang app
StackReact 19 + Zustand with persistReact 18 + Context/useReducer, zero runtime deps
Stylinghand-rolled design system, no Tailwind, 565 lines of CSSTailwind 3 + PostCSS + 233 lines custom CSS
Total LOC2,093 across 22 files1,554 across 19 files
App code only~1,090~1,099
Item operationsadd, delete, inline rename, reorder, nest, expandadd, delete, nest, expand
Listscreate, rename, delete, switch, export/import JSONcreate, delete, switch, per-list color and icon
Extrasprogress ring, reduced-motion fallbacks, theme gateprogress bar, sample-data loader, README
Tests27 vitest, all green from one command8 unit + 6 real-browser e2e, test command broken as delivered
State todayTypeScript clean, lint clean, build okTypeScript clean, build ok, one broken test command

I specifically listed the row "app code only". Nearly identical. Ornith's 3x token spend did not cause the app to be bigger. It bought twice the test code, a bespoke design system, and a deeper plan. BigBang's motto is leanness but it backfired in the way the app looks.

Both codebases came out tidy, no hacks, defensive handling of saved data in both. Ornith's is the more complete product. BigBang's is the more rushed one.

So who won?

Based on both tests, this and the Voxel one, it’s not a spectacular victory

BigBang’s strong sides are: speed, efficiency, and verification fidelity. 2.2x faster wall clock, a third of the tokens, a quarter of the reasoning, half the peak context, zero tool failures, and the only real-browser verification in the series. If you are paying per task or watching the clock, it is not close.

Ornith wins overall: deliverable completeness, solid code, and process. More features, a deeper plan, 27 green tests behind one working command, zero residual defects, zero lint findings, and it asked before assuming. Its one weakness, the 193K context growth, is a tuning knob, not a character flaw.

The thing I find most interesting is that under an identical harness, the two models produced two opposite personalities. Ornith thinks ~5x more per call and it shows in plan quality and self-review. BigBang decides the simple route, plus more like run it and see. One asked me three questions and caught a bug by reading. The other asked nothing and caught its bugs by breaking things. Both strategies finished the job but these are just different jobs.

If I could combine them, I would want BigBang's instincts inside Ornith's review loop. Move fast, install the missing tool, and then re-read your work the way Ornith does.

Next

I won’t test BigBang any more for now as I have other models to cross check - KAT Coder and Tiel Coder.

I will keep the weights for now and decide if I want to go back to this model and do some other things.

I enlarged Ornith’s thinking budget already as it was hitting the cap. I had this limitation as I noticed other models of that side to use thinking purely for prose and restating 50 times what they want to do.

There is a tradeoff, because reasoning tokens get re-sent in history, so deeper thinking could accelerate exactly the context-growth slowdown that already cost Ornith the tail of this run.

I will check if I can make thinking not getting leaked to the main chat, that is also on my to do list.

Plus the next round of testing most likely will be also around harnesses. I ditched Zcode, have an article on that comparing my previous Zcode runs and Grok Build ones.

But the natural choice for such models, is a harness with the smallest amount of bloat in the system prompt and everything around - Pi Agent.

Hope You enjoyed this one, I try to improve with time, so If you expect some other details, want some other style of the article or I should cut some part of the content to make it shorter, let me know. 🫡😎

Originally published on X · Sep 9, 2026
Next article · 03Don't let the voxel image fool you. Ornith 1.5B 35B vs Bigbang-v1 voxel challenge