
The setup
prism-ml/Ternary-Bonsai-2-27B-mlx-2bit is a 27B model quantized to roughly two bits per weight, about 9 GB on disk. It ran on an M2 Max 64 GB, served by MLX Core (mlx-serve), driven by Grok Build as the agent harness. The task was a two‑turn prompt : plan in plan mode, then build an Apple‑styled todo app with lists, items, nested subitems and local saving.
Round one, thinking off: eight sessions, no app
Eight sessions over about 14.5 hours never produced a compiling app. The final state had 52 TypeScript errors.
The dominant failure was copy‑echo collapse. Grok's search_replace tool asks the model to emit the exact existing text and then the modified version. At two bits, Bonsai kept emitting both strings identically, so the tool rejected the edit and the model tried again. Sixty‑four identical edits were rejected across the round; the worst case was a 122‑minute oscillation on one file, adding and removing the same two lines 42 times while the real compile error sat twelve lines below. The same pattern hit web search in a later attempt: the identical query sixteen times, the same CSS file fetched five times, because that tool never rejects a repeat.
Five of the eight sessions died from my own server configuration, which is worth spelling out (skill issue strikes again):
- The prefix cache 'ate' the context window. mlx-serve keeps already‑processed prompt prefixes in memory so repeated context isn't re‑read every turn. I reserved 16 GB for it. The server then sizes the context window from what memory is left, so it came out at about 47k tokens instead of 131k. Once the conversation grew past 47k, which happens within minutes in an agentic session, every request was refused with HTTP 400 "prompt too long".
- Agent and server disagreed about context size. Grok's config said 155k; the server had 47k. Grok kept sending prompts it thought were fine.
- KV‑cache quantization backfired. Quantizing the key/value cache to 8 bits cut decode from ~29 to 10.8 tokens per second, and the 'dequantization' spikes on every prefill chunk got the Metal command buffer killed by macOS.
- A 155k dense context didn't fit. An 85k‑token prompt needs 10+ GB of working memory on top of the weights; the server refused four such requests.
- A stale process ran the old config. Grok reads its config at startup. My first "thinking on" attempt ran on a process started before the config change, so every request went out with thinking off.
One thing I'd highlight: about 89% of server time was prompt processing on 60–120k contexts, not generation. Decode was steady at ~29 tokens per second. On this hardware the expense of an agentic session is re‑reading the conversation.
What I've changed
The decisive change was turning thinking on: the Grok profile now sends reasoning_effort: "high", which mlx-serve translates to thinking=true. I intended a 1024‑token reasoning cap, but it never applied: MLX Core's budget setting only affects its own chat window, and the app‑managed server that served the run ran unlimited.
The server was pinned to what the machine sustains: 131,072 context, dense KV, 5 GB memory plus 10 GB disk prefix cache, 4096‑token prefill chunks, 1800 s timeout. Grok's auto‑compaction was set to trigger at 75% (~98k tokens) so it fires below the server limit. Result: zero server errors across 170 requests, 96.4% cache hit rate.
Round two, thinking on: an app in an hour, then four hours of nothing
One session, same prompt. About 20 minutes in plan mode, roughly 25 minutes writing the app, 10 more writing tests: 24 files, 1,813 lines, React 19, TypeScript, Vite, Tailwind 4 wired through the official plugin, a reducer‑and‑context store, a persistence layer, Vitest.
No loops. Two identical search_replace attempts, each rejected once and never retried, against 64 the round before. Zero web searches. The transcript holds 167 coherent reasoning blocks, about 185k characters. Final state: tsc clean, vite build clean, 41 of 41 tests passing, zero lint errors.
Then the last project‑file edit landed at 61 minutes, and the remaining 4.3 hours, 84% of the session, went into trying to drive headless Chrome over raw DevTools Protocol websockets to get a screenshot, until I cancelled it. It left two orphaned servers behind.
And it shipped a crash. ThemeProvider was written but never wrapped around App, so the first useTheme() call threw and the page rendered white. All 41 tests passed because none of them mounted App. The video below is the app Bonsai produced with this one issue fixed in order to have it properly built to showcase the final product - no other bugs were fixed here.
Why it fought the browser: mostly the harness, partly the model
This is the finding I got wrong at first. Grok Build's own system prompt contains a <browser_verification> block: "you MUST verify your work in the browser before finishing, whenever browser tools are available."
Neither session had a browser tool. I checked the tool definitions: 27 tools, none for browsing or screenshots, no browser MCP connected. So the harness told both models they must use a browser it never gave them. As for a comparison with a different, moe model - Ornith, which ran the same task in the past, it's reasoning shows it inventorying its tools, concluding "no browser automation available", quoting the escape clause and writing a mount test. Bonsai, two hours into the fight, wrote in its own reasoning: "Option 3: accept that the app works and document the verification method. The user's rule is clear: I MUST verify in the browser." It listed the correct exit and rejected it, because at two bits it had lost the second half of the rule.
Two more twists. Bonsai's homemade CDP rig actually worked well enough to screenshot, and every screenshot was a white page with an empty #root, which was the real ThemeProvider crash. It spent hours blaming headless Chrome for a bug its own verification had found. And the reason it was so sure the code was fine was a compaction summary stating "home‑page screenshot confirms Apple‑like UI renders", a claim with no supporting evidence in the pre‑compaction transcript. Lossy compaction handed the model a false premise, and it reasoned from it for the rest of the session.
So the tail is model weakness (instruction comprehension, sunk‑cost persistence) multiplied by harness weakness (an unsatisfiable mandate, no off‑ramp, a compaction that invented a fact). The fix for future runs is either a browser MCP so the MUST is satisfiable, or a short rule placed first: "No browser tools here. Verify UI with a vitest mount test and stop, but most likely
I will just use a different harness.
Run metrics, successful session
5 h 20 m wall clock, 266 min inside model requests. 167 model calls, 171 tool calls, 4 compactions (one hit Grok's 300 s wall‑clock limit and retried, taking 456 s to compress 98.7k tokens to 13.6k). 9.3 M input tokens, 93.7% served from cache. 106.6k output tokens, about half reasoning. Median decode 26.7 tokens per second, prompts from 19k to 99k tokens, zero server errors. Productive core: 40–55 minutes.
As an example, Ornith did the same task in about 105 minutes of API time with 166 tool calls, no compactions and a real 1024‑token reasoning cap. Tool counts nearly identical; the time gap is decode speed and the verification tail.
The apps: Ornith shipped the better one
I verified Ornith's app before writing this: it mounts, typechecks clean, and its 27 tests pass, including four full UI flows. Side by side, the result is not close.
As a comparison, here is the Ornith 1.5 35B 4bit MLX built app, same prompt, same harness.
- Ornith's app worked on first load. Bonsai's crashed to a white page. For a user that alone decides it.
- Ornith got subitems right. Bonsai didn't. The prompt was built around nested subitems. Ornith's countItems walks every level, so progress, completion and sidebar counts include subtasks, with tests to prove it. Bonsai's selectors count only the top level, so a list reads "2 of 2 done" with open subtasks inside it; its subitem badge is shown only when expanded, the reverse of useful; every row carries a permanent "add sub‑item" control that clutters deep trees; and its expand animation is capped at a fixed 2000 px, so large subtrees clip.
- Ornith did the boring parts. plan.md in the project root, as asked. Bonsai's plan never landed there.
- Ornith shipped more real function. JSON export and import, item reordering at any depth, delete animation, collapse state held in one place. Bonsai's extras, toasts, a follow‑system theme mode and a progress pill, are cosmetic, and the theme was the thing that crashed it.
- Ornith's tests test the product. Bonsai's test the plumbing. 27 versus 41 flatters Bonsai. Ornith's App.test.tsx creates a list, nests subitems, completes them, checks localStorage and deletes. Bonsai's 41 are all store, selector and storage units, and they coexisted with an app that could not render.
- Both apps share the same solid foundations: schema‑versioned localStorage with defensive normalisation of corrupt data, light and dark themes, inline editing, unlimited nesting in the data model, clean typing.
What Bonsai earns credit for, stated precisely: its code is well organised (clean reducer, immutable recursive updates, careful storage validation), readable, and it compiles. For a 2‑bit model that could not compile anything a week earlier, that is a real step. But well‑structured code and a working app are different deliverables, and only Ornith delivered the second.
Takeaway
Bonsai being 2bit and roughly 9GB size did well for it's size but not enough to beat a 4 bit moe model. I wouldn't expect it to beat it but I chose this comparison for a reason - if you have enough memory to support a 4 bit 35B moe - this will be a better choice over a 2 bit model like Bonsai.
I think there might be a use case for Bonsai, under a different harness (Pi or OMP maybe) and probably some further settings finetune plus agents.md specifying the testing and verification protocol, ensuring the bugs get caught.
