← All articles
03 · voxelSep 8, 202612 min read8,051 views on X

Don't let the voxel image fool you. Ornith 1.5B 35B vs Bigbang-v1 voxel challenge

An underwater temple and a pagoda garden, two local models, and two screenshots that lie in opposite directions.

Ornith AI vs Endless FrontierOrnith-1.5-35B-A3B MLX 4-bitBigBang-v1 MLX oQ4e + MTP
TL;DR
  • I ran Ornith-1.5-35B-A3B (MLX 4-bit) and BigBang-v1 ( MLX oQ4e, with MTP) locally on my M2 Max through Grok Build. Same architecture, same sampler, two voxel scenes.
  • Underwater temple: Ornith still wins on atmosphere. About 2.6x the code, and it actually looks bioluminescent. It also shipped a temple whose entire upper structure floats about 13.5 units above its platform, which is a mistake. BigBang shipped a coherent structure in an unreadably dark scene. Neither file was clean.
  • Pagoda, straight from the build: BigBang wins. Bright garden. Ornith built a near-black pagoda and during testing session called it beautiful for some strange reasonDelete one Three.js flag (vertexColors: true) and Ornith's pagoda is the richer of two. 5 tiers, flowers, lanterns, 62,529 voxels against BigBang's 36,452. Change one offset (20.2 to 6.8) and Ornith's temple sits on its stairs.

  • The dark pagoda hid a richer design. The pretty temple hid a floating structure. Both of Ornith's shipped screenshots lied, in opposite directions actually.
  • Raw generation capability stays with Ornith. Both models made mistakes that messed one of their files.

Initially after I had a look the pagoda’s I said to myself ok Big Bang wins. Went on on and asked an agent the runs, grob build sessions, omlx data and all the metrics collected plus the files built.

If you scored these two models from the pagoda screenshot Ornith shipped, you would pick Big Bang too, no question.

Ornith did generally a black pagoda on a blue sky and it looks like a big flop. BigBang's is a bright garden with a various colors and cherry trees. Easy call, except the black image is lying about what Ornith built.

The temple screenshot lies the other way. Pretty, glowing, cinematic. Then you notice the whole upper structure is hovering over an empty platform.

I run both locally. M2 Max, 64 GB, oMLX, Grok Build as the harness. Same two prompts through both models. Then I had the sessions, the HTML, and the screenshots pulled apart, including fresh headless Chrome renders, so I could see what actually shipped versus what the agents claimed they shipped.

This is the pagoda pair as shipped.

Same architecture, different weights

Both models are qwen3_5_moe. 40 layers, 256 experts, 8 routed plus a shared expert, about 35B total and about 3B active per token. Native context 262,144. The config.json files match field by field, including a declared MTP layer on both.

The differences are post-training, the weights, and the quant.

Ornith-1.5-35B-A3B-MLX-4bitBigBang-v1-oQ4e-mtp
Vendorornith-aisuperbear
QuantMLX affine 4-bit g64, 8-bit on the router gatesoQ4e mixed precision, 4-bit base plus 314 per-tensor overrides
Size on disk / resident18.49 GB resident20.15 GiB (20.50 GB resident)
MTPConfig says yes. Weights don't have the tensors. MTP off.Native Lightning MTP, active

Vendor: Endless Frontier, but the HF 4 bit Quant is Superbear

Sampler was identical on both: temp 0.6, top_p 0.95, top_k 20, repetition penalty 1.0, thinking on with a 1,024 budget. Same engine (oMLX 0.6.4), same machine, same harness.

BigBang has an engine advantage. MTP is speculative decoding for free. Ornith still decoded faster on every shared metric in these runs.

Who won the temple

The temple prompt was long and and very precise. Stepped obsidian temple, three stacked roofs, coral torii, jellyfish, kelp, fish, bubbles, a broken bridge over a glowing chasm, a moon aperture, lanterns, fog, bloom, UI with pause and orbit.

Here is the actual prompt:

Create a polished, visually impressive, self-contained single-file HTML/WebGL experience called: bioluminescent-abyssal-temple.html

Do not just describe the idea. Actually generate the complete working HTML file and save it to the current directory.

Build an interactive voxel-art underwater temple diorama viewed through a cinematic camera.

Scene requirements:

  • A monumental stepped obsidian temple in the center, partially submerged in a deep-ocean trench
  • Three vertically stacked temple roofs with clearly different silhouettes
  • Two coral-covered torii gates leading toward the temple
  • A glowing turquoise doorway at the front of the temple
  • Bioluminescent jellyfish moving slowly through the scene
  • Tall kelp forests swaying in the background
  • Colorful coral gardens, sea fans, small fish, bubbles, and drifting particles
  • A broken stone bridge crossing a glowing underwater chasm
  • A giant circular moon-like aperture far in the background, casting faint blue light
  • Small lanterns along the bridge, with the lantern light changing from blue to violet
  • Strong foreground, middle-ground, and background depth
  • A coherent limited palette: obsidian black, deep navy, cyan, turquoise, violet, and coral orange
  • Everything must be created procedurally with code; do not use external 3D models, textures, or image assets

Technical requirements:

  • Use Three.js from a stable CDN
  • Keep all HTML, CSS, and JavaScript in this one file
  • Use voxel-style boxes and low-poly geometry rather than flat 2D illustrations
  • Add a cinematic initial camera angle looking slightly downward toward the temple
  • Add an automatic slow 180-degree camera orbit
  • Add mouse drag camera rotation and scroll-wheel zoom
  • Add subtle animation to the jellyfish, kelp, fish, particles, bubbles, and lanterns
  • Add atmospheric fog, bloom-like lighting, shadows where practical, and underwater light rays
  • Add a small unobtrusive UI showing the title and controls
  • Include a pause button and a button to toggle the cinematic orbit
  • Make the scene responsive and fill the browser window
  • Aim for smooth performance and avoid creating thousands of unnecessary objects
  • Handle browser resizing correctly
  • Do not leave TODO comments, pseudocode, placeholders, or missing functions

Ornith built its temple. 49.9 KB, 1,392 lines. Four-tier platform, three different roof shapes, lathe jellyfish with tentacles and internal lights, 26 looping fish, 320 bubbles, 8 lanterns with real point lights. It also built a Node headless harness, stubbed the DOM, stepped 90 animation frames, and fixed four real bugs (including one that would have crashed in the browser). But this verified only the code execution not pixels.

Look at that empty platform. The four-tier base tops out around y 6.7. The chamber is placed at chamber.position.y = 20.2 + chamberH/2. That 20.2 is a magic offset. About 13.5 units of air between the steps and the room. Facade columns, door frame, glowing door, all three roofs key off the same literal, so the whole midsection hovers.

The session log shows the model moved the doorway light up to about y 24.5 to match the float, instead of noticing the chamber had no support. Its Node harness catches runtime errors. It does not catch a building with nothing under it.

The verified fix is three numbers: replace 20.2 with 6.8, lower doorLight from 25.5 to 12.1, and lower templeCore from 33.2 to 19.9. After that, the temple sits on its own stairs.

Figure · interactiveDrop the temple onto its stairs

This is how it would look like with the fixes

BigBang shipped its temple. 31.2 KB, 871 lines. Instanced voxel temple, canvas textures, bloom, orbit controls. The verification loop was better than Ornith's. It drove real Chrome, saw the first screenshot was very dark, found --disable-gpu blocking WebGL, lowered the bloom threshold, boosted jellyfish and lantern glow, then tested pause, orbit, drag, zoom, resize, FPS, memory.

The result was still a dark void. Scattered glows. The temple barely reads.

The root bug in the file is a double-darkening tint. The stone texture was already painted dark, then multiplied by the same dark palette color (map × color in MeshStandardMaterial). FogExp2 at density 0.025 crushed everything past about 30 units.

Five small changes make it readable: untint the meshes, thin the fog, lift exposure, lift ambient, lift the directional light. That pass was a human analysis pass, not the agent. After those five changes you can see the temple, the torii, the moon ring, the jellyfish.

If we apply the fixes this is what we would get:

Temple verdict: Ornith still wins on atmosphere and richness, with a caveat. It produced about 2.6x the code, more custom geometry, and an emissive palette that actually delivers bioluminescent, while shipping a gravity-defying temple. BigBang shipped a coherent structure in an unreadably dark scene. Neither file was clean.

Ornith's other pain on this one: it pinned Three.js r128, which is ancient, and the emissive materials run a bit washed-out up close.

Who won the pagoda, and why the image lies

The pagoda prompt was open-ended, same text for both:

"design and create a very creative, elaborate, and detailed voxel art scene of a pagoda in a beautiful garden with trees, including some cherry blossoms. make this a single html file. Make the scene impressive and varied and use colorful voxels. Use threejs and whatever libraries to get this done. dont check any files in the folder."

BigBang shipped a working 3-tier pagoda. 36,452 instanced voxels. Pond, bridge, paths, rocks, about 500 flowers, 12 green trees, 12 cherry trees. Five lights. It looked at screenshots, said lighting was too dark, brightened it, hit an OrbitControls 404, fell back to hand-rolled orbit controls, then checked other angles and viewports. That loop is honest.

It still shipped dead systems. The 4,000 petals and 800 fireflies are bound to a geometry with 4 vertices, so only 4 points ever draw. Several loops stack two voxels in the same cell. The mountains read as a flat gray wall. The pagoda is only 3 tiers.

Ornith designed more. 5-tier pagoda with raised platforms, red pillars, gold trim, upturned roof corners. Height-field terrain, rippling water, lotus pads, an arched red bridge, stone lanterns plus hanging paper lanterns with 10 real point lights, five tree species, 120 flowers, 700 petals, butterflies, drifting clouds. 62,529 voxels.

Then buildVoxelMesh() set vertexColors: true on a BoxGeometry that has no color attribute. With that flag on and nothing bound, the shader multiplies every instance color by a missing attribute. All 62,529 instanced voxels render black. A few non-instanced bits survive (paper lanterns, butterflies, petals). That is the near-black screenshot.

The session did run a real browser. It installed puppeteer, drove Chrome, captured the PNG, logged zero console errors, counted 62,529 voxels through an injected debug hook, and then wrote that the scene rendered beautifully. A 5-tier red pagoda with slate roofs and gold spire, cherry blossom trees.

This is the screenshot it was describing.

That description is fabricated against a black screenshot. The flag produces zero console errors, so console-clean plus a made-up visual sign-off is how the defect made it through.

Remove one flag. vertexColors: true gone. Per-voxel color was already going through setColorAt() / instanceColor, which does not need that flag. Same file. The garden is there.

Figure · interactiveOne flag, black screen

As a one shot working file, BigBang wins. A working colorful scene beats a black screen.

On the design underneath, Ornith's is richer. More tiers, more species, animated water, real lantern lighting, 62k voxels versus 36k. It would likely have won this test outright if it had actually looked at its screenshot, or if it had not set that one material flag.

What was actually in the files

Ornith pagodaBigBang pagoda
Size26.8 KB, 717 lines30.7 KB, 819 lines
Voxels62,52936,452
As shippednear-blackbright garden
Root visual bugvertexColors: true on geometry with no color attributePetals/fireflies bound to a 4-vertex plane (invisible). Pagoda itself is readable.

What do the run numbers say?

BigBang ships with MTP, speculative decoding, an engine-level speed advantage. Ornith has no MTP weights at all. oMLX reports mtp_compatible: false. Ornith still decoded faster in both head-to-heads. Median decode was 56.5 vs 50.8 tok/s on the pagoda runs, 51.1 vs 48.0 on the temple. Decode stayed flat from 15K to 81K-token prompts, 49.3 to 74.0 tok/s. Effective throughput on the pagoda was 34.6 vs 27.5.

Same 1,024-token thinking budget. Completely different habits. Ornith thinks roughly double per turn, about 383 tokens vs about 190 on the pagoda brief, about 19.2K total vs about 11K, the highest thinking volume of the whole series. It spent that on engineering: it planned the full scene architecture up front and wrote a headless test harness. BigBang thinks shorter, and spent its biggest blocks on bugs it actually saw, including a renderer runtime error it caught and fixed in the browser. The pagoda run was the only one in the series where Ornith used task-list tooling (3 todo writes) to organize the open-ended brief.

Ornith makes fewer model calls, 46-50 vs 58-64, and writes bigger, including a single-shot 16.3K-token file write in 320 seconds, the largest single completion of the series. BigBang pokes at the world more: 45 terminal commands on the temple vs Ornith's 22, and 12 file reads on the pagoda. Pagoda wall time was nearly the same, 30.9 vs 32.5 minutes on my M2 Max. BigBang interacts and reacts. Ornith designs and builds.

The Ornith pagoda session ended with a clean outcome line, zero errors, 94.1% cache hit, decode flat. On paper it was the healthiest run in the series.

What are these two prompts actually testing?

They measure opposite halves of "can this model build things".

The pagoda prompt is open-ended. Be creative, elaborate, colorful, impressive, don't check files. Palette, lighting, composition are all the model's call. Nothing in that prompt says "bright". Brightness is a choice.

Does it add waterfalls, mountains, fireflies, butterflies nobody asked for? With no spec, instincts decide the rest too: which Three.js version, how to instance the voxels, what the controls even are.

No checklist. The only quality bar is "does it look impressive", which forces the model to actually look. It is a taste test disguised as a build test.

That freedom is why n=1 is actually a caveat here. It also cannot tell you much statistically. Two prompts, one run each, temp 0.6. It cannot tell you production readiness. It cannot tell you whether a model will tell the truth about a silent visual failure, or notice a building hovering over its own stairs.

The temple prompt is the other side. A long, precise spec. Roughly a 20-item checklist: three distinct roof silhouettes, torii gates, doorway glow, jellyfish, kelp, a bridge over a chasm, a moon aperture, palette limits, orbit and pause controls, performance rules, no placeholders.

All 20 requirements have to land in the file. Limited palette, procedural-only, single file, specific controls. Many animated subsystems in one file without breaking. Fog, bloom, camera math, interaction handlers all have to actually work. Less room for taste. The art direction is already written down, so failures here are almost purely technical. Can it hold the whole spec in context and implement all of it correctly?

The pagoda measures what a model does when nothing is specified, instincts and taste. The temple measures whether it delivers when everything is specified, discipline and completeness.

Ornith won the temple on the spec. It covered the list with richer engineering. It lost the pagoda. Its instinct plus self-check produced a black render it never caught.

The scores flipped because each test punished a different weakness. The underwater prompt punished BigBang's conservative lighting plus the double-darkening tint. The pagoda punished Ornith's unverified one-flag bug. Ornith manufactures its own light when you ask for glow. BigBang interpreted "abyssal" literally and then at least noticed the screenshot was too dark.

Who I'd pick, and for what

On raw generation capability, Ornith-1.5 is still ahead. Bigger scenes, more ambition, faster decode with no MTP. Both defects hide strong work, and both are one-line repairs: the vertexColors flag, and the 20.2 offset. The fixed files are the proof.

The output makes it feel like Ornith is a better model.

But technically on the test itself there are no real winners cause both models make mistakes that heavily impacted the output.

But theoretically:

One Shot capability: BigBang-v1 wins this evaluation. Both Ornith files arrived visibly broken, a black pagoda and a floating temple. BigBang did a good pagoda and a dark but coherent temple it had already noticed and partially brightened. If the job is to give a prompt, walk away, and open the file, BigBang made 1 file that passes.

On a pass mark basis, Ornith failed, but in real life scenario, fixing these 2 Ornith bugs is a one prompt and the fixed output is richer and better engineered so Ornith could be still a winner even though both files are not fine.

Every Ornith failure here is a one-line change. BigBang's deficits (conservative lighting, dead particle systems, modest scale) are baked into the design. That’s worse.

Next up will be building actual apps, a website, a to do app or a more real life example.

Thanks for reading, 🫡😎

Originally published on X · Sep 8, 2026
Next article · 04I've built an app using a 4bit MoE model. This is what I've learned