Twelve models participated, including GPT-5.6's Sol, Terra, and Luna, as well as Meta's Muse Spark 1.1. The experiment compared operational verification and costs by having each model generate four different apps five times each.

In the first-person maze exploration task, GPT-5.6 received high marks across all its tiers. Claude unexpectedly struggled, while Grok 4.5 was considered practical relative to its price point. Muse Spark also produced surprising results in cases where it functioned.

For the 3D Rubik's Cube, GPT-5.6 received unexpectedly low evaluations. In contrast, Claude Fable 5 succeeded in flawless calculations 5 out of 5 times, earning high praise. It was reported that Opus failed to achieve a complete solution even once.

In the calculator app task, Claude delivered the best results. Both Opus and Fable functioned correctly 5 out of 5 times, with Fable leaving a positive impression regarding design. GPT-5.6 Sol attempted 3D rendering, but it was reported that the experience was poor due to inconsistent styling.

For Conway's Game of Life, open-weight models achieved high-quality results at a low cost. Analysis suggests that Qwen 3.7 Plus and GLM-5.2 benefited from the abundance of existing code.

In SVG generation, Claude Fable produced high-quality and humorous results. It was pointed out that the GPT-5.6 series lacked clean rendering and showed little room for growth.


Source: GPT-5.6, Grok 4.5, Claude, and Muse Spark build the same 4 apps(HN 159pt・89コメント) (HN Search (backfill), 2026-07-11)