Milkbench
All GPT-5.6 models and reasoning levels, tested for taste.
I created a prompt to one-shot a website for milk as if it was some sort of new revolutionary tech product. But it’s just milk.
We're tasting the milk.
Then I took that one prompt and, from a clean session on a clean slate, gave it to Codex across all different models and reasoning levels:
Luna: Light, Medium, High, Extra High, Max.
Terra: Light, Medium, High, Extra High, Max, Ultra.
Sol: Light, Medium, High, Extra High, Max, Ultra.
17 sites. Same prompt. Same product.
The results are pretty expected, but there were some interesting findings.
Sol Ultra is 32% of the total cost. One out of 17 sites was a third of the cost. Sol as a family is 73% of the total API cost, even though it only has 6 of the 17 sites.
Meanwhile, all five Luna runs together cost $3.97. Luna High and Extra High cost less than a dollar and take about 15 minutes. The value there is pretty outstanding.
Some things didn’t correlate at all. More thinking doesn’t mean a better site, turning up the effort doesn’t always mean it creates more work, and the size of the sites really didn’t have any relevance at all.
The heaviest site is 45x larger than the lightest: Terra Ultra is 10.9 MB, while Luna Max is only 0.24 MB. Same prompt, same product, but widely different shipping weight. But Sol consistently shipped lightweight sites with a premium feel, and Terra always shipped heavy sites.
Sol Ultra also made up about 24% of all token spend, using 1.33 million tokens and taking 70 minutes. Across all the runs, the lowest was Terra Medium at 59k tokens, while the highest was Sol Ultra at 1.33 million.
Terra Medium is a real outlier. It’s a higher effort than Terra Low, but it took less time, used fewer tokens, and was significantly cheaper. It also made a pretty terrible site.
Sol’s effort ladder isn’t a clean scale either. Medium is slower and has more token spend than High, while High is slower than Extra High.
The carton is the real skill check, and pretty much all the models fail. Sol Ultra does the best here, but it also spent 22 times more tokens than Terra Medium.
The hero section is another regular problem. The carton on top of the headline shows up across multiple models and effort levels. Sol Extra High is the only one that actually put the carton behind the headline, but then it put white text on a white carton.
Luna Light really outperformed its price tag, while Sol Light already feels very premium at a fraction of the cost. Sol Ultra is probably the best one-shot, but it’s also the worst value.
The family breakdown is pretty interesting.
Luna has the best effort-to-taste until you hit Max. Terra is the chaotic middle child. Sol is the premium one, but it comes with 73% of the total API cost.
Turning the reasoning up is not a quality slider. There’s a lot of variability.
The higher efforts fairly consistently reduce catastrophic failures and unlock animations and interactivity, but the peak taste is often around the mid-high levels, not at the absolute top of the dial.
Overall, Luna High and Extra High are the best value. Sol Ultra is the best website overall. And the best premium taste is Sol, especially once you break into Sol High.
The Milkbench overview has all the benchmark details and sites available to browse.