"Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok

tryai.dev

226 points by hershyb_ 16 hours ago


NichoPaolucci - 14 hours ago

As I looked through the images I was unimpressed entirely, at first. But, then I started thinking, these look a little... "childish" to me.

Childish as in... A newish artist who is drawing a concept rather than light / forms (Which is something artists typically do as they understand drawing more and more).

The rose in the vase specifically - some models understood that there was supposed to be shading, reflections, the concept of refraction - others just drew "blue = glass" and "green = stem" and "red = rose".

Really odd to look at, considering if I saw any of these drawings from a human kid, I would say "good job buddy" and put it on the fridge. I'm expecting these to get better as models improve, and perhaps the artistic progression will be there along with it...

jnathsf - 12 hours ago

GPT 5.6 Sol had the best two drawings (rose and starry nights) but even more impressive was how efficient it was RE cost/time/tokens vs Fable (3.4M vs 14.6M / $7.74 vs $161!). OpenAI has quietly innovated around inference - this is will be a growing differentiator even against open models.

bdcravens - 11 hours ago

The Grok ones are amusing, almost comically bad. However, whenever I've tried to pass an image creation request to any of the Opus models, it's been far worse, like first week of using Microsoft Paint bad (while ChatGPT would create social media quality images using the same prompts)

ksd482 - 16 hours ago

Grok! LOL!

Seriously, what's going on there ? Why is it so different from others? Is it just behind technologically/training wise or it's using something fundamentally different?

traes - 5 hours ago

Perhaps there was something in the prompt or settings preventing this, but I'm surprised (and slightly disappointed) that none of the models approached this the way I would: download an image of the Mona Lisa, run all of the drawing functions many times to construct a forward model of the drawing implements, and attempt to solve some kind of explicit inverse problem through either ML or a classical algorithm to minimize some difference metric. Were they just restricted from running code or are the models unindustrious without a very specific prompt?

fastball - 7 hours ago

The most interesting result for me is that they apparently prompted the models to optimize for SSIM, but many of the models trend worse over time. I suppose because viewing the canvas always comes after drawing, and they didn't give "revert to previous" capability as part of the toolkit.

Which in turn kinda jives with my experience of using these models for code: to some extent they only seem to have a concept of "forward", which invariably leads to "write more code to fix previous problems created", rather than taking a step back and removing broken things entirely.

tylerrobinson - 14 hours ago

The Grok ones are so weird that they cross into uncanny and surreal. Truly bizarre and capable of eliciting feelings from me, if only bad feelings…

hombre_fatal - 12 hours ago

GPT-5.6 Sol is the knock out here. Some of those results are really human/charming.

The rectangular smudge tool is a weird tool in the first place, but it's cute to see the models try to use it.

tulio_ribeiro - 6 hours ago

Weird choice of SSIM/RMSE. By feeding these back into the model, the agent is actively degrading its artistic output. This is only valid for the target reference I think?

Alternatively, I think much better results could be had by computing the cosine similarly using DINOv2 ONNX (@xenova/transformers).

My guess is that the results are greatly limited by relying on the current metrics.

Even better, attach toModelOutput directly to drawTool so it returns the rendered canvas image immediately upon drawing.

ms7892 - 16 hours ago

Fix this signup error please

https://www.tryai.dev/?error=server_error&error_code=unexpec...

NiloCK - 7 hours ago

For capabilities reference:

I made a lower effort but similar scaffold for LLMs to do iterative drawing in Nov 2024, with Sonnet 3.5 as the artist: https://paritybits.me/llm-drawing-with-eyes-open/

Quite a difference.

sashank_1509 - 7 hours ago

Where do these models even have this data to learn from. There must be massive computer use datasets? I have a hard time believing it’s emergent if it’s able to do something this good.

LastTrain - 13 hours ago

Why include Grok when it is clearly not even in the same category as the other three?

taf2 - 14 hours ago

Would love to see the results if they had used /goal or similar

MitziMoto - 10 hours ago

Grok's look like a truly disturbed child. Like the drawings from that kid in "The Ring".

dizzard - 13 hours ago

I wonder how much better a harness could get for drawing

pavel_lishin - 16 hours ago

I think this is just an ad.

eth0up - 15 hours ago

Claude was clearly 'pushing back' on the coziness of the cabin. But I think it did best with the cat. Grok, I fear, is making a case for euthanasia. It's suffering and I think it would be cruel to let it continue. Someone pull the plug. .

mbmbn - 6 hours ago

Shows about the same level of technical mastery that present days painters.

AlienRobot - 15 hours ago

The difference in cost is pretty incredible.

wolttam - 9 hours ago

Would be great to see Kimi here.

roguedemon - 8 hours ago

That was a good testcase. Pretty cool.

reactordev - 12 hours ago

Should have compared against FLUX Klein and Z-Image Turbo...

fwip - 13 hours ago

Article is AI-written.

maxall4 - 11 hours ago

I very much enjoyed Grok 4.5's rendition of the Mona Lisa as Elon Musk with tentacles.

SXX - 14 hours ago

Now add Deepseek, GLM and Kimi :-)

Razengan - 14 hours ago

Now this is a cool test. Sol's the best in each one, while Grok looks like the work of that person who "fixed" that Jesus painting..

flintapi - 6 hours ago

[flagged]

- 13 hours ago
[deleted]
geroge_kyaw - 16 hours ago

Useless They are not image generation models.

KolinFirz - 16 hours ago

Tell them to draw LeBron James or Heisenberg. Everyone will refuse.

leobg - 6 hours ago

I'm missing some context here.

First of all, this is a marketing piece, made for engagement.

Second, in production, you would adjust your approach to the model's strengths and weaknesses. For example, if you found that a model does not respond to the tools you are offering, then you would find a way to make it more responsive. Also, we know that different models handle prompts and system prompts differently. So comparing several models from different providers against exactly the same prompt is quite a naive way of exploring the capabilities of those models or comparing them to each other. You may have just happened to “speak the language“ of one model better than that of another. Why would that be the model’s fault and not yours?

Third, none of these results mean anything if you do not run each process on each model several times to check for consistency. Your model might just have gotten lucky. This point is even more true when you're talking about a progressive task where one step builds on the previous one. The model might just have painted itself into a corner.