Show HN: Pelican-bicycle alternatives

gally.net

129 points by tkgally 2 days ago


In November and December 2025, inspired by Simon Willison’s pelican-riding-a-bicycle benchmark, I had some then-current LLMs create SVGs from thirty similar prompts, such as “Generate an SVG of an octopus operating a pipe organ.” Simon mentioned that experiment on his blog [1].

Nine months have passed and much stronger models have been released, so I tried the experiment again today. The linked site shows the results.

Running ten of the prompts through six models at OpenRouter cost about twenty dollars, so I stopped there for now.

[1] https://simonwillison.net/2025/Nov/25/

svcrunch - 2 days ago

I'd like to mention the Little Dorrit Benchmark [1] which I have been running for a couple of years now. It has a few nice features:

1. It tests visual reasoning and structured output in a single task.

2. It seems to sort correctly on advancing general intelligence. As a counterexample, if I'm not misremembering, artificialanalysis.ai made some changes to their benchmark recently after Astra ranked below several older models.

3. While models have gotten significantly better in the past 2 years, the top model is still at 0.78 F1, so the test is not yet saturated. As a reference point, when I started, the top models were in the [0.1, 0.2] range.

[1] https://dorrit.pairsys.ai/

vova_hn2 - 2 days ago

Website looks very cool, Fable's octopus-organist looks very cute, but I feel like this benchmark (generate an SVG by a short and slightly ridiculous description) in general has been completely Goodharted [0].

I think they all just added a bunch of similar tasks to their training sets, so we cannot judge true emergent capabilities of the models anymore.

[0] https://en.wikipedia.org/wiki/Goodhart%27s_law

samayashar - 2 days ago

All models are pretty good now at generating these images. Back in the day, I remember experimenting with the pelican images and most of the models couldn't align the legs with the wheels. Right now as well, GPT messed up an octopus leg by originating it through the instrument rather than the octopus itself.

I think that intertwining two entities (living/non-living) is still challenging but overall they're pretty sound.

outlore - 2 days ago

Anyone else surprised the generations look so remarkably similar? All of these models have “independently” generalized that the moose should roughly be standing at the same position (left) or that the giraffe should have a certain color palette.

meta-level - 21 hours ago

Qwen 3.8 Flash-Next is really missing here (because it's small enough to run it locally on a (formerly) affordable machine). This is what Qwen3.8-Flash-Next-UD-Q4_K_XL gives me for "an octopus operating a pipe organ": https://imgur.com/a/zHyHIqI Seems very similar to the output of Qwen 3.8 Max to me

kennywinker - 2 days ago

Would love to see Qwen3.8-27b here, since that is the model most people are running locally.

ianberdin - 2 days ago

https://playcode.io/blog/macbook-svg-benchmark

MacBook Pro 3D in SVG for me the most helpful one.

theshrike79 - a day ago

What's interesting to me is that the SVG versions don't have the "AI Image generation hates negative space" issue as badly as generated images do. They kinda stay on point and don't fill every single empty bit with some pattern.

steinvakt2 - 2 days ago

Feels like google has a different training set than the others?

sajithdilshan - 2 days ago

Interesting, out of all examples Gemini 3.8 is the best for me. Also the image style is different and more vibrant than others

- 2 days ago
[deleted]
zaphar - 2 days ago

I notice none of the octopi seem to be actually facing the organ.

andy_ppp - 2 days ago

Love the output from Qwen 3.8 it seems very impressive for the cost! Why does Gemini 3.8 flash blur everything? What are Google playing at!

miohtama - 2 days ago

I hope there would be parameter iterations allowing especially cheaper models on inspect and fix their output via rendering

sceptic123 - 2 days ago

> An elephant typing on a typewriter

A monkey, surely?

pohl - 2 days ago

Gemini 2.5 Pro is the only model with a sense of where a ferris wheel operator would be.

eddytrex_ - 2 days ago

Does a test of instructions how to fold origami figures in a SVG/jpeg exist? Or could be useful?

BrokenCogs - 2 days ago

Gemini 3.8 flash seems to (subjectively) be the outlier in terms of performance to cost ratio?

neilellis - 2 days ago

Well that benchmark is now saturated, what next. How fast you can hack the pentagon?

dustfinger - 2 days ago

It is interesting how similar the designs are across the models.

qiine - 2 days ago

Asking to animate it add an interesting layer of difficulty

water-drummer - 2 days ago

Ok the Grok ones are cute

mock-possum - 2 days ago

Try asking an LLM to draw you the cool S.

dcreater - 2 days ago

Why is this a good test?

- 2 days ago
[deleted]
simonw - 2 days ago

I love these.

bicepjai - 2 days ago

[dead]

aaron695 - a day ago

[dead]

villish - 2 days ago

The 3 US models have their own style.

Qwen3.8 is very clearly distilled from Claude models.