The Pelican Benchmark: 2 Years of AI Evolution

In October 2024, Simon Willison gave sixteen LLMs one prompt: "Generate an SVG of a pelican riding a bicycle."

He has run this same prompt on every new model for two years. He now has 103 drawings from 66 different models.

We collected these images into one place: pelicanzoo.ai.

Here is what this experiment teaches us about AI:

  • SVG is code, not pixels. LLMs cannot see, but they can write code. They must build a picture using only coordinates and radii.
  • The task is hard. Bicycles are difficult for humans to draw from memory. Pelicans do not have the right shape to ride them.
  • Early models failed in funny ways. Some drew triangles and circles. Others drew shapes that looked like sinks or tractors.
  • Capability tracking changed. In the first year, better pelicans meant better coding skills. Now, that link is gone.
  • Small models can win. A small model running on a laptop once drew a better pelican than a massive flagship model.

Why keep running a "stupid" test?

It serves two practical purposes:

  • It proves a model works. Posting a pelican shows you actually ran the model instead of just reading the news.
  • It reveals hidden data. You can see how many tokens a model uses and how much it costs to generate a single simple image.

The pelican test is no longer a way to rank the best AI. It is a ritual. It shows you the personality of a model. You might see one model add a little hat to a bird, while another refuses to draw a bike at all because birds cannot ride them.

One model even used 13,000 tokens and 25 cents just to draw one bird.

The drawings show the history of AI through code.

Visit pelicanzoo.ai to see the collection.

Source: https://dev.to/_94be737e156beb4d74df2/two-years-of-pelicans-on-bicycles-103-svgs-66-models-and-what-the-guy-who-invented-the-benchmark-4c7e