Show HN: Pelican-bicycle alternatives (updated for 2026)

(gally.net)

42 points | by tkgally 3 hours ago ago

13 comments

  • svcrunch 34 minutes ago

    I'd like to mention the Little Dorrit Benchmark [1] which I have been running for a couple of years now. It has a few nice features:

    1. It tests visual reasoning and structured output in a single task.

    2. It seems to sort correctly on advancing general intelligence. As a counterexample, if I'm not misremembering, artificialanalysis.ai made some changes to their benchmark recently after Astra ranked below several older models.

    3. While models have gotten significantly better in the past 2 years, the top model is still at 0.78 F1, so the test is not yet saturated. As a reference point, when I started, the top models were in the [0.1, 0.2] range.

    [1] https://dorrit.pairsys.ai/

  • vova_hn2 an hour ago

    Website looks very cool, Fable's octopus-organist looks very cute, but I feel like this benchmark (generate an SVG by a short and slightly ridiculous description) in general has been completely Goodharted [0].

    I think they all just added a bunch of similar tasks to their training sets, so we cannot judge true emergent capabilities of the models anymore.

    [0] https://en.wikipedia.org/wiki/Goodhart%27s_law

    • GaggiX 5 minutes ago

      I don't think a bunch of similar tasks can really saturate the "create a SVG of X", because the model should have a quite good spatial understanding of the world and how everything interacts.

      For example Gemini 3.8 Flash seems very impressive at first glance but the results are not actually very coherent, this shows that its "world model" is not particularly great (compare to SOTA models).

  • samayashar an hour ago

    All models are pretty good now at generating these images. Back in the day, I remember experimenting with the pelican images and most of the models couldn't align the legs with the wheels. Right now as well, GPT messed up an octopus leg by originating it through the instrument rather than the octopus itself.

    I think that intertwining two entities (living/non-living) is still challenging but overall they're pretty sound.

    • CamperBob2 25 minutes ago

      All models are pretty good now at generating these images.

      Not zebras. If you want to see how bad SVG output still is, ask for a zebra riding a scooter.

  • BrokenCogs 15 minutes ago

    Gemini 3.8 flash seems to (subjectively) be the outlier in terms of performance to cost ratio?

    • GaggiX 11 minutes ago

      Gemini 3.8 Flash results are often not very coherent but it does put a lot of shading and details to hide the fact.

  • neilellis 15 minutes ago

    Well that benchmark is now saturated, what next. How fast you can hack the pentagon?

  • eddytrex_ 24 minutes ago

    Does a test of instructions how to fold origami figures in a SVG/jpeg exist? Or could be useful?

  • steinvakt2 2 hours ago

    Feels like google has a different training set than the others?

  • sajithdilshan an hour ago

    Interesting, out of all examples Gemini 3.8 is the best for me. Also the image style is different and more vibrant than others

  • sceptic123 an hour ago

    > An elephant typing on a typewriter

    A monkey, surely?

  • qiine 41 minutes ago

    Asking to animate it add an interesting layer of difficulty