Single-prompt “vibe” benchmarks get mocked, and mostly they earn it. But the Habsburg-frog benchmark is a sharper instrument than it looks: one fixed prompt — generate an SVG of a frog with a Habsburg jaw — run against 14 models, three tries each per month. Forty-two runs, forty-two produced valid SVG. The scoring isn’t “which frog is prettiest”; it’s what that prompt forces a model to do.

This is the pelican-on-a-bicycle lineage, and it works for the same reason. Saturated benchmarks stop discriminating, so people reach for weird, underspecified prompts where instruction-following and spatial composition still separate the field. It’s no substitute for a real eval harness — no rubric, tiny n, no statistical power. But as a cheap, reproducible probe you can rerun monthly, it beats chasing another MMLU point. The HN discussion argues over contamination and taste. My question: how long until a model trains specifically on frog-with-a-Habsburg-jaw and the probe goes dark?

tags: [ llm-ops ] [ research ]