Simon Willison on X: "Playing with a new LLM benchmark: how good are they at drawing an SVG of a pelican on a bicycle?
Here's the difference between Claude 3.5 Sonnet new and Claude 3.5 Sonnet previous"
Playing with a new LLM benchmark: how good are they at drawing an SVG of a pelican on a bicycle?
Here's the difference between Claude 3.5 Sonnet new and Claude 3.5 Sonnet previous
Playing with a new LLM benchmark: how good are they at drawing an SVG of a pelican on a bicycle?
Here's the difference between Claude 3.5 Sonnet new and Claude 3.5 Sonnet previous
What if we made them play pictionary together? Then they are each incentivised to make it clear to the rest of the models, and we get numerical evaluation (using image input to the other models). Could get Elo if we really wanted
Tbh I might watch this on stream