Skip to main content
Choosing the Right AI Model

LESSON 2 OF 4

Test it on your own work

BY THE END OF THIS LESSON

Run a comparison that tells you something about your tasks, not someone else’s.

Why a leaderboard cannot answer your question

The weak version and the strong one.
AvoidA published rankinga fixed set of problems
Do thisYour materialyour work is not on that list

A published ranking measures performance on a fixed set of problems. Your work is not on that list. A tool that scores well on the tests and badly on your material is entirely possible, and you would never know from the score.

The ranking is someone else’s answer to someone else’s question.

Build a small set of your own

1 of 5 to check carefully.
  • Safe: A typical one
  • Safe: An awkward one
  • Safe: A long one
  • Safe: One where the tone matters
  • Check this: One where being wrong would cost you

Take five things you have actually done: a typical one, an awkward one, a long one, one where the tone matters, and one where being wrong would cost you something.

Five is enough to separate tools and small enough that you will really run it. Keep the inputs and keep the outputs — the set is only useful if you can run it again later.

Judge it before you look

3 stages, each leading to the next.
  1. Decide what good looks like
  2. Then run it
  3. Judge by the criteria

Decide what a good answer looks like before you run anything. Otherwise you read the output and then decide, which is how you end up preferring whichever one sounded most confident.

Write the criteria down: did it follow the constraints, is every fact checkable, would I send this after one edit or three?

WORKED EXAMPLE

DECIDING BY IMPRESSION

Tried the same question in both. The second one felt better, so I switched.

DECIDING BY CRITERIA

Same five real tasks in both. Criteria written first: obeyed the word limit, no invented figures, kept our banned words out, usable after one edit. Result: one obeyed the word limit every time and the other did not, and that was the whole difference for my work.

"Felt better" is unreproducible and usually means "sounded more confident". Criteria written in advance give you a result you can act on, and can re-run in six months when both tools have changed.

YOUR TURN

Build the five-task set before you compare anything.

RUN THIS

Help me build a small evaluation set for my own work. Here is what I do: [describe your work]. Suggest five specific tasks I should use — typical, awkward, long, tone-sensitive, and one where being wrong would cost me something. For each, write the criteria I should judge the answer against, in a form I could tick off.

What a good result looks like

Every criterion should be checkable by looking — "under 120 words", "no figure I cannot verify". Anything like "reads well" needs replacing with what you would actually check.

KNOWLEDGE CHECK

Answer all 3 correctly to complete this lesson.

  1. 1. Why is a published ranking a weak basis for your choice?

  2. 2. Put a fair comparison in order.

    Click them in order, first to last.

  3. 3. Why write the criteria before running the comparison?

REMEMBER THIS

Five of your own tasks, judged against criteria written in advance, beats any ranking.