THE TOP OF THE LADDER · AN EXPERIMENT IN PUBLIC

Three models, one judge.

I gave the exact same real-world task to Claude, ChatGPT, and Grok: audit a staffing agency invoice against our timeclock records and find every discrepancy. Same brief, word for word, fresh chats, no coaching. Their full answers will be published here — and then I grade them.

Because the real skill isn't using AI. It's knowing when to trust what comes back.

C

Claude

Anthropic

SITTING THE EXAM

C

ChatGPT

OpenAI

SITTING THE EXAM

G

Grok

xAI

SITTING THE EXAM

The rules

One task, three doors. A synthetic-but-realistic invoice audit — eight workers, one week, an invoice that doesn't add up. Every model gets the identical brief pasted into a fresh chat.

First answer counts. No retries, no coaching, no follow-up hints. If a model asks a question instead of answering, that gets noted too.

The grading is the product. I built the dataset, so I know every planted error — including one designed to test honesty: a discrepancy in our favor, not the vendor's. Models that only find the exciting overbilling will be marked accordingly.

Full disclosure. Claude helped build this site and this test — so its exam happens in a fresh chat with no memory of the answer key, same as the others.

RESULTS PENDING

The exams are out. The judge's bench is empty — for now.

This page ships before the results exist, because that's how this whole site works: I'm flying the plane while I build it. Come back to see three answer sheets, red ink and all.

New to all of this? Start at the shallow end: the five sentences you already know.