THE TOP OF THE LADDER · AN EXPERIMENT IN PUBLIC
Three models, one judge.
I gave the exact same real-world task to Claude, ChatGPT, and Grok: audit a staffing agency invoice against our timeclock records and find every discrepancy. Same brief, word for word, fresh chats, no coaching. Their full answers will be published here — and then I grade them.
Because the real skill isn't using AI. It's knowing when to trust what comes back.
Claude
Anthropic
SITTING THE EXAM
ChatGPT
OpenAI
SITTING THE EXAM
Grok
xAI
SITTING THE EXAM
The rules
One task, three doors. A synthetic-but-realistic invoice audit — eight workers, one week, an invoice that doesn't add up. Every model gets the identical brief pasted into a fresh chat.
First answer counts. No retries, no coaching, no follow-up hints. If a model asks a question instead of answering, that gets noted too.
The grading is the product. I built the dataset, so I know every planted error — including one designed to test honesty: a discrepancy in our favor, not the vendor's. Models that only find the exciting overbilling will be marked accordingly.
Full disclosure. Claude helped build this site and this test — so its exam happens in a fresh chat with no memory of the answer key, same as the others.
RESULTS PENDING
The exams are out. The judge's bench is empty — for now.
This page ships before the results exist, because that's how this whole site works: I'm flying the plane while I build it. Come back to see three answer sheets, red ink and all.
New to all of this? Start at the shallow end: the five sentences you already know.