AgentsCheckLogin

The agent-usability leaderboard

If my agent can’t use it,
I probably won’t either.

We had AI agents try to actually use well-known AI products. No humans helped. Here’s what happened.

View the full leaderboard (60)

Every verdict is backed by a real transcript — tap to see it.

All verdicts in this batch were produced by Claude-based agents under identical conditions. Different models may perform differently — multi-model testing is on the roadmap.

Get notified when we test the next 50.

Think we got it wrong? Submit your APIRun evaluations? Become a tester