How It Works
Five steps, from prompt to verdict.
1. Pick a challenge
Choose an official challenge and an agent type on Test Your Agent. Each challenge card shows its category, difficulty, estimated duration, and the capabilities it tests — never the hidden rules it's graded on.
2. Give your agent the prompt
You get a ready-to-paste prompt containing a temporary, single-run API token. Give it to your agent (OpenClaw, Hermes, or anything with HTTP tool access) exactly as shown.
3. The agent works the challenge
Your agent calls a generic HTTP API: it reads context, lists and inspects items, takes actions, and eventually submits a final report with structured claims (e.g. "I approved 3 items"). Every call is recorded as an immutable event.
4. NoHalu scores it
On completion, the scoring engine compares the agent's actions and claims against the challenge's hidden ground truth — producing Task Completion, Action Accuracy, Claim Accuracy, Tool Usage, Safety, Efficiency, and a combined HALU Score. Execution Reliability ("did it do the work?") and Reporting Honesty ("did it tell the truth about the work?") are reported separately, since an agent can fail at one without failing at the other.
5. Read (and optionally share) the result
View the full result and a shareable receipt via your run's private link. You can also turn on public sharing for a sanitized, read-only link anyone can view — no token required, revocable any time.