The Audit

The $1 audit that reads your work order like a hostile reviewer

Most agent failures are decided before the agent ever runs. The goal is mushy, the checks are unverifiable, and the boundaries are missing. The audit finds those problems while the only thing at risk is one dollar.

What the audit does

You hand it a work order: a goal, acceptance criteria, non-goals, and a budget. The auditor reads it like a reviewer who is trying to break it. It looks for the specific failure modes that make agent runs burn time and money:

  • Acceptance criteria that cannot be proven by a command, URL, file, or human check
  • Goals that describe activity instead of a finished state
  • Missing boundaries that let the agent drift into extra work
  • Checks that the agent could pass vacuously without doing the job
  • A contract too vague for any verifier to grade honestly

What you get back

A graded dry run with an exit code, notes on every acceptance criterion, and a repaired work order. When the auditor emits a repaired form, the site shows a Use audit repairs button that drops the cleaned contract straight back into the form. Review it, edit it, and you are ready to run.

Example audit output

This is what an audit verdict looks like. The auditor grades the contract, predicts the exit code, and returns repaired criteria you can apply in one click.

<verdict>
  grade: A
  predicted_exit: 0 SUCCESS
  weakest_link: All criteria cite external sources; verification is possible.
  one_fix: Add a check for wavelength-dependent scattering intensity formula.
</verdict>

--- Repaired Criteria ---

[AUTO] Cite peer-reviewed source (Bohren & Clothiaux 2006 or equivalent 
       physics textbook) defining Rayleigh scattering
[AUTO] State the 1/wavelength^4 relationship for scattering intensity
[AUTO] Explain cone-cell sensitivity difference: human eyes are more 
       sensitive to blue (450-495nm) than violet (380-450nm)

--- Dry Run Notes ---

Iteration 1: Found NASA Science page on blue skies. Partial evidence.
Iteration 2: Located HyperPhysics Rayleigh scattering page. Formula confirmed.
Iteration 3: CIE 1931 cone sensitivity data found. All criteria satisfied.

Predicted exit: SUCCESS. Contract is tight enough for a MECHA run.

The audit does not execute the job. It stress-tests the contract so you can fix problems before spending on a MECHA run.

Why a $1 dry run beats a $10 burn

A MECHA run is real execution: a swarm of agents, a sandbox, model costs, minutes of your time. Running that on a contract nobody checked is how you get a confident-looking report about the wrong thing. The audit is the cheap pass that catches the contract problems while the only thing at risk is a dollar.

What the audit is not

  • Not a code review of your repo. It reviews the work order, not the codebase.
  • Not a guarantee. It raises the odds that a run verifies cleanly; it cannot make a bad job good.
  • Not a human consultant. It is an AI reviewer with a strict protocol and an honest failure mode.

Chat with Grok to build a work order, then run the $1 audit before you spend on a MECHA run.

Put your next agent task through the press

Talk to Grok, get a tight work order, and hit it with a $1 audit before anything expensive runs.