Skip to main content

What is Yardstick?

Measure the value your AI agents create, rank it, and govern the spend behind it.

Yardstick is one platform that measures the value AI agents create, ranks it, and governs the spend behind it. Instead of guessing whether an AI coding tool is worth it, you see the work it actually shipped and what that work cost.

The classic example is engineering: Yardstick ties the merged pull requests an AI coding agent helped ship to the tokens it burned, so you get a real number like cost per merged PR and a clear AI-versus-human comparison. The same model extends to support bots, SDRs, and research workflows, because every one of them is an agent that consumes tokens and produces outcomes.

Yardstick has three parts that build on each other:

  • Observe measures the value your agents create (outcomes, cost, scores).

  • Optimize turns those scores into better agents with ranked changes and

experiments. It is rolling out now and is included on Pro & Scale as each tool ships.

  • Treasury governs the budget behind it all with budgets, burn tracking,

alerts, and the use-it-or-cash-it allowance.

Yardstick is in private beta. Join the waitlist to get access and a design-partner slot.

Did this answer your question?