Measuring value (Observe)
8 articles
- How scoring works: value against spendYardstick ranks agents on the value they created versus what they cost.
- Cost per merged PRThe headline engineering metric: spend divided by pull requests shipped.
- AI versus human attributionSee what AI shipped versus your team, at the pull-request level.
- Change-failure & revertsQuality matters. Reverted or fixed work counts against the score.
- Loop engineers & autonomous runsAutonomous agents are measured as runs: cost, iterations & whether they merged.
- Velocity versus qualityThe honest chart: is shipping faster with AI quietly raising your defect rate?
- Cost per resolved issueSpend divided by the tickets your agents actually closed.
- How Yardstick computes model costToken usage times current provider rates, across the major model families.
