Skip to main content

How Yardstick computes model cost

Token usage times current provider rates, across the major model families.

For every measured call, Yardstick multiplies token usage by the price for that model. We maintain rates for the major families, including Anthropic Claude, OpenAI (GPT & the o-series), Google Gemini, Mistral, DeepSeek, Cohere, xAI Grok, Llama, and Qwen, so calls captured through the gateway or OpenTelemetry get a real dollar cost rather than zero. Rates for fast-moving providers are best-estimates we keep current; if you see a cost that looks off, contact us and we will check the rate.

Did this answer your question?