For every measured call, Yardstick multiplies token usage by the price for that model. We maintain rates for the major families, including Anthropic Claude, OpenAI (GPT & the o-series), Google Gemini, Mistral, DeepSeek, Cohere, xAI Grok, Llama, and Qwen, so calls captured through the gateway or OpenTelemetry get a real dollar cost rather than zero. Rates for fast-moving providers are best-estimates we keep current; if you see a cost that looks off, contact us and we will check the rate.
How Yardstick computes model cost
Token usage times current provider rates, across the major model families.
