Tokens per second can be misleading for AI agents. Prefill, KV cache, context length, tool calls, and latency often matter more than raw generation speed.