Time to first token
Review this signal against your actual prompt distribution and provider documentation.
Compare operational model signals, pricing, context, and available performance evidence for latency-sensitive applications. Values below come from the canonical AIBuzzHub model profile and are marked as not verified when unavailable.
Latency-sensitive systems should evaluate time to first token, sustained throughput, provider region, and workload variance alongside price.
Review this signal against your actual prompt distribution and provider documentation.
Review this signal against your actual prompt distribution and provider documentation.
Review this signal against your actual prompt distribution and provider documentation.
Operational fields and scores are only decision aids. Confirm methodology, model version, region, and workload-specific behavior before shipping.
Paid-tier introductory list pricing through 2026-12-31; Google lists $1.50 input and $7.50 output beginning 2027-01-01. Free-tier pricing is not represented in this tracker row.