AI Glossary term

Latency

How long a model takes to respond - often measured as time to the first token, then tokens per second.

Also called: speed, time to first token

How long a model takes to respond - often measured as time to the first token, then tokens per second.

Why it mattersA model that's 5% smarter but three times slower loses in most real products.

See also Inference, Streaming

Explore more AI terms

Browse the full glossary for plain-English definitions across models, agents, data, and safety.

Back to all terms