Inference & Serving · Established · Beginner
TPS
Also known as: Tokens Per Second
An industry-standard acronym standing for Tokens Per Second within inference & serving.
What TPS is
TPS (Tokens Per Second) is a fundamental terminology concept used throughout inference & serving.
How it works
Serves as shorthand in technical documentation, research papers, and engineering discussions.
Why it matters
Mastering TPS helps developers and researchers communicate effectively across the AI ecosystem.
Common uses
- →Inference & Serving terminology
- →Technical documentation
- →Research literature
More in this collection
Browse all AI A–Z