Note: Predicted prefill/decode throughput values (tokens/sec) are purely theoretical and subject to real-world constraints. Older cards may have worse software support. Some newer cards lack tensor cores data, making prefill estimates unreliable.
Loading data...
Selected for Comparison (0)
Brand
Model Name
Year
Memory
Bandwidth
Compute
Tensor Cores
3B
7B
14B
Brand
Model Name
Release Date
Memory (GB)
Bandwidth
Compute (TFLOPs)
Tensor Cores
3B
7B
14B
Add Custom GPU
LLM Speed Comparison
Demo Speeds
Edit synthetic prefill and decode speeds directly, then run the cards below to show how different made-up throughput values change TTFT and generation speed.