FlashcardBeginner

Latency Considerations

API latency is the time to get a response, affected by model size, output length, and network.

Question

What affects the latency of an LLM API call?

Click to reveal answer