84 items

1

Measuring Time-to-First-Token

Code Quiz
2

Cost Estimation for LLM Workloads

Slides / Video
3

Setting max_tokens to Control Cost

Code Quiz
4

Deploying an LLM Endpoint

Code Quiz
5

Concurrent Async Requests for Throughput

Code Quiz
6

Fallback to Alternate Model

Code Quiz
7

Tracking Per-Request Cost Metrics

Code Quiz
8

Semantic Cache with Embedding Similarity

Code Quiz
9

Choosing Embedding Batch Size

Code Quiz
10

Using Prompt Caching for Context

Code Quiz
11

Model Routing for Simpler Tasks

Code Quiz
12

Handling 429 Rate Limit Errors

Code Quiz
13

Truncating Conversation History

Code Quiz
14

Load Balancing Across API Keys

Code Quiz
15

Batching Requests to Reduce Cost

Code Quiz
16

Estimating Cost from Token Usage

Code Quiz
17

Retry with Exponential Backoff

Code Quiz
18

Token Bucket Rate Limiting

Code Quiz
19

Reducing Latency with Streaming

Code Quiz
20

Response Cache for Repeated Prompts

Code Quiz
21

Prompt Compression to Reduce Tokens

Code Quiz
22

Reusing the HTTP Client

Code Quiz
23

Adding Request Timeouts

Code Quiz
24

Cost/Latency Observability

Quiz
25

Throughput and Tokens-per-Second

Quiz
26

Prompt Length Impact

Quiz
27

GPU Requirements for Self-Hosting

Quiz
28

Output Token Length Impact

Quiz
29

Model Size vs Cost Tradeoff

Quiz
30

Handling Rate Limits

Quiz
31

Prompt Caching

Quiz
32

Semantic Response Caching

Quiz
33

Timeouts and Graceful Degradation

Quiz
34

Provisioned vs On-Demand Pricing

Quiz
35

Choosing a Cheaper Model

Quiz
36

TTFT vs Total Latency

Quiz
37

Fallback and Model Routing

Quiz
38

Load Balancing Across Endpoints

Quiz
39

Autoscaling LLM Workloads

Quiz
40

Inference Servers Overview

Quiz
41

Quantization for Cheaper Inference

Quiz
42

Self-Hosting vs Managed API

Quiz
43

Concurrency and Parallel Requests

Quiz
44

Estimating Token Costs at Scale

Quiz
45

Batching for Throughput

Quiz
46

Reducing Cost via Prompt Compression

Quiz
47

Monitoring and Observability for LLM Systems

Slides / Video
48

Retry, Timeout & Fallback Strategies

Slides / Video
49

Quantization for Faster, Cheaper Inference

Slides / Video
50

Model Cascading and Routing

Slides / Video
51

Semantic Caching for Similar Queries

Slides / Video
52

Caching Responses and Prompt Caching

Slides / Video
53

Prompt Compression to Cut Tokens

Slides / Video
54

Choosing Smaller Models to Cut Costs

Slides / Video
55

Autoscaling Based on Traffic Demand

Slides / Video
56

Horizontal Scaling & Load Balancing Replicas

Slides / Video
57

Batching Requests to Improve Throughput

Slides / Video
58

Throughput vs Latency Trade-offs

Slides / Video
59

KV Cache

Flashcard
60

Token Optimization for Cost

Flashcard
61

Inference Servers

Flashcard
62

Monitoring and Observability

Flashcard
63

Model Cascading and Routing

Flashcard
64

Cold Starts and Warm Pools

Flashcard
65

Speculative Decoding

Flashcard
66

Provisioned Throughput vs Pay-Per-Token

Flashcard
67

Output Length Control

Flashcard
68

Horizontal Scaling and Load Balancing

Flashcard
69

GPU Selection

Flashcard
70

Prompt Caching

Flashcard
71

Throughput Metrics

Flashcard
72

Key Latency Metrics

Flashcard
73

Quantization

Flashcard
74

Self-Hosting vs Managed API

Flashcard
75

Autoscaling for Traffic Spikes

Flashcard
76

Cost/Latency/Quality Triangle

Flashcard
77

Model Selection for Cost/Latency

Flashcard
78

Semantic Caching

Flashcard
79

Context Window Size Impact

Flashcard
80

Batching and Continuous Batching

Flashcard
81

TTFT and Tokens Per Second

Slides / Video
82

GPU Selection for LLM Serving

Slides / Video
83

Model Serving Infrastructure & Inference Servers

Slides / Video
84

Hosted API vs Self-Hosted Models

Slides / Video
Slides / VideoIntermediate

GPU Selection for LLM Serving

Learn how to choose GPUs and hardware for serving LLMs based on memory, throughput, and cost.

Slide 1
1 / 6