Archive
All Posts
Every post by AI engineer HaRyeom Jang (18), grouped by year.
202618
- InferenceGPU Memory Fragmentation in LLM Serving: Causes and Measurement
- InferenceMixture of Experts vs Mixture of Tokens: Two Sparsity Strategies Compared
- InferenceWhy Chunked Prefill Exists: How Prefill Stalls Decode and the Fix
- InferenceKV Cache Quantization: Why It's Different from Weight Quantization
- InferencePrefix Caching Deep Dive: Why TTFT Drops but Throughput Doesn't
- InferenceWhy TTFT, TPOT, and p99 Latency Move Independently in LLM Serving
- InferenceContinuous vs Static Batching: Throughput Gains and When They Reverse
- InferenceMixture of Tokens vs MoE: Two Paths to Sparse Activation
- InferenceLLM Serving Schedulers: Who Decides When Requests Hit the GPU
- InferenceWhere Quantization Loses Accuracy: W4A16, W8A8, and FP8 Compared
- LLMMixture of Experts vs Mixture of Tokens: Key Differences and Use Cases
- InferenceKV Cache Memory Calculations: Context Length vs. Batch Size
- InferenceWhy LLM Serving Needs Separate Prefill and Decode Stages
- InferencePrefill-Decode Disaggregation: Why Mixing Both Phases on One GPU Pool Hurts
- LLMRoPE: Rotary Position Embedding in Transformers
- LLMSinusoidal Positional Encoding in Transformers
- LLMTransformer Deep Dive: Token Embeddings Explained
- LLMHow Transformers Work, Part 1: Tokenizers