Guest on AI Engineer.

The Geopolitics of AI Infrastructure - Dylan Patel, SemiAnalysis
Jun 19, 2025 · 18:29
Dylan Patel of SemiAnalysis argues that despite US sanctions, Huawei has engineered a 384-chip cluster (Cloud Matrix 384) that Nvidia failed to deploy, while accessing TSMC via Softgo and HBM from Samsung via shell companies — all legally. China's SMIC will soon produce 7nm AI chips in high volumes, debunking the notion that China lacks compute. Meanwhile, Middle East players like G42 (UAE) and Datavolt (Saudi Arabia) are building multi-gigawatt data centers, with G42's deal letting it keep 20% of 500,000 GPUs yearly for itself while 80% goes to US companies like OpenAI. Patel highlights the US's 63-gigawatt power shortfall vs. 100 GW of planned data centers, explaining why US companies rely on Middle East capacity and why China's superior power buildout gives it a geopolitical edge.

System Design for Next-Gen Frontier Models — Dylan Patel, SemiAnalysis
Feb 11, 2025 · 18:29
Dylan Patel of SemiAnalysis breaks down the inference challenges for next-generation frontier models like GPT-4 (1.8 trillion parameters) and upcoming models trained on 100,000+ GPU clusters. He emphasizes that prefill (prompt processing) is compute-intensive while decode (token generation) is memory bandwidth-intensive, creating a systems problem where serving 64 users at 30 tokens/second requires 60 terabytes/second of memory bandwidth. Patel details engineering strategies such as continuous batching to improve batch utilization by 10-100x, disaggregated prefill to isolate noisy neighbors and maintain time-to-first-token SLAs, and context caching (like Google's) to cache KV cache on CPU/storage instead of GPU memory, dramatically reducing prefill costs. He warns that open-source tools like LLaMA.cpp lack these optimizations, making high-performance serving of models like LLaMA 405b infeasible without libraries like vLLM or TensorRT-LLM. On scaling, Patel notes that 100,000 GPU clusters (e.g., Microsoft's Arizona data center consuming 150 MW) face reliability issues — optical transceivers fail every five minutes — and straggler chips (silicon lottery) can degrade training…
Powered by PodHood