CUDA 3 Removing the Legacy TensorRT Backend from TensorRT-LLM: A Retrospective Jul 23, 2026 Piecewise CUDA Graph and Breakable CUDA Graph for LLM Inference Jul 6, 2026 When NVLS Is "Supported" on H100, Why Does AllReduce Take Another Path? Jun 23, 2026