<feed xmlns="http://www.w3.org/2005/Atom"> <id>https://wanli-jiang.github.io/</id><title>Wanli Jiang</title><subtitle>A blog to record my learning and research.</subtitle> <updated>2026-08-18T13:11:33+00:00</updated> <author> <name>Wanli Jiang</name> <uri>https://wanli-jiang.github.io/</uri> </author><link rel="self" type="application/atom+xml" href="https://wanli-jiang.github.io/feed.xml"/><link rel="alternate" type="text/html" hreflang="en" href="https://wanli-jiang.github.io/"/> <generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator> <rights> © 2026 Wanli Jiang </rights> <icon>/assets/img/favicons/favicon.ico</icon> <logo>/assets/img/favicons/favicon-96x96.png</logo> <entry><title>廉价生成之后：AI Agent 与再生式软件工程</title><link href="https://wanli-jiang.github.io/posts/AI-agent-Cheap-generation-and-regenerative-software-engineering/" rel="alternate" type="text/html" title="廉价生成之后：AI Agent 与再生式软件工程" /><published>2026-08-18T00:00:00+00:00</published> <updated>2026-08-18T00:00:00+00:00</updated> <id>https://wanli-jiang.github.io/posts/AI-agent-Cheap-generation-and-regenerative-software-engineering/</id> <content type="text/html" src="https://wanli-jiang.github.io/posts/AI-agent-Cheap-generation-and-regenerative-software-engineering/" /> <author> <name>Wanli Jiang</name> </author> <category term="AI Agent" /> <category term="methodology" /> <summary>当代码、实验和方案的生成成本快速下降，工程的稀缺资源会转向反馈、验证、责任与长期判断力。本文提出“再生式软件工程”：一次自动化不仅要完成任务，还应积累验证资本、人类能力资本与架构选择权，让下一次变化更容易理解、验证、恢复和重新选择。</summary> </entry> <entry><title>Removing the Legacy TensorRT Backend from TensorRT-LLM: A Retrospective</title><link href="https://wanli-jiang.github.io/posts/trtllm-maintain-Removing-legacy-TensorRT-backend-from-TensorRT-LLM/" rel="alternate" type="text/html" title="Removing the Legacy TensorRT Backend from TensorRT-LLM: A Retrospective" /><published>2026-07-23T00:00:00+00:00</published> <updated>2026-07-23T00:00:00+00:00</updated> <id>https://wanli-jiang.github.io/posts/trtllm-maintain-Removing-legacy-TensorRT-backend-from-TensorRT-LLM/</id> <content type="text/html" src="https://wanli-jiang.github.io/posts/trtllm-maintain-Removing-legacy-TensorRT-backend-from-TensorRT-LLM/" /> <author> <name>Wanli Jiang</name> </author> <category term="LLM" /> <category term="engineering" /> <summary>A retrospective on removing the legacy TensorRT engine backend from TensorRT-LLM through small, reviewable, always-compilable pull requests.</summary> </entry> <entry><title>Piecewise CUDA Graph and Breakable CUDA Graph for LLM Inference</title><link href="https://wanli-jiang.github.io/posts/LLM-optimization-Piecewise-CUDA-Graph-and-Breakable-CUDA-Graph-for-LLM-Inference/" rel="alternate" type="text/html" title="Piecewise CUDA Graph and Breakable CUDA Graph for LLM Inference" /><published>2026-07-06T00:00:00+00:00</published> <updated>2026-07-25T00:52:58+00:00</updated> <id>https://wanli-jiang.github.io/posts/LLM-optimization-Piecewise-CUDA-Graph-and-Breakable-CUDA-Graph-for-LLM-Inference/</id> <content type="text/html" src="https://wanli-jiang.github.io/posts/LLM-optimization-Piecewise-CUDA-Graph-and-Breakable-CUDA-Graph-for-LLM-Inference/" /> <author> <name>Wanli Jiang</name> </author> <category term="LLM" /> <category term="optimization" /> <summary>A technical note on full, piecewise, and breakable CUDA Graphs in modern LLM and VLM inference systems.</summary> </entry> <entry><title>When NVLS Is "Supported" on H100, Why Does AllReduce Take Another Path?</title><link href="https://wanli-jiang.github.io/posts/LLM-optimization-NVLS-supported-on-H100-why-AllReduce-takes-another-path/" rel="alternate" type="text/html" title="When NVLS Is &amp;quot;Supported&amp;quot; on H100, Why Does AllReduce Take Another Path?" /><published>2026-06-23T00:00:00+00:00</published> <updated>2026-06-23T00:00:00+00:00</updated> <id>https://wanli-jiang.github.io/posts/LLM-optimization-NVLS-supported-on-H100-why-AllReduce-takes-another-path/</id> <content type="text/html" src="https://wanli-jiang.github.io/posts/LLM-optimization-NVLS-supported-on-H100-why-AllReduce-takes-another-path/" /> <author> <name>Wanli Jiang</name> </author> <category term="LLM" /> <category term="optimization" /> <summary>A TensorRT-LLM debugging story about separating static NVLS capability from fabric handle usability and single-node POSIX-FD allocation.</summary> </entry> <entry><title>GDPval Bench introduction</title><link href="https://wanli-jiang.github.io/posts/LLM-bench-GDPval-Bench-introduction/" rel="alternate" type="text/html" title="GDPval Bench introduction" /><published>2026-06-01T00:00:00+00:00</published> <updated>2026-06-01T00:00:00+00:00</updated> <id>https://wanli-jiang.github.io/posts/LLM-bench-GDPval-Bench-introduction/</id> <content type="text/html" src="https://wanli-jiang.github.io/posts/LLM-bench-GDPval-Bench-introduction/" /> <author> <name>Wanli Jiang</name> </author> <category term="LLM" /> <category term="bench" /> <summary>An exploration of GDPval, OpenAI's benchmark for evaluating AI models on realistic, economically valuable knowledge-work tasks.</summary> </entry> </feed>
