廉价生成之后:AI Agent 与再生式软件工程
当代码、实验和方案的生成成本快速下降,工程的稀缺资源会转向反馈、验证、责任与长期判断力。本文提出“再生式软件工程”:一次自动化不仅要完成任务,还应积累验证资本、人类能力资本与架构选择权,让下一次变化更容易理解、验证、恢复和重新选择。
当代码、实验和方案的生成成本快速下降,工程的稀缺资源会转向反馈、验证、责任与长期判断力。本文提出“再生式软件工程”:一次自动化不仅要完成任务,还应积累验证资本、人类能力资本与架构选择权,让下一次变化更容易理解、验证、恢复和重新选择。
A retrospective on removing the legacy TensorRT engine backend from TensorRT-LLM through small, reviewable, always-compilable pull requests.
A technical note on full, piecewise, and breakable CUDA Graphs in modern LLM and VLM inference systems.
A TensorRT-LLM debugging story about separating static NVLS capability from fabric handle usability and single-node POSIX-FD allocation.
An exploration of GDPval, OpenAI's benchmark for evaluating AI models on realistic, economically valuable knowledge-work tasks.
A quick survey of the agent-facing skills and AGENTS.md guides shipped by vLLM, SGLang, FlashInfer, and TensorRT-LLM, so coding agents (and humans) can be productive in each codebase.
A comprehensive technical guide — from collective communication fundamentals to fused kernel internals, with Nemotron-H as a running example.
A work log to record the optimization steps for Nemotron v3 Super in TRTLLM.
An introduction to TerminalBench, a benchmark designed to evaluate how well AI agents complete complex, real-world tasks in terminal environments.
Goal Setup personal blog on github.io so that I can share my personal experiences about deep learning, large language model, and AI-infra optimization. Setup about the blog To set up this blog, ...