《GPU Zen 4: Advanced Rendering Techniques》第 1 章中文译稿。原文作者:Dominik Lazarek、Philip Hammer、Jean Geffroy。
系列导航: 第一篇:从 Forward+ 到 Triangle Visibility Buffer|第二篇:Pixel Command 与 Deferred Lighting|第三篇:Variable-Rate Compute Shaders
译者导读
前两篇梳理了从 Triangle Visibility Buffer 到 Deferred Texturing、Deferred G-Buffer Update 和 Deferred Lighting 的整条主线。这一篇转向专为 Compute Workload 设计的 Variable-Rate Compute Shaders(VRCS):Shading Rate Image 如何生成,Primary Pixel 如何集中到连续 wave,Duplicate Pixel 又如何回填。最后再结合性能与显存数据,看看整套 idTech 8 Pipeline 的实际代价和收益。
1.10 Variable-Rate Compute Shaders(VRCS)
Variable-Rate Compute Shaders(VRCS)最早由 [Fuller 19] 提出,当时称为 Sparse Lighting,之后又见于 [Fuller 22]。它是 Variable Rate Shading(VRS)针对 Compute-Based Rendering Workload 的变体,与 idTech 8 的新 Pipeline 十分契合。
2021 年,《毁灭战士:永恒》在 Xbox Series X|S 更新中加入 Hardware VRS,具体实现见 [Fuller et al. 22]。VRS 很适合当时的 Forward+ Rendering Pipeline,能节省大量 GPU 时间,使团队得以提高画质与目标分辨率。
Hardware VRS 只适用于传统 Vertex–Fragment Pipeline,并非所有平台都支持,基础版 PlayStation 5 就是一个例子。idTech 8 的新 Pipeline 又以 Compute 为主,传统 VRS 能覆盖的只剩 Transparent Object 和少量仍走 Forward+ 的 model。减少 Particle 等内容的 Fragment Shader Invocation 当然依然有价值,但《毁灭战士:黑暗时代》的主要开销已经转移到 Compute-Based Opaque Rendering。开发团队需要的是一套能适配新 Pipeline 的 VRS 方案。
VRCS 沿用了 VRS 的基本观察:并非所有 object 或 screen region 都需要相同的 Shading Rate(SR)。这里的 Shading Rate,表示多少像素共享一次独立的 Shading Calculation。Full-Rate Shading(SR 1 x 1)会为每个可见像素计算一次;Coarse Shading 的计算次数则少于输出像素数,例如四像素 quad 只计算一到两次。
晴朗天空或平坦墙面没有多少高频细节,通常不必逐像素计算;有些区域只对一半或四分之一的像素执行独立着色就够了。没有 VRS 时,每个 screen pixel 都要运行一个独立 Shader Thread。启用 Coarse Shading 后,一个 pixel quad 只执行一到两个线程,再把结果复制给不再单独计算的邻近像素。计算量由此显著下降,但输出带宽不会减少,因为每个像素仍然要写入。VRS 的完整规范可参见 [Microsoft 25]。
VRCS 的细节足以单独写一篇文章,完整说明可参见 [Fuller 22] 与 [Hammer and Fuller 25]。这里先交代理解方案所需的概念,再看《毁灭战士:黑暗时代》如何使用 VRCS、哪些 Compute Pass 从中受益,以及开发团队从中得到了哪些经验。
1.10.1 生成 Shading Rate Image
Shading Rate 由一张 screen-space image 指定,图中的每个值对应一组像素,也就是一个 tile。Compute Shader 可以生成与 16 x 16、8 x 8 或 2 x 2 等 tile size 对应的图像。Hardware VRS 的 tile 通常由硬件规定为 16 x 16 或 8 x 8;完全由软件实现的 VRCS 则采用 2 x 2。这张图像就是 Shading Rate Image(SRI)。
图 1.16 展示 Hebeth 场景。左上是最终画面,右上是 luminance;下排分别是 8 x 8 tile 的 SRI,以及 VRCS 使用的更细 2 x 2 SRI。8 x 8 SRI 供启用 Hardware VRS 的 Raster Pass 使用,2 x 2 SRI 则供部分 Compute Shader Workload 使用。
图像格式为 R8_UINT,不同 Shading Rate 编码在低位。调试颜色直接对应这些 Rate:红色为 1 x 1,绿色和青色分别为 1 x 2 与 2 x 1,蓝色为 2 x 2。

图 1.17 展示 tile Shading Rate 与 pixel quad 配置之间的映射,以及 Primary Pixel 和 Duplicate Pixel 的关系。最终实际计算哪些像素,由 Shading Rate 与 triangle coverage 共同决定。

红色 1 x 1 tile 逐像素计算,成本最高;蓝色 2 x 2 tile 每四个像素只执行一次独立着色计算,成本最低;绿色与青色的 1 x 2 / 2 x 1 Half-Rate Tile 介于两者之间。
8 x 8 VRS SRI 与 2 x 2 VRCS SRI 由同一个 Compute Shader 生成。Pass 分析上一帧的 Luminance Image 与当前帧 Depth Buffer,判断各 tile 需要的 Shading Rate。SRI 有多种生成方式;这里使用的是基于 [Fuller 19] Sobel Filter 的改进算法,分析 screen luminance 的 gradient delta,同时考虑 depth discontinuity。算法可以调整 Shading Rate 与 Depth Discontinuity Tolerance,也能加入额外 Rate 处理边缘情况。
为了避免与 Async Queue 争用,上一帧的 luminance 会在最后一个 Post-Processing Step 中写入独立的 Image Buffer,写入时机位于 Antialiasing 和 Tonemapping 之后。
1.10.2 把 Pixel Thread 打包进 Compute Wave
SRI 已经确定每个 screen quad 的 Shading Rate,接下来必须把它用于实际 workload。最朴素的做法是:
- Shader 开始时采样 SRI,确定当前 pixel quad 的 Rate。
- 如果是 1 x 1,照常执行。
- 如果 Rate 更低,只计算一到两个 Primary Pixel,把结果复制给两到三个 Duplicate Pixel,并让 Duplicate Pixel 对应的线程 early-out。
这套逻辑能工作,却不会自然转化为明显的性能收益。GPU 以 wave 为单位执行一组线程,并不单独调度每个线程。只跳过 wave 中的一部分线程,通常仍要承担整条 wave 执行完整 Shader 的成本。
要真正从 Coarse Shading 获益,必须把负责 Primary Pixel 的线程与 Duplicate Pixel 的线程分到不同 wave,如图 1.18。Primary Pixel 执行实际工作并把结果复制给 Duplicate Pixel;Duplicate Pixel 运行同一 Shader,但会很早退出。

只包含 Duplicate Pixel 的完整 wave 会一致命中 early-out,跳过昂贵计算,并很快由 GPU Command Processor 回收到可用池中。系统仍会 Dispatch 足够多的 wave,以覆盖所有像素都需要处理的最坏情况;其中没有多少工作的 wave 则会提前结束。
为了避免输出图像出现空洞,Primary Pixel Thread 会在结束前把结果复制到位置已知的邻近 Duplicate Pixel。也可以完全省掉复制,以节省 Memory Bandwidth;代价是 SSDO 或其他 Image Filtering 等所有可能读到缺失像素的后续 Shader,都必须能处理 Coarse Shading Tile。团队测试后确实看到了更多收益,但受开发时间限制,这套做法没有随游戏发布。
屏幕被划分为 16 x 16 pixel 的 tile grid,每个 tile 正好由八个 wave32 计算,threadgroup dimension 为 8 x 4(图 1.19)。每个 tile 内的 Primary Pixel Coordinate 必须被重映射到连续 wave(图 1.20)。所有启用 VRCS 的 Pass 都必须使用相同的 wave32 + 8 x 4 threadgroup 布局,否则无法还原原始 Thread Position。
选择 wave32,是因为现代 GPU 普遍支持,而且相比 wave64,部分填充 wave 的浪费较少。Intel 等原生 wave16 GPU 需要模拟 wave32,例如使用 Vulkan 的 VkPipelineShaderStageRequiredSubgroupSizeCreateInfo。


以包含 132 个 Primary Pixel、124 个 Duplicate Pixel 的 16 x 16 tile 为例,线程可排成:四个由 32 个 Primary Pixel 填满的 wave;一个包含 4 个 Primary Pixel 与 28 个 Duplicate Pixel 的部分填充 wave;三个全部由 Duplicate Pixel 组成的完整 wave。
图 1.21 是未重映射时按 screen coordinate 调度的默认布局。Primary 与 Duplicate Pixel 混在每条 wave 中,即使 VRCS 已经启用,仍需支付全部 wave 成本。

图 1.22 是重映射后的布局。四个完整 Primary wave 仍需支付完整 Shader 成本;三个只含 Duplicate Pixel 的 wave 可以立即退出并由 GPU 快速复用。部分填充的 wave 4 虽然只有 4/32 线程做有效工作,性能特征仍接近完整 wave。

Pixel Coordinate Remapping 需要在 GPU Buffer 中额外保存逐 tile Primary Pixel Count,以及逐像素 payload vrcsData:
- 4 bit 保存 tile-relative x coordinate。
- 4 bit 保存 tile-relative y coordinate。
- 3 bit 保存水平、垂直、对角线复制指令。
执行实际 workload 时,先以 global thread ID,也就是 HLSL 的 SV_DispatchThreadID 为基准,再通过 tile-relative coordinate 还原 16 x 16 VRCS tile 内真正的 Pixel Location。
图 1.23 的调试视图中,黄色块表示整个 16 x 16 tile 的全部线程都是 Primary Pixel;绿色是 Primary Pixel;蓝色是从邻近 Primary Pixel 复制结果的 Duplicate Pixel;粉色区域表示整个 VRCS tile 都采用 2 x 2 Shading Rate,这种情况下无需额外保存像素坐标,只需通过三个 Copy Flag 即可解析。


局部放大后,可以清楚看到绿色 Primary Thread 与蓝色 Duplicate Thread 在 screen space 中的排列。共享 Shading Result 时还会检查 triangle coverage:结果只在同一三角形内部共享,或者在 face angle 与 depth 十分接近的三角形之间共享。图 1.25 展示了根据 Visibility Buffer 中 Primitive ID 得到的 Triangle Coverage;实际实现可以高效计算二值的水平与垂直 Primitive ID Gradient。

1.10.3 把 VRCS 应用于 Compute Workload
所有使用 VRCS 的 Workload Shader 都必须以 wave32 模式和 8 x 4 threadgroup layout 进行 Dispatch,并加入两个步骤:Shader 开头调用 Prefix Function,计算实际 screen-space pixel position 并取得 Copy Bit;Shader 结束时调用 Postfix Function,将结果复制到邻近像素以填充空洞。
#define VRCS_TILE_SKY 0
#define VRCS_TILE_LIGHT_ALL_PIXELS 1
#define VRCS_TILE_WORK_FOR_THREAD 2
#define VRCS_TILE_NO_WORK_FOR_THREAD 3
#define VRCS_TILE_SIZE 16
#define VRCS_DATA_GET_X(vrcsData) ((vrcsData >> 0) & 0xf)
#define VRCS_DATA_GET_Y(vrcsData) ((vrcsData >> 4) & 0xf)
#define VRCS_DATA_COPY_H(vrcsData) ((vrcsData & 0x100) > 0)
#define VRCS_DATA_COPY_V(vrcsData) ((vrcsData & 0x200) > 0)
#define VRCS_DATA_COPY_D(vrcsData) ((vrcsData & 0x400) > 0)
uint VRCSPrefix(inout int2 coord, inout uint vrcsData) {
// There are 8 wave32s per 16x16 VRCS tile, calculate which 16x16 tile
// we're dealing with
const uint2 tileId = coord.xy / VRCS_TILE_SIZE;
const uint tileIndex =
tileId.y * constants.vrcsBufferWidth + tileId.x;
const uint vrcsCount = constants.vrcsCountBuffer[tileIndex];
const uint coordIndex = ((coord.y & 15) * 16) + (coord.x & 15);
BRANCH if (vrcsCount == 0) {
return VRCS_TILE_SKY;
}
BRANCH if (vrcsCount > (256 - 32)) {
return VRCS_TILE_LIGHT_ALL_PIXELS;
}
BRANCH if (coordIndex >= vrcsCount) {
return VRCS_TILE_NO_WORK_FOR_THREAD;
}
vrcsData = vrcsCoordsBuffer[
tileIndex * MAX_COORDINATES_PER_SCREEN_TILE + coordIndex];
coord.x =
int(VRCS_DATA_GET_X(vrcsData) + (VRCS_TILE_SIZE * tileId.x));
coord.y =
int(VRCS_DATA_GET_Y(vrcsData) + (VRCS_TILE_SIZE * tileId.y));
return VRCS_TILE_WORK_FOR_THREAD;
}
void VRCSPostFix(int2 coord, uint vrcsData, float4 outputValue) {
imageStore(targetImageUAV, coord, outputValue);
int2 copyCoord = coord.xy ^ 0x1;
// Do the pixel copies (hole filling)
if (VRCS_DATA_COPY_H(vrcsData)) {
imageStore(targetImageUAV, int2(copyCoord.x, coord.y),
outputValue);
}
if (VRCS_DATA_COPY_V(vrcsData)) {
imageStore(targetImageUAV, int2(coord.x, copyCoord.y),
outputValue);
}
if (VRCS_DATA_COPY_D(vrcsData)) {
imageStore(targetImageUAV, int2(copyCoord.x, copyCoord.y),
outputValue);
}
}
void main() {
int2 pixelLocation = compute.globalInvocationID.xy;
uint vrcsData = 0;
if (constants.vrcsEnabled == true) {
uint vrcsCode = VRCSPrefix(pixelLocation, vrcsData);
// Early-out directly if there is nothing to do at all
if ((vrcsCode == VRCS_TILE_SKY) ||
(vrcsCode == VRCS_TILE_NO_WORK_FOR_THREAD)) {
return;
}
}
float4 outputValue;
// Do shader work, e.g., lighting calculations, etc.
// ...
if (constants.vrcsEnabled == true) {
VRCSPostFix(pixelLocation, vrcsData, outputValue);
}
}
清单 1.13|把 VRCS 应用于 Compute Workload。
SRI 与 VRCS Data 每帧只生成一次,随后便可供多个 workload 共用。游戏主要把 VRCS 用在 Deferred Lighting、Deferred G-Buffer Update、Deferred Texturing 和 Deferred Composite。
Deferred Lighting、Deferred G-Buffer Update,以及最初启用 VRCS 的 Deferred Composite 都属于 Fullscreen Compute Pass,仍以 8 x 4 threadgroup 常规覆盖整个屏幕。虽然 Dispatch 了覆盖全屏的 wave,但完全由 Duplicate Pixel 组成的 wave 可以很早退出。清单 1.13 的主要作用就是这一点。
**译注:**原书此处句首疑似缺失;本文根据前后文恢复主语。
Deferred Composite 负责 fog、probe、reflection 等场景合成,其中使用了 Temporal Blue Noise Dithering。实践中,VRCS 很难稳定处理这种 noisy signal,常常会出问题,因此团队在发布后的 Patch 中关闭了该 Pass 的 VRCS。Noisy Image 与 VRCS 该如何结合,仍需进一步研究。
Deferred Texturing 可以更进一步:生成 Pixel Command 时完全不为 Duplicate Pixel 添加命令(见图 1.8)。Command Count 与 Generation Shader 仍需为 100% Primary Pixel 的最坏情况 Dispatch 足够多的 wave,但最终 Pixel Command List 只包含 Primary Pixel,因此不会出现运行 Duplicate Pixel 的 wave。
此时只需在 Command Count 和 Generation Shader 中调用 VRCSPrefix;Deferred Texturing Shader 本身只处理 Primary Pixel,再调用 VRCSPostFix 填充 Duplicate Pixel。整个屏幕最多只留下一个部分填充 wave,而不是每个 tile 一个,因此比 Screen-Space Dispatch 更高效。未来团队还希望把 Pixel Command 用于其他 Deferred Pass。

1.10.4 Deblocking
VRCS 和其他 Variable Rate Shading 技术一样,把结果复制到 Coarse Shading Quad 后可能产生块状 artifact。为此可以增加一个 Deblocking Pass。idTech 8 的 VRCS Deblocking 大致基于 [Drobot 20]:为 VRCS Coordinate Buffer 中的每个 entry 启动一个线程,再执行 Pixel Neighborhood Filtering。
- 2 x 1 Rate:水平方向平均两个像素。
- 1 x 2 Rate:垂直方向平均两个像素。
- 2 x 2 Rate:对角平均四个像素。

Deblocking 的更多细节见 [Hammer and Fuller 25]。
1.10.5 VRCS 性能结果
VRCS 平均可以节省约 1.0–2.0 ms GPU 时间,不过实际收益很看平台、Internal Render Resolution 和当前画面内容。为了保证首发时达到目标帧率,团队最初把实现与优化重点放在主机上;目标分辨率较高时收益最好,例如 Xbox Series X 的 1440p。
游戏发布后,PC 版也加入了 VRCS 开关,但开启后并不总会更快。使用 FSR、DLSS 或 XeSS 等第三方 upscaler 时,VRCS 面对的 Internal Resolution 较低,收益也会随之递减。分辨率低到一定程度后,可省掉的内部像素不够多,VRCS 自身的开销甚至可能反超收益。关闭 Async Compute,或者 PC GPU 的 Async Compute 效率不高时,这个问题会更加明显。所以下文主要看 Xbox Series X 的数据。

| Resolution Scale | 状态 | Frame | Deferred Texturing | Deferred G-Buffer | Deferred Lighting |
|---|---|---|---|---|---|
| 100% | VRCS On(ms) | 16.88 | 1.561 | 0.93 | 1.54 |
| 100% | VRCS Off(ms) | 18.63 | 2.536 | 1.193 | 2.377 |
| 100% | 节省(ms) | 1.75 | 0.975 | 0.263 | 0.837 |
| 100% | 耗时降低 | 9.39% | 38.45% | 22.05% | 35.21% |
| 85% | VRCS On(ms) | 13.64 | 1.24 | 0.63 | 1.245 |
| 85% | VRCS Off(ms) | 15.74 | 2.196 | 1.05 | 2.04 |
| 85% | 节省(ms) | 2.1 | 0.956 | 0.42 | 0.795 |
| 85% | 耗时降低 | 13.34% | 43.53% | 40.00% | 38.97% |
| 50% | VRCS On(ms) | 8.9 | 0.671 | 0.461 | 0.729 |
| 50% | VRCS Off(ms) | 10.13 | 1.327 | 0.641 | 1.205 |
| 50% | 节省(ms) | 1.23 | 0.656 | 0.18 | 0.476 |
| 50% | 耗时降低 | 12.14% | 49.43% | 28.08% | 39.50% |
表 1.5|最佳场景下的 VRCS 性能结果。
图 1.28 的画面对比度和空间频率都较低,因此可以使用更多 Half-Rate 与 Quarter-Rate Tile。最高收益并不一定出现在最低或最高分辨率;对这个场景而言,约 85% Dynamic Resolution Scale 才是 sweet spot。

| Resolution Scale | 状态 | Frame | Deferred Texturing | Deferred G-Buffer | Deferred Lighting |
|---|---|---|---|---|---|
| 100% | VRCS On(ms) | 18.7 | 1.984 | 1.04 | 1.61 |
| 100% | VRCS Off(ms) | 20.02 | 2.61 | 1.263 | 2.366 |
| 100% | 节省(ms) | 1.32 | 0.626 | 0.223 | 0.756 |
| 100% | 耗时降低 | 6.59% | 23.98% | 17.66% | 31.95% |
| 85% | VRCS On(ms) | 15.9 | 1.473 | 0.897 | 1.352 |
| 85% | VRCS Off(ms) | 17.7 | 2.287 | 1.118 | 2.053 |
| 85% | 节省(ms) | 1.8 | 0.814 | 0.221 | 0.701 |
| 85% | 耗时降低 | 10.17% | 35.59% | 19.77% | 34.15% |
| 50% | VRCS On(ms) | 10.31 | 0.718 | 0.539 | 0.801 |
| 50% | VRCS Off(ms) | 11.27 | 1.334 | 0.68 | 1.226 |
| 50% | 节省(ms) | 0.96 | 0.616 | 0.141 | 0.425 |
| 50% | 耗时降低 | 8.52% | 46.18% | 20.74% | 34.67% |
表 1.6|平均场景下的 VRCS 性能结果。
图 1.29 的画面细节更多、空间频率更高,因此整体 Shading Rate 也更高,节省幅度低于最佳场景。
1.11 详细结果
优化后的 idTech 8 Opaque Geometry Rendering Pipeline,选在游戏第四个任务 Siege 的一个场景中做评估。这里视距很远,还有多个 NPC、密集植被和大量小三角形,是一个典型的 GPU 性能热点。

表 1.7 的数据来自 Xbox Series X,Internal Render Resolution 为 2560 x 1440,关闭 Dynamic Resolution Scaling。为分别展示各阶段成本,大多数数值关闭了 Async Compute;“假定 Async 免费”的 Overall 值则根据整体 Frame Time 差异推导,正确计入 Async Compute 的重叠。
| 阶段(ms) | idTech 7 | idTech 8 | idTech 8,无 VRCS | idTech 8,无 VRCS / Tile Classification |
|---|---|---|---|---|
| Pixel Commands 与 Dispatch Arguments(Async) | — | 0.35 | 0.35 | 0.35 |
| Deferred Attribute Interpolation(Async) | — | 0.43 | 0.43 | 0.43 |
| VRCS Shading Rate(Async) | — | 0.41 | — | — |
| Deferred Texturing(Compute Dispatch) | — | 3.79 | 4.20 | 4.20 |
| Tile Classification | — | 0.01 | 0.01 | — |
| G-Buffer Update(Tiled) | — | 0.99 | 1.10 | 1.30 |
| Deferred Lighting(Tiled) | — | 1.32 | 1.70 | 1.80 |
| Forward+ Opaque Drawing | 9.43 | — | — | — |
| Overall,关闭 Async Compute | 9.43 | 7.30 | 7.79 | 8.08 |
| Overall,假定 Async“免费” | 9.43 | 7.0 | 7.5 | 7.9 |
表 1.7|Xbox Series X 在 Siege 地图中的性能,单位为毫秒。
在全部优化与 Async Compute 同时启用的最佳情况下,Opaque Geometry Rendering 从 idTech 7 的 9.43 ms 降到 idTech 8 的 7.0 ms,耗时最多降低约 25%。即使不使用 Async Compute,耗时降幅仍有 22%。
Triangle Visibility Buffer 与 Deferred Texturing / Lighting 是最大贡献者,相比 idTech 7 最多节省 1.53 ms;VRCS 在该视角另外节省约 0.5 ms;Tile Classification 通过限制修改 G-Buffer 与计算 Lighting 的像素数量,再节省约 0.4 ms。VRCS 与 Triangle Visibility Buffer Pipeline 的具体收益都会随可见内容变化。
Dynamic Resolution Scaling 在游戏中覆盖总像素数的 100%–50%。早期目标之一,是在 Resolution Scale 降低时获得更接近像素数量变化的性能收益,使游戏平均维持更高分辨率。
| Pipeline | 100% Resolution | 50% Resolution |
|---|---|---|
| idTech 7 Forward+ | 2560 x 1440:9.43 ms | 1811 x 1019:6.4 ms(67%) |
| idTech 8 Deferred | 2560 x 1440:7.5 ms | 1911 x 1019:3.9 ms(52%) |
表 1.8|Dynamic Resolution Scaling 的有效性;括号为相对 full-scale 的耗时比例。
**译注:**原书表 1.8 将 idTech 8 的 50% Resolution 记为
1911 x 1019,本文按原书保留;同章前文表 1.1 的对应项为1811 x 1019,两处存在不一致。
idTech 8 的 Opaque Draw Cost 几乎随像素数量线性变化;idTech 7 的 Forward+ 在降至一半像素后仍保留 67% 的耗时,主要原因是 Rasterization Overhead 与更低的 quad utilization。
表 1.9 汇总了这些优化的显存成本。3840 x 2160 配置在表中的合计为 286.03 MiB;原文正文同时称,4K 下显存成本最高可达约 330 MiB。原文没有进一步解释两者的口径差异,这里按原样保留。由于这些优化直到项目后期才完成,团队还没有系统研究如何复用此前 Pass 的 Buffer / Image,或进一步压缩资源。
| Technique(MiB) | 3840 x 2160 | 2560 x 1440 |
|---|---|---|
| Visibility Buffer + Compute Dispatch | ||
| Visibility Buffer Image | 63.3 | 28.1 |
| Compute Dispatch Buffer(含 Indirect Argument) | 12 | 5 |
| Pixel Command + Counter Buffer(与 Strands Hair alias) | 46 | 20 |
| Strands Hair Simulation Buffer(aliasing deduction) | -32 | -20 |
| Vertex Barycentric Derivatives | 63.3 | 28.1 |
| Vertex Barycentrics | 31.6 | 14 |
| Vertex Tangent Frames | 63.3 | 28.1 |
| Visibility Buffer + Compute Dispatch 小计 | 247.5 | 103.3 |
| VRCS | ||
| VRS Shading Rate Image | 0.03 | 0.01 |
| VRCS Shading Rate Image | 1.9 | 0.87 |
| VRS Luma Image | 7.9 | 3.5 |
| VRS Luma Image,Post-Blended | 7.9 | 3.5 |
| VRCS Coordinate Buffer | 13.8 | 6.2 |
| VRCS 小计 | 31.53 | 14.08 |
| Tile Classification Buffer(含 Indirect Argument) | 7 | 3.5 |
| 总计 | 286.03 | 120.88 |
表 1.9|idTech 8 全部优化技术的显存成本,单位为 MiB。
目前较明确的复用是让 Pixel Command Buffer 与帧后段才使用的 Strands Hair Simulation Buffer alias 到同一段 GPU Memory。团队希望未来继续寻找类似机会。
1.12 结论
Triangle Visibility Buffer 与 Deferred Texturing 提高了 Shading Efficiency,尤其适合大量小三角形导致传统 Forward+ Geometry Pass 的 quad utilization 成为瓶颈的场景。
把 Deferred Lighting 拆成 G-Buffer Modification 与 Lighting 两个专用 Compute Pass 后,Shader 更小、目标更明确,Register Usage 更低,也更容易优化。Tile Classification 进一步把昂贵 Shader Variant 限制在真正需要它的区域。VRCS 则允许邻近像素共享结果,在适合的屏幕区域跳过昂贵 Lighting Calculation。
几项优化叠加后,Opaque Geometry Rendering 与 Lighting 的 GPU 时间降低了约 25%,Pipeline 也更有能力承载包含更多小三角形的高质量 model。
这些收益并非没有代价。当前实现占用的显存不少,尤其在较高的 Internal Render Resolution 下,未必适合 GPU Memory Budget 已经吃紧的项目。Variable Rate Shading 也会带来 artifact 和 temporal instability;为一类场景做的修复,还可能影响另一类场景的画质。
此外,Forward+ Geometry Pass 与 Compute-Based Deferred Texturing / Lighting 必须并存,mesh 甚至可能逐帧切换路径。代码库因此更复杂,后续修改 Rendering Pipeline 时也更难判断影响范围。即便要付出这些代价,团队仍然认为这次改造是成功的:游戏最终得以在当代主机与 PC 上以至少 60 Hz 发布。
1.13 后续工作
Deferred Texturing 已经完全基于 Compute,却仍运行在 Graphics Pipeline。将其移到 Async Compute 很有吸引力,但前提是 Graphics Queue 上有可以与它重叠的工作,同时两条 Queue 不会争抢 GPU Resource。
构建 Pixel Command List 和执行 Attribute Interpolation 时,Async Compute 的收益也低于预期,原因是同一时段还有 SSDO 等 Compute-Intensive Work 被安排在 Graphics Queue。后续需要重新组织这些 GPU Job,让两条 Queue 形成更有效的重叠。
显存占用同样需要继续优化。Pixel Command Buffer 已经能够与帧后段使用的 Strands Hair Simulation Buffer alias,但表 1.9 中其他 GPU Resource 尚未充分探索类似机会。
Visibility Buffer Generation 目前仍绑定在 Geometry Depth Pass 上。此时全部 Triangle Data 已经以大型 GPU Buffer 的形式常驻,Triangle Index 也经过了 GPU Triangle Culling。以 Software Rasterization 替代或补充 Hardware Depth Prepass,或许能进一步提高未来可支持的 Triangle Density。
Deferred Texturing 与 Deferred Lighting / G-Buffer Update 的 Dispatch 策略差异很大。前者使用 Fine-Grained Dispatch,可以把专用 Shader 精确调度到单个像素;后者则采用少量 Ubershader Variant 和较粗的 Tile Dispatch。下一步可以研究让 Lighting 与 G-Buffer Update 也采用类似的 Per-Pixel Dispatch。如果分组的 Material Granularity 能接近 Deferred Texturing,Lighting Shader 理论上也可以重新利用一部分 uniformity assumption。
VRCS 将 Shading Sample 与 Internal Render Resolution 解耦,整体取舍令人满意,但它对画质的影响会随场景和 Lighting Environment 大幅变化。针对一个场景调整 VRCS Parameter 与 Shading-Rate Threshold,很容易让另一个场景变差。更自动化的测试系统可以覆盖更广泛的用例,并尽早暴露 Visual Regression,这也是后续值得投入的方向。
性能上,如果把 Pixel Command List 扩展到其他 Deferred Pass,并设法省掉 Hole Filling 带来的 Memory Bandwidth Cost,VRCS 应该还有继续优化的空间。
致谢
作者感谢 id Software 工程团队持续维护 idTech 的各个部分,并在其他功能与问题上分担工作,让他们得以在发售前不到一年尝试大规模改造核心渲染方式。特别感谢 Allen Bogue、Billy Khan、Bogdan Coroi、Carson Fee、Ian Malerich、John Roberts、Johan Donderwinkel、Mel-Frederic Fidorra、Oliver Fallows、Peeter Parna 博士、Regan Carver、Seth Hawkins、Stefan Pientka、Thorsten Lange、Tiago Sousa 与 Yixin Wang;也感谢 id Software 管理层 Marty Stratton 和 Hugo Martin 对这项高风险优化工作的信任。
作者还感谢 Microsoft ATG 团队提供早期反馈,尤其感谢 Martin Fuller 对 VRCS 的早期工作与大量直接贡献。他在工作室参与开发的时间,为整个团队带来了宝贵经验。
参考文献
[Burns and Hunt 13] Christopher A. Burns and Warren A. Hunt. “The Visibility Buffer: A Cache-Friendly Approach to Deferred Shading.” Journal of Computer Graphics Techniques (JCGT) 2:2 (2013), 55–69. https://jcgt.org/published/0002/02/04/
[Doghramachi and Bucci 17] Hawar Doghramachi and Jean-Normand Bucci. “Deferred: Next-Gen Culling and Rendering for Dawn Engine.” In GPU Zen: Advanced Rendering Techniques, edited by Wolfgang Engel. Black Cat Publishing, 2017.
[Drobot 20] Michal Drobot. “Software-Based Variable Rate Shading in Call of Duty: Modern Warfare.” Presented at SIGGRAPH, 2020. https://research.activision.com/publications/2020-09/software-based-variable-rate-shading-in-call-of-duty–modern-war
[Drobot 21] Michal Drobot. “Geometry Rendering Pipeline Architecture at Activision.” Rendering Engine Architecture course, SIGGRAPH, 2021. https://enginearchitecture.org/downloads/reac2021/geometry_pipeline_rendering_architecture.pptx
[Fuller et al. 22] Martin Fuller, Christopher Wallis, and Philip Hammer. “Variable Rate Shading Update Xbox Series Consoles.” Microsoft Game Dev channel, YouTube, 2022. https://www.youtube.com/watch?v=pPyN9r5QNbs
[Fuller 19] Martin Fuller. “Variable Rate Shading, A Deep Dive | Game Developers Conference 2019.” Microsoft Game Dev channel, YouTube, 2019. https://www.youtube.com/watch?v=2vKnKbaOwxk
[Fuller 22] Martin Fuller. “Variable Rate Compute Shaders—Halving Deferred Lighting Time.” Microsoft Game Dev channel, YouTube, 2022. https://www.youtube.com/watch?v=Sswuj7BFjGo
[Geffroy et al. 20] Jean Geffroy, Axel Gneiting, and Yixin Wang. “Rendering the Hellscape of DOOM Eternal.” Presented at SIGGRAPH, 2020. https://advances.realtimerendering.com/s2020/RenderingDoomEternal.pdf
[Hable 21] John Hable. “Visibility Buffer Rendering with Material Graphs.” Filmic Worlds Blog, July 5, 2021. http://filmicworlds.com/blog/visibility-buffer-rendering-with-material-graphs/
[Hammer and Fuller 25] Philip Hammer and Martin Fuller. “Variable-Rate Compute Shaders in DOOM: The Dark Ages.” Presented at Graphics Programming Conference, 2025. https://graphicsprogrammingconference.com/archive/2025/#variable-rate-compute-shaders-in-doom-the-dark-ages
[Hecker 95] Chris Hecker. “Perspective Texture Mapping.” Game Developer Magazine, April/May 1995, 16–25. https://www.chrishecker.com/images/4/41/Gdmtex1.pdf
[Karis et al. 21] Brian Karis, Rune Stubbe, and Graham Wihlidal. “A Deep Dive into Nanite Virtualized Geometry.” Presented at SIGGRAPH, 2021. https://advances.realtimerendering.com/s2021/Karis_Nanite_SIGGRAPH_Advances_2021_final.pdf
[Kulkarni and Engel 24] Manas Kulkarni and Wolfgang Engel. “Triangle Visibility Buffer 2.0.” In GPU Zen 3: Advanced Rendering Techniques, edited by Wolfgang Engel, 99–109. Black Cat Publishing, 2024.
[Lazarek and Hammer 25] Dominik Lazarek and Philip Hammer. “Visibility Buffer and Deferred Rendering in DOOM: The Dark Ages.” Presented at Graphics Programming Conference, 2025. https://graphicsprogrammingconference.com/archive/2025/#visibility-buffer-and-deferred-rendering-in-doom-the-dark-ages
[McLaren 22] James McLaren. “Adventures with Deferred Texturing in Horizon Forbidden West.” Presented at Game Developers Conference, 2022. https://www.gdcvault.com/play/1027553/Adventures-with-Deferred-Texturing-in
[Microsoft 25] Microsoft. “Variable Rate Shading.” DirectX-Specs, GitHub, 2025. https://microsoft.github.io/DirectX-Specs/d3d/VariableRateShading.html
[Saito and Takahashi 90] Takafumi Saito and Tokiichiro Takahashi. “Comprehensible Rendering of 3-D Shapes.” Computer Graphics 24:4 (1990), 197–206. https://dl.acm.org/doi/pdf/10.1145/97880.97901
[Wihlidal 24] Graham Wihlidal. “Nanite GPU-Driven Materials.” Presented at Game Developers Conference, 2024. https://gdcvault.com/play/1034407/Nanite-GPU-Driven