Implementing Custom Terrain in UE5.8: From Height Data to Rendering and Collision

Overview

This terrain system takes a heightfield as input and builds separate representations for visual geometry, surface materials, and physics collision in UE5.8. The article follows the data through height encoding, CDLOD node selection and geomorphing, weight clipmaps, Dynamic Texture Arrays, and RVT, then covers shared collision edges, triangle holes, and project integration. Implementation versions and acceptance records appear in the appendices.

The pseudocode summarizes the current implementation's main control flow. It uses descriptive names in place of some UE types and omits logging, resource creation, and repeated validation. It is a guide to the source, not compilable code. The relevant sections describe configuration branches and asynchronous completion conditions.

Terrain running in UE5.8
Figure 1 | Terrain preview in UE5.8. Roads, slopes, and surface layers use Basic template preview lighting.

1. Terrain Representations and Data Layout

A large terrain contains many samples, but only a subset needs drawing each frame. Nearby areas require a dense mesh; distant areas can use wider sample spacing. Surface textures have their own detail requirements: rock, grass, and roads still need appropriate sampling and blending as the mesh becomes coarser. Character collision also needs a stable physical surface that does not change with visual LOD.

These requirements lead to three terrain representations. The heightfield defines the shape, and CDLOD selects drawing resolution by distance. Weights and texture arrays describe surface layers, while RVT caches attributes composited at world positions. Physics builds separate geometry from heights and triangle visibility. The three paths share coordinates and valid bounds but use different partitions and update rules. Figure 2 shows their data relationships.

Terrain data flow
Figure 2 | Terrain data flow. Heights feed CDLOD, weights and layer textures feed material composition, and triangle visibility constrains both rendering and collision.
The hill view running in UE5.8
Figure 3 | Roads and slopes in the hill terrain, shown with Basic template preview lighting.

1.1 Runtime Payload and Triangle Visibility

The terrain payload has separate Height, Weight, Hole, MaterialCatalog, and Collision sections. Its header stores block counts, samples per block, sample spacing, height scale, and section offsets. The loader parses the file into a shared payload structure, which the rendering, weight, and physics paths then consume.

Data states distinguish Unknown, Valid, Empty, OutOfRange, Failed, and Pending. An empty block must remain distinct from a failed load; otherwise, zero heights, default colors, or old pages can be mistaken for valid results. Callers inspect LoadError, LastError, and diagnostic state rather than checking only whether a resource object exists.

Triangle visibility is stored in a matching .visibility sidecar deployed with the main payload. It contains triangle visibility for 110 source-node groups and the valid terrain bounds, shared by rendering and collision. Triangle data is separate because the main payload's whole-cell hole format cannot represent a surface that retains only half a cell.

1.2 Data Blocks, Render Patches, and Collision Blocks

Partition Current Configuration Purpose
Payload data block 64×64 blocks, each containing 128×128 samples File organization and source-data addressing
CDLOD render patch Each node reuses a regular 32×32-cell mesh Changes world-space coverage with LOD
Collision block 129×129 vertices, 128×128 cells Stable physical surface with shared edges
Weight cache page Defined by clipmap page size and mip Updates the physical texture cache

Data blocks define file sections and addressing; render patches define the grid covered by one draw; collision blocks define the extent of physical geometry. They can share height samples without using the same dimensions. The component default for PatchSize is 16, while this scene uses 32. Payload blocks still contain 128 samples per axis.

Three spatial partitions
Figure 4 | Three spatial partitions. Payload blocks organize samples, CDLOD patches reuse a mesh over different extents, and collision blocks read the final row and column from neighbors. Grid counts in the diagram are schematic.

1.3 Coordinate Conversion and Block Addressing

The data paths use world coordinates, terrain-local coordinates, full-map sample coordinates, and block-local coordinates. World coordinates locate cameras and characters; local coordinates are relative to the terrain origin; full-map sample coordinates convert distances into sample indices; block-local coordinates address the corresponding payload array. Applying modulo too early or omitting the origin can produce valid memory accesses to the wrong terrain region.

For a global integer index GX, sample spacing Spacing, and CellX horizontal samples per block, addressing is:

PatchX = GX / CellX
LocalX = GX % CellX
SourceIndex = LocalY × CellX + LocalX

These equations apply to the current nonnegative source indices. World positions are converted to page coordinates with floor. For negative coordinates, truncation toward zero and rounding down can produce different page indices. Section 5 covers negative-coordinate clipmap wrapping. Source access still requires state and bounds checks; ring-cache modulo does not change the source data's valid extent.

For example, each data block currently stores 128 samples per axis. Full-map sample GX=128 is LocalX=0 in the next block, not a duplicated sample at the end of the preceding block. Collision construction uses this addressing to read the additional shared edge from neighboring blocks. Both render-atlas assembly and collision edge completion follow this storage layout.

Coordinates and addressing
Figure 5 | Coordinates and addressing. GX=128 selects the first sample of the next block. Collision shared edges use full-map indices to read neighboring blocks.

1.4 File Versions and Error States

The loader reads the file into memory, then uses a bounds-checked Reader to parse the magic value, version, header, and data sections. The Reader's Seek(), Read(), and ReadBytes() check file bounds and record the position and requested length on failure. Section offsets allow independent addressing, but the current implementation still loads the entire file immediately. Asynchronous section reads are not implemented.

File version 2 adds HeightScaleWU; version 1 uses the default value of 100. This field participates in height decoding, so version differences directly affect vertex heights. A nonpositive height scale falls back to the default, meaning the loader does not strictly reject every invalid field. A bad magic value, unsupported version, or out-of-bounds read sets LoadError and stops the affected path.

Valid empty regions, unfinished requests, and failed pages may use the same placeholder pixels. Block state distinguishes their runtime meaning. Empty means processing finished with no content; Pending means the result has not arrived; Failed means the current attempt cannot provide data. Callers use these states to wait, display fallback data, or report an error.

2. Preparing the Heightfield for the GPU

The vertex shader samples heights by terrain coordinates, so the input is first assembled into a 2D atlas. The terrain component reads the main payload, and BuildHeightNormalTexture() stitches its blocks contiguously. The 64×128 samples along each axis produce an 8192×8192 atlas containing 67,108,864 heights. This path loads and retains the full atlas; it does not request height pages by camera position.

The stitching stride is CellCount, which is 128 in this configuration. Adjacent payload blocks do not duplicate boundary samples. Using 127 instead would introduce a one-texel displacement at every block boundary, causing sampling misalignment where the terrain varies in height.

The atlas uses PF_B8G8R8A8, with sRGB disabled and ordinary texture streaming disabled. Logical R/G channels store the high and low height bytes, and B/A store normal XY. Height encoding and decoding are:

H16 = clamp(round(height_m × 128 + 32768), 0, 65535)
height_m = (H16 - 32768) / 128
height_wu = height_m × HeightScaleWU

Here, height_m is the height section's value in meters. HeightScaleWU converts it to the world units used by vertices. This dataset uses 125 to preserve the source terrain's vertical scale. The collision section's RawHeights already includes this conversion, so physics construction does not multiply by 125 again. Both paths must produce the same height after conversion.

Height padding is -256, which becomes -32000 world units after multiplication by 125. This value still generates geometry. Holes and terrain boundaries require visibility and bounds clipping; lowering vertices alone leaves triangles and collision surfaces below the terrain.

Normals are generated from central height differences, with both height and horizontal spacing expressed in meters. The current implementation preserves the source normal convention and does not recompute normals after scaling vertex heights by ×1.25. This retains the reference sampling semantics. Final lighting also depends on material normals, roughness, and scene settings.

The four-byte 8192² atlas itself occupies about 256 MiB. This estimate covers atlas pixels only, excluding CPU height arrays, payloads, temporary buffers, and other textures. CDLOD reduces the geometry drawn, but the full-map atlas's VRAM use still grows with data resolution.

Heights on the GPU
Figure 6 | Heights on the GPU. R/G pack height and B/A store normal XY. Rendering multiplies decoded height by HeightScaleWU; the collision section has already converted units.

2.1 Height Encoding and Quantization Range

After a 16-bit height is packed into two eight-bit channels, the shader must restore the contribution of each byte. If sampled R and G are normalized to [0,1], multiplying each by 255 and substituting into the decoding equation gives:

height_m = R × 510 + G × 1.9921875 - 256

The Vertex Factory uses this form. Multiplying the high byte by 256 and dividing by 128 gives the coefficient 510; the low-byte coefficient is 255/128. The channels therefore hold different bit ranges of one linear number. sRGB conversion would change that relationship and must remain disabled.

The encoding step is 1/128 m. With this dataset's height scale, that corresponds to a world-height interval of about 0.9765625 cm. This is the rendering height encoding's resolution. Collision uses separate quantization, and ray references also involve triangle interpolation, so collision error must be measured at runtime. The later error figures are measured results.

Heights outside the encoding range are clamped. Source extraction, the payload header, and bounding boxes must agree on the range. Enlarging a bounding box cannot recover height lost during encoding. Conversely, correct heights with an undersized bounding box can cause the engine to cull a visible mountain prematurely.

2.2 Normal Generation and Resource Costs

Normal XY in the height atlas is computed using central differences. Interior samples compare left/right and upper/lower neighbors. Boundary sample coordinates are clamped to the atlas, and the denominator uses the actual distance between the samples. Normalized XY is written to B/A; the shader maps those channels back to [-1,1] and reconstructs a nonnegative Z component.

Terrain normals and material detail normals describe different spatial frequencies. The former represent heightfield slope; the latter come from grass, rock, and other layer textures. Detail normals can still vary the pixel-stage surface when geometry becomes coarser, but they cannot change the silhouette or replace collision geometry. Distant-view differences require separate checks of the mesh, normals, and surface layers.

Full-map loading allocates several data copies in sequence. The payload retains block data, the component builds a floating-point height grid, and a four-byte pixel array is generated and uploaded. These objects may coexist during initialization, so peak memory and long-term residency must be measured separately. The Clipmap Actor also invokes the payload loader independently. There is no shared payload service, so the full memory budget must account for duplicate loads and data retained by each module.

Full residency permits direct access to any height and helps validate coordinates and sampling against the source. A future streaming height cache must preserve the same decoding and coordinate contract and provide adequate height ranges for node bounds. The quadtree still depends on those ranges; a smaller height texture does not remove that dependency.

WorldNormal visualization in UE5.8
Figure 7 | WorldNormal visualization. Colors show the material’s world-space normals, including slope direction and detail normals. This is a display image, not raw GBuffer data.

3. CDLOD: CPU Node Selection and GPU Geometry

Continuous Distance-Dependent Level of Detail (CDLOD) selects terrain regions through a quadtree and reuses a regular grid to draw the heightfield at different resolutions. The author's CDLOD project distinguishes fully resident and streaming versions. This project's height path currently uses full residency.

3.1 Quadtree Selection and Partial Replacement

The quadtree is built recursively. Each node stores its center, size, LOD, minimum and maximum height, children, and visibility offset. Parents aggregate their children's height ranges to construct bounding boxes that contain the actual terrain.

TraverseNode() computes the squared XY distance from the viewpoint to a node's bounding box and compares it with per-level thresholds. Nodes outside the current level's range are covered by a coarser level; entering a finer level's range triggers recursion into the four children. CPU selection uses planar distance and does not include camera height.

If only some children are selected, the parent still submits one draw and uses quarter-cull bits to suppress the regions covered by those children. When all four children cover the region, the parent is omitted. This preserves full coverage without drawing the same area twice.

Distance thresholds combine LODDistanceFactor, SampleSpacing, LODDistanceAdditions, and a step that doubles at each level. The acceptance configuration's additional distances are [6400, 34000, 37200, 74400]. A single doubling formula cannot replace these source-defined offsets.

Selection uses the nearest distance to a 2D AABB. For each axis, subtract the node's half-extent from the absolute camera-to-center distance and clamp negative results to zero:

dx = max(abs(view_x - center_x) - extent_x, 0)
dy = max(abs(view_y - center_y) - extent_y, 0)
DistanceSquared = dx² + dy²

The distance is zero inside the node's horizontal bounds. Outside, it measures the nearest edge or corner. This detects node edges already close to the camera and avoids the bias of measuring only center distance. Squared distance avoids square roots during traversal.

Traversal returns whether the current subregion is covered by this level or its descendants. If only two children enter the finer LOD range, the parent sets the corresponding two quarter bits to cull them and draws the other two quarters with its own grid. These four bits carry the coverage relationship into the vertex path.

CPU selection and GPU morphing use the same distance parameters but operate on different objects. The CPU decides coverage at node scale; the GPU computes grid-point convergence at vertex scale. Selection determines the available topology, while vertex evaluation provides the continuous transition. Neither operation replaces the other.

The selection process is shown below. selected stores nodes and their quarter-cull masks. Returning true means the region is covered by the node or its descendants.

function SelectNode(node, lod, viewXY, selected):
    d2 = DistanceSquaredToAABBXY(viewXY, node.bounds)
    if d2 >= LODDistance[lod]²:
        return false

    if lod == 0 or d2 >= LODDistance[lod - 1]²:
        selected.append(node, lod, quarterCullMask = 0)
        return true

    quarterCullMask = 0
    for q in 0..3:
        if SelectNode(node.children[q], lod - 1, viewXY, selected):
            quarterCullMask |= 1 << q

    if quarterCullMask != 0b1111:
        selected.append(node, lod, quarterCullMask)
    return true

The parent remains selected until all four quarters are covered by children. Culling the whole parent whenever any child is selected would leave gaps during partial refinement.

CDLOD selection and geomorphing
Figure 8 | CDLOD selection and geomorphing. The parent retains only quarters not covered by children. The GPU gradually moves odd grid points toward neighboring coarse points while even points remain fixed. Regions and levels are schematic.

3.2 Instance Data and the Vertex Factory

Each selected node supplies a position offset, grid spacing, LOD, texture coordinates, quarter-cull bits, and triangle visibility offset. SceneProxy groups this data into batches and submits FMeshBatch. The custom Vertex Factory unpacks instance data in the vertex shader.

The regular grid defines horizontal sample positions. The GPU positions vertices using each node's grid spacing, samples the height atlas, and writes decoded height into Z. Nodes reuse the same topology without retaining a highest-resolution vertex mesh for the entire map.

This Vertex Factory retains the implementation's translation and scale conventions. Acceptance tests have not covered arbitrary Actor rotation, nonuniform transforms, or multiple-camera combinations. It should not be integrated as a general terrain component supporting arbitrary transforms without further validation.

The implementation has ordinary and hole-aware grid topologies. The ordinary topology duplicates vertices along the middle row and column so each quarter can carry its own culling information. With 32 cells per axis, the logical grid is still 33×33, but the ordinary vertex buffer contains 34×34 vertices after duplication. Duplicated vertices occupy the same positions and separate attribute ownership without adding terrain area.

The hole-aware topology gives each cell's two triangles separate vertices, totaling 32×32×6. Two triangles sharing a grid point may have different visibility. If they shared a vertex with one PrimId, the vertex stage could not degenerate either triangle independently. Separate vertices assign a consistent triangle ID to all three vertices of each triangle.

SceneProxy separates ordinary and hole-aware instances, selects the matching Vertex Factory variant, and batches them according to Uniform Buffer instance capacity. Each instance uses two float4 rows: one stores node position and UV offset; the other stores packed flags, inverse grid spacing, texel step, and visibility offset. Some integer flags are bit-reinterpreted as float and restored with asuint. This is not an approximate numeric integer-to-float conversion.

UV offsets must be normalized by atlas size. The node's minimum corner is divided by horizontal sample spacing to obtain texel coordinates, then multiplied by inverse atlas dimensions. Adding texel counts directly to normalized UVs causes the Clamp sampler to restrict most requests to the atlas boundary. Flat test data can hide this because different positions return the same height; terrain variation makes shape and LOD-boundary misalignment more apparent.

3.3 Geomorphing and Height Sampling

Adjacent LODs use different grid spacing. Switching meshes directly makes fine-grid vertices jump. Geomorphing gradually moves them toward coarse-grid points over a distance interval. The shader uses each vertex's distance to the reference viewpoint:

MorphK = clamp((Distance - MorphStart) / MorphDistance, 0.0, 1.0);
MorphVector = -2.0 * MorphK * frac(InputPosition * 0.5);
MorphedInputPosition = InputPosition + MorphVector;

MorphStart is the transition's start distance, MorphDistance is its width, and MorphK is the convergence factor. Even grid points have a zero frac term and remain fixed. Odd points move by at most one current grid interval. At MorphK equal to 1, fine-grid points align with the coarse grid.

Height sampling must follow the same transition. The shader explicitly samples the four corners of the current LOD cell from the full-map atlas and interpolates within that cell. If position has converged to the coarse grid but height still comes from fine detail at a fractional world position, the transition can depart from the actual coarse surface.

Packed heights do not rely on ordinary texture mips for automatic averaging. Explicit LOD-grid sampling preserves the source sample relationship.

Actual LOD selection in UE5.8
Figure 9 | Actual CDLOD selection. Grayscale encodes LOD levels; boundaries follow the coverage of selected nodes.

3.4 Grid and Height Examples in the Transition Region

The CDLOD diagram shows convergence along one axis at MorphK=0,0.5,1. Positions 1 and 3 eventually coincide with 0 and 2 respectively. Some fine-grid triangles gradually degenerate, leaving a coarse grid with twice the spacing.

Continuous positions do not guarantee correct heights. A transitioning vertex lies at a fractional grid position, and its height must come from adjacent samples in the current LOD grid. The implementation floors the morphed grid coordinate and reads four corners; horizontal and vertical frac terms determine bilinear weights. All four reads use height-atlas mip0, but corner spacing increases with LOD, so coarser levels read a sparser set of source samples.

This differs from sampling an ordinary downsampled height mip. Ordinary mips often average a region's heights, whereas this CDLOD path retains values at the corresponding coarse grid points. Both can look smoother but need not produce the same coarse surface. Reproducing the source algorithm requires preserving the coarse-grid sampling relationship.

LOD parameters also control transition width. Distance thresholds determine when coverage passes to finer or coarser levels, while MorphDistanceFactor sets each level's transition band. A wider band spreads geometric change over a larger distance without adding source detail. If distance offsets and morph intervals disagree, coverage selection must be corrected; increasing the smoothing factor cannot remove that error.

The pseudocode below shows grid convergence followed by height sampling. grid is the LOD-grid coordinate clipped to the valid terrain extent. AtlasUV() includes the node UV offset, LOD texel step, and half-texel positioning.

function EvaluateTerrainVertex(inputGrid, node, referenceView):
    originalXY = node.origin + inputGrid * node.vertexSpacing
    distance = DistanceXY(ToWorld(originalXY), referenceView)
    k = Saturate((distance - MorphStart[node.lod]) / MorphWidth[node.lod])
    morphedGrid = inputGrid - 2 * k * Frac(inputGrid * 0.5)

    xy = ClampToTerrain(node.origin + morphedGrid * node.vertexSpacing)
    grid = (xy - node.origin) / node.vertexSpacing
    base = Floor(grid)
    f = Frac(grid)

    h00 = SampleHeightNormal(AtlasUV(node, base + (0, 0)), mip = 0)
    h10 = SampleHeightNormal(AtlasUV(node, base + (1, 0)), mip = 0)
    h01 = SampleHeightNormal(AtlasUV(node, base + (0, 1)), mip = 0)
    h11 = SampleHeightNormal(AtlasUV(node, base + (1, 1)), mip = 0)
    packed = Lerp(Lerp(h00, h10, f.x), Lerp(h01, h11, f.x), f.y)
    z = DecodeHeight(packed.rg) * HeightScaleWU
    return Position(xy, z)

This pseudocode assumes valid node data and omits triangle degeneration and missing-page diagnostics. The source computes displacement before and after clamp, then subtracts it from the grid coordinate. The pseudocode combines those steps by deriving grid coordinates from clipped XY, keeping samples aligned with positions.

4. Triangle Visibility and Terrain Coverage

Each grid cell contains two triangles. Per-triangle visibility allows only half a cell to render. Grid vertices carry a Primitive ID, and the Vertex Factory looks up its bit using the node's visibility offset. Invisible triangles are degenerated and contribute no effective surface area.

Source visibility uses X-major ordering. For a 32×32-cell LOD0 node:

PrimId = 2 × (x × 32 + y) + triangle
triangle ∈ {0, 1}

The height array uses row-major ordering, y × width + x, so the indexing conventions differ. Reusing height indices for visibility selects the wrong triangles and can transpose hole locations. Rendering, collision, and acceptance references must each convert to the correct format.

The valid terrain extent is also independent of the payload rectangle. This configuration uses world bounds [44200, 775000]² cm. Rendering constrains node positions to that extent and adjusts height UVs accordingly; collision removes cells outside it. Full-map padding remains in the data but does not create valid surface coverage.

4.1 Bitmasks, Node Offsets, and Index Units

Visibility is stored as bits; one uint32 describes 32 triangles. Shifting the Primitive ID right by five gives the word index, while its low five bits select the bit within that word. Each node stores the beginning of its visibility data, so the shader's complete index adds the node start to the local PrimId.

CPU node offsets are measured in words. Before entering the vertex path they are multiplied by 32 to become bit offsets. Omitting this unit conversion makes subsequent nodes read another node's visibility data.

Sidecar loading also validates node coordinates, LOD, grid span, uniqueness, and data length. Collision retains only fine-grained LOD0 masks. Their two bits describe nearby cells: 00 is fully invisible, 01 or 10 retains half a cell, and 11 retains both triangles. Valid nodes without a separate mask remain fully visible; the absence of a special hole record is not treated as missing terrain.

Triangle visibility lookup
Figure 10 | Triangle visibility lookup. The local PrimId and node bit offset locate visibility. Rendering and collision apply the same coverage constraint through their respective paths.

4.2 Padding, Valid Extent, and Bounding Boxes

The data rectangle contains addressable samples and may extend beyond the terrain visible in the game. The current visible bounds run from 44200 cm to 775000 cm, while the payload retains outer padding. Rendering clamps local vertex positions and converts the displacement back into grid and UV corrections. Height reads must follow the clipped position; otherwise, a boundary vertex can lie inside the valid region but use a height from outside it.

Bounding boxes control renderer primitive culling. The source builds component Bounds from payload dimensions and vertical Extent, then transforms them into world space. If Z bounds retain a fixed value from an early flat-terrain prototype, mountains may extend beyond Bounds and be partially culled at oblique angles. Triangle visibility controls local topology; Bounds controls whether the primitive reaches the draw path. Both must match the actual data.

5. Weight Clipmaps: A Fixed Cache and a Moving Window

Surface shading involves three resource types. Clipmaps manage weight pages within a spatial window, DTA maps layer IDs to texture slices, and RVT caches material attributes composited by world position. The main material still retains full-map weights and source-array branches. The following sections describe both window caching and the resources actually bound to the material.

Surface weights specify the layers and blend ratio at each position. The payload supports two encodings: HD uses three bytes for two 8-bit IDs and a ratio; Normal packs two 5-bit IDs and a 3-bit ratio into two bytes. Ratios can be interpolated, while material IDs identify discrete layers. Linear filtering of IDs as colors can produce an ID for a layer that did not contribute to the region.

A clipmap stores local weights at different mips in a fixed-size physical texture. Successive levels cover larger world-space extents. Camera movement shifts the logical window, while modulo addressing reuses physical slots in a ring.

The clipmap core stores the logical world center CenterPatchId separately from the physical texture center CenterPatchIdInTexture. When the camera crosses a page boundary, center displacement updates the physical center:

physical_center = positive_mod(physical_center + logical_delta, LayerPatchCount)

Movement in a negative direction can produce a negative ordinary remainder, so the implementation uses positive modulo. Moving the logical window does not require copying the entire texture; new logical pages are assigned to reusable slots.

Each mip computes its logical center independently from world position. If the level-zero page step is S0, the step at level m is S0×2^m:

C_m = floor((ViewPosition - EffectiveLeftTop) / (S0 × 2^m))
Delta_m = C_m - PreviousC_m
T_m = positive_mod(PreviousT_m + Delta_m, N)

C_m is the source logical center, T_m is the physical center, and N is the number of physical slots per axis. For the same camera displacement, fine levels typically cross more pages, while a coarse center may remain unchanged. Levels therefore do not need to shift by the same number of slots.

In a one-dimensional example with N=8, a physical center at slot 6 moves to slot 0 when the camera moves two pages in the positive direction. The result 0 is positive_mod(6+2,8): the texture's starting slots are reused while the logical source window keeps advancing. Moving one page in the opposite direction from slot 6 gives slot 5. Positive modulo keeps the result within the physical slot range. The value 8 illustrates wrapping and does not redefine the current scene configuration.

The physical center also has a continuous UV position. While the camera remains in one logical page, the integer center stays fixed but the within-page offset changes sampling. The source divides the camera's residual distance from the current page center by the page's world step, adds it to the physical center's half-page offset, then divides by the slot count. Integer state handles page changes; fractional state provides continuous positioning within a page.

Clipmap wrapping
Figure 11 | Clipmap wrapping. The physical center moves from slot 6 to slot 0. Overlapping parts of the window retain their pages while source page indices continue forward.

5.1 Page Requests and Residency Checks

Request generation visits each mip's active inner region first, then appends the outer retention ring. The Actor processes the resulting request list over multiple frames according to MaxUploadPerFrame. A camera-position update replaces the old list with a new one.

The list covers the required region, and residency checks exclude pages that do not need another upload. ProcessOneRequest() checks a physical slot's ExistPatchs. If the same logical page is already resident, it is skipped; an upload occurs only when the slot content must change.

A missing source page is recorded as Failed, while the physical texture may retain old pixels. Consumers must check the requested page's validity and readiness rather than infer successful residency from those pixels.

Mapping Source Pages to Physical Slots

A request identifies source page (x,y,mip). ComputeTexPatchId() computes its difference from the current logical center, corrects wrapping direction using the source mip's page count, adds the result to the physical center, and takes positive modulo by LayerPatchCount.

Source-map wrapping and physical-cache wrapping use different dimensions. The former uses the source page count at that mip; the latter uses the physical slot count. A source mip may contain only a few pages while the physical window retains a fixed number of slots. Using one divisor for both can yield an in-bounds texture address that refers to the wrong cached page.

The coarsest level also retains the physical position corresponding to its initial upper-left corner and uses it to correct later alignment. This keeps the source origin of coarse fallback coverage stable as the logical center changes. The coarse image's physical reference must not drift arbitrarily. These coarse alignment semantics are retained from the port and differ from ordinary fine-level tracking.

The Update Budget Counts Processed Requests

ProcessRequests() increments its processed count for every request removed, including requests skipped because they are already resident. Thus MaxUploadPerFrame actually limits requests processed per frame, not completed texture uploads or uploaded bytes. Inner requests appear first, but visiting already-resident requests still consumes the frame's allowance.

Page upload copies source rows into an independently owned buffer captured by the render command. Local arrays may be destroyed after the game-thread function returns, so their temporary pointers cannot be passed directly to a later UpdateTexture2D. The current heap buffer is freed after the render command consumes it, keeping data valid during the read.

This path updates ExistPatchs and upload statistics after enqueueing. It does not wait for a separate GPU fence before exposing those counters. Command ordering connects uploads to later rendering, but Uploaded in the log means the update was submitted, not that GPU execution time was measured independently. Cross-frame asynchronous consumers would require separate definitions for enqueueing, resource readability, and statistical completion.

Page addressing and request processing are separate steps. The first step finds a physical slot at the requested mip; the second consumes the request budget and uploads pages whose contents must change.

function ResolvePhysicalSlot(sourcePage, mip):
    delta = SignedRemainder(sourcePage - LogicalCenter[mip], SourcePageCount[mip])
    for axis in (x, y):
        if Abs(delta[axis]) > PhysicalSlotCount / 2:
            delta[axis] -= Sign(delta[axis]) * SourcePageCount[mip][axis]
    return PositiveMod(PhysicalCenter[mip] + delta, PhysicalSlotCount)

function ProcessPageRequests(budget):
    processed = 0
    while Requests.notEmpty() and processed < budget:
        page = Requests.popFront()
        processed += 1                 // Resident requests also count toward the budget
        slot = ResolvePhysicalSlot(page.xy, page.mip)
        if ExistingPage[slot, page.mip] == page:
            continue
        if not SourcePageIsValid(page):
            ExistingPage[slot, page.mip] = page
            FailedPages.add(page)      // Keep the old pixels in the slot
            continue
        if not RHITextureIsReady(page.mip):
            FailedPages.add(page)
            continue
        pixels = CopySourcePage(page)
        EnqueueRegionUpload(slot, page.mip, Own(pixels))
        ExistingPage[slot, page.mip] = page

SignedRemainder() corresponds to the source's signed % operation; PositiveMod() keeps physical slots nonnegative. The pseudocode omits source-pixel bounds checks and statistics. Source validity, failure sets, and residency records remain distinct. The missing-page branch changes the logical record without replacing pixels, so callers must also inspect failure state rather than rely only on ExistingPage or texture content.

5.2 Re-Encoding Weight Mips by Layer Contribution

Weight downsampling cannot average IDs directly. Downsampling collects 2×2 child texels, assigns each texel's 255 - Ratio to ID1 and Ratio to ID2, and accumulates those contributions by layer. It then selects the two largest contributors and re-encodes their IDs and blend ratio.

Accumulation preserves each layer's contribution before reducing the result to a top-2 approximation. Contributions from further layers are discarded, so coarse weight mips are constrained approximations rather than lossless representations. Compared with choosing the child texel with the largest Ratio, this avoids mistaking one texel's high ID2 fraction for the dominant contribution across the whole region.

Consider four child texels: grass 75% / soil 25%, grass 25% / soil 75%, and two entirely rock texels. Their accumulated contributions are 1 for grass, 1 for soil, and 2 for rock. The coarse texel retains rock and selects one of the tied grass and soil layers according to the implementation's scan order, then renormalizes the two retained contributions.

The top-2 format cannot retain the true proportions of all three materials. The discarded third layer no longer contributes to the coarse texel, and the retained pair is renormalized. Narrow snow ridges, water edges, or road blends may change as a result. Choosing source mips or rebuilding approximate mips requires comparison at the same distant camera, beyond checks of array dimensions and encoding validity.

The current function accumulates contributions across 32 layers, visits IDs in a fixed order, and updates maxima with a strict greater-than comparison. Tie results are therefore deterministic. The function supports at most 32 layers.

The 2×2 downsampling process follows. TopTwo orders by descending contribution, then ascending ID for ties, matching the current ascending-ID scan with strict greater-than comparisons.

function DownsampleWeight(children2x2):
    mass[0..31] = 0
    for pixel in children2x2:
        if pixel.id1 < 32:
            mass[pixel.id1] += 255 - pixel.ratio
        if pixel.id2 < 32:
            mass[pixel.id2] += pixel.ratio

    id1, id2 = TopTwo(mass, tieBreak = SmallerIdFirst)
    if mass[id1] == 0:
        return Weight(id1 = 0, id2 = 0, ratio = 0)
    if mass[id2] == 0:
        return Weight(id1 = id1, id2 = id1, ratio = 0)
    ratio = Round(255 * mass[id2] / (mass[id1] + mass[id2]))
    return Weight(id1, id2, ratio)

Contributions are accumulated before selecting the two layers. Selecting a dominant layer for each child first and counting those selections would lose secondary contributions and produce a different approximation.

Source-weight diagnostics in UE5.8
Figure 12 | Source weights and final shading. Upper left shows ID1 decoded from R, upper right ID2 decoded from G, lower left the B-channel blend ratio, and lower right final shading at the same camera. The diagnostic material reads full-map weight mip0 with Texture.Load, using a shared false-color palette without linear ID filtering. These views do not show coarse-weight mip selection, DTA slots, or RVT page residency.

ID1 and ID2 identify the two discrete layers participating in surface blending. Their IDs cannot be inferred from similar BaseColor values. The final image also includes road colors and other overlay branches, so road boundaries need not coincide exactly with source-weight ID boundaries. The diagnostic reads the full-map weights actually bound to the material. Coarse-weight re-encoding and mip selection still follow the rules described above.

5.3 Binding Weight Textures to the Material

PushMIDParameters() binds WeightMap to the full-map weight texture and WeightMapFar to the coarse texture. The current scene additionally binds SourceWeightMipAtlasV104 to retain source-mip weight selection.

The ring cache manages local weight pages, but the material still binds full-map data and a source-mip atlas. The ring cache has fixed capacity; the complete weight resource set does not yet have bounded residency.

BaseColor visualization in UE5.8
Figure 13 | BaseColor visualization. Roads, rock, and grass show the composited material’s color distribution. This view does not show layer IDs, DTA slots, or RVT page residency. The camera matches the WorldNormal view.

6. DTA: From Material IDs to Array Slices

The Dynamic Texture Array (DTA) manages source textures for surface layers. Weights specify the layers and blend ratio at a position; DTA resolves layer IDs to slices in color and normal arrays.

Weight GlobalId
    → IndirectionTex resolves SlotId
    → Texture2DArray samples color and normal slices
    → LayerInfoTex supplies layer parameters

The DTA core stores both GlobalId → SlotId and SlotId → GlobalId mappings, along with demand, request, pending-upload, reference-count, and failure state. Both color and normal uploads must complete before NotifyArrayUploaded() exposes a layer's mapping and refreshes patch readiness.

Color and normal slices describe one layer even though they belong to separate arrays. Publishing SlotId as soon as color completes could pair the new layer's color with the old layer's normal. Group readiness exposes the two uploads as one visible state. The core rejects unknown, stale, or duplicate callbacks to prevent old requests from overwriting new mapping state.

The core determines resize capacity, and the resource layer copies valid slices in batches. After copying, it swaps resources and rebuilds mappings. ResizeCollectBatch(), NotifyResizeCopied(), and FinalizeResize() separate copying from publication. Shrinking also requires a stable interval to avoid repeated capacity changes.

Currently, shading initialization expands the demand window to the full map and calls UploadAllNecessaryIds() to upload required layers synchronously. The acceptance scene uses 32 slots, 512×512 slices, and a complete mip chain. Per-frame Tick mainly updates the material's main-camera parameters; it does not continuously drive DTA demand from camera position.

Full-map initialization keeps every layer required by this dataset accessible, but the resident set does not shrink with camera movement. Continuous replacement through the cache core still requires runtime demand updates and coordination between upload budgets, completion state, and material readiness.

6.1 Requests, Residency, and Bidirectional Mappings

Layer demand is the union of IDs across patches in the window. Many patches can share a layer, which is requested only once when added to demand. After a window change, RebuildTextureRequest() collects the set again, sorts it by GlobalId, and assigns slots to required layers. Resident layers still in demand are removed from requests. Slots being uploaded are reserved and cannot be assigned again as idle slots.

Core state is represented by requested, uploading, and resident sets rather than a single additional enum. RequestTextureId identifies uploads not yet started; PendingTextureId identifies requests with an assigned slot waiting for all arrays; bidirectional mappings represent exposed residency. A separate set records Failed. Residency requires validating both mappings; absence from the request list alone does not establish it.

IsResidentId() checks both directions. A GlobalId can retain an old SlotId after that slot contains a new layer. Checking only the ID-to-slot lookup would incorrectly treat the old layer as resident. The slot-to-ID entry must point back to the same GlobalId. This invariant, together with material indirection updates, keeps logical IDs aligned with readable texture layers.

6.2 Group Publication and Partially Completed Cancellation

A bitmask accumulates completion across both arrays. For color and normal, completion bits are 01 and 10. Only the combined 11 state calls SetMapping() and refreshes patch readiness. A patch becomes eligible for DTA rendering after every layer it requires has been exposed.

Canceling a partial upload must handle mismatches between slot contents and mappings. Suppose a slot originally belongs to layer A and is reused for B. B's color has been written, but its normal is unfinished. Because B is not group-ready, the logical mapping may still name A even though the physical slice has changed. Treating A as resident would sample B's color with A's normal.

AssignIdleSlots() checks DoneMask when canceling such requests. A slot with only some arrays completed is evicted from the core mapping. If the old layer is still required, later requests upload it again. A request that never started does not require the same physical-content invalidation; a fully completed request awaiting collection must not be mistaken for a partial upload and evicted.

Unknown or duplicate completion callbacks are rejected so a canceled ID cannot expose its mapping later. The current interface mainly checks whether the ID remains in Pending; it has not established an independent generation-number protocol for every asynchronous request. More complex continuous asynchronous loading would require checking overlap between old and new tasks during cancellation, same-ID retries, and resource swaps.

DTA upload ordering
Figure 14 | DTA upload ordering. Mapping publication requires matching color and normal content. After partial-upload cancellation, the old layer cannot continue to identify a slot whose contents have already changed.

The following pseudocode shows completion callbacks and cancellation separately. Upload completion first sets a bitmask; group completion exposes the mapping. Cancellation removes the old mapping if the physical slot has been partially overwritten.

function OnArrayUploaded(globalId, arrayIndex):
    if arrayIndex is invalid or globalId not in Pending:
        return Rejected
    bit = 1 << arrayIndex
    if DoneMask[globalId] & bit:
        return Rejected                 // Duplicate completion callback

    DoneMask[globalId] |= bit
    fullMask = (1 << ArrayCount) - 1
    if DoneMask[globalId] == fullMask:
        slot = Pending[globalId].slot
        if not SetMapping(slot, globalId):
            return Rejected
        RefreshPatchReadyStates()
    return Accepted

Pending entries are removed during later request processing. Exposing a mapping and cleaning up its completed request object are separate operations.

function CancelPendingOutsideDemand(requestedIds):
    fullMask = AllArrayBits()
    for globalId in CopyOf(Pending.keys):
        if globalId in requestedIds:
            continue
        slot = Pending[globalId].slot
        done = DoneMask[globalId]
        Pending.remove(globalId)
        DoneMask.remove(globalId)

        if done != 0 and done != fullMask:
            oldId = GlobalIdPerSlot[slot]
            if oldId is valid:
                SlotPerGlobalId[oldId] = Unmapped
            GlobalIdPerSlot[slot] = Unmapped

Cancellation uses the DoneMask recorded by the core. The resource layer must accurately report which arrays have completed. Asynchronous upload would also require defining the order between cancellation and physical resource writes; deleting a request entry alone is insufficient.

6.3 Copying and Swapping During Resize

Texture-array capacity is defined at resource creation, so growing usually requires a new resource. When demand exceeds capacity, the core marks expansion and supplies a copy table for slices that must be retained. The resource layer copies in batches with a default per-batch entry limit to distribute scheduling costs.

Copy completion and resource publication remain separate. The resource layer reports completed destination slots, and finishing all entries enters SwapPending. After the actual texture-resource swap, FinalizeResize() rebuilds core mappings and subsequent requests. Rebuilding mappings while the MID still references the old arrays would interpret old slices through new SlotIds. Swap ordering must prevent that state.

Shrinking waits until all demand is resident and a stable interval has accumulated. Excess capacity must also satisfy NeedReduceDTAMemory(). The default stable interval is 2 seconds; DTAMin and the excess-slot count participate in the decision. These heuristics reduce repeated resource creation as cameras move around demand boundaries, but do not guarantee optimal capacity in every scene.

Old and new arrays coexist during copying, so peak VRAM exceeds final resident-array storage. Multi-frame copying limits scheduled work but does not allocate the new resource one slice at a time; the complete resource is generally already allocated. Measurements should separately record copy time, old-resource release timing, and simultaneously resident capacity.

6.4 UTexture Integration, Mip Copying, and Resource Lifetime

Material parameters receive UTexture objects, while the cache core manages slots and requests. The texture-array wrapper exposes cache resources to materials: its UObject supplies the referenced texture object, and TextureResource creates the RHI array. The array resource manager is not a UObject. The implementation uses AddToRoot to retain the texture wrapper and RemoveFromRoot when releasing resources, preventing garbage collection from removing objects still used by the resource layer.

When uploading from UTexture2D, the source can be larger than the array slice, but its dimensions must be a power-of-two multiple of the slice dimensions. The source selects a mip with MipBias=log2(SourceSize/SliceSize). A 2048 source texture and 512 slice give MipBias 2: source mip2 becomes destination mip0, and subsequent levels correspond until the destination chain ends. This path copies existing source mips rather than resizing source mip0.

A 512 slice has a complete chain of 10 levels from 512 to 1. The source needs at least MipBias+NumMips levels. Missing tail mips cause upload rejection, preventing uninitialized higher mips in the array. CPU pixel-buffer uploads follow another path and generate later mips using box downsampling. Their source data and filtering differ, so CPU test uploads are not equivalent replacements for native asset mips.

Initialization and uploads repeatedly call FlushRenderingCommands() before obtaining resources or continuing synchronous work. This waits for queued render-thread work, which is not equivalent to explicitly waiting for a GPU fence. The implementation relies on RHI command ordering and existing resource-use paths. An asynchronous version must define resource readability more precisely.

With the current four-byte format, one 512² slice at mip0 is about 1 MiB. Two 32-slot arrays total about 64 MiB at mip0, or about 85.3 MiB with full 2D mip chains. This estimates array texel storage only, excluding source UTexture2D assets, temporary textures, driver alignment, and old/new resource overlap during resize. Memory budgets must include slot count, slice dimensions, and array count.

7. Material Composition and RVT

The shading manager creates a dynamic material instance and binds weights, arrays, indirection, and layer parameters to the CDLOD component. Basic shading selects two texture layers from the weights, applies tiling parameters, and blends color, normals, and related attributes.

The reference material also includes sampling from the source seven-layer hardware texture array, source weight mips, resident-level lookup, page-source selection, and road color/opacity overlays. Together these conditions determine final layer sampling. DTA indirection describes basic resource organization; source arrays and page selection add further material input paths.

Weight, DTA, material, and RVT relationships
Figure 15 | Weight, DTA, material, and RVT relationships. The left side separates layer selection, source textures, and spatially composited results. The right side shows conditional invalidation after group readiness. RVT pages rebuild during later production, not synchronously inside the DTA completion callback.

Surface Sampling Coordinates and Layer Parameters

Height UVs locate samples across the map, while grass or rock UVs tile textures repeatedly in world space. Multiplying normalized map coordinates by one fixed coefficient would make each texture cover a larger world area as the map grows, lowering world-space texture density.

The fallback in ResolveLayerParams() adjusts tiling by the ratio of actual terrain length to reference length. This dataset spans 819200 world units versus the 102400 reference, a factor of eight. Tiling frequency increases accordingly to preserve world-space texture density. Individual layers can override values through LayerParams, so GlobalTiling is not the final tiling parameter for every layer.

Layer parameters are stored in a one-row floating-point texture containing Tiling, Specular, and FarTillingSwitch. GlobalId-to-SlotId indirection uses a one-row eight-bit texture, with 255 meaning unmapped. Both disable sRGB and use point filtering because they store IDs and parameter records; adjacent texels do not represent adjacent spatial samples that should interpolate.

Basic two-layer color blending can be illustrated as C=(1-r)C1+rC2, where HD weight r=Ratio/255 is the second layer's fraction. This explains weight semantics. The final material can also include height blending, near/far tiling, normal processing, and source road overlays. Final per-pixel layer selection requires inspecting the actual material asset's parameters and Shader.

ID debug colors differ from BaseColor. A diagnostic assigns an artificial color to each ID to reveal regions, whereas BaseColor comes from texture sampling and multilayer blending. Debug colors do not represent the material's actual base colors.

Final shading and Roughness at the same camera
Figure 16 | Final shading and Roughness at the same camera. The left shows the material under preview lighting; the right displays roughness in grayscale. These views do not show RVT page provenance or cache state.

7.1 DTA and RVT Responsibilities

Runtime Virtual Texture (RVT) stores spatial results of surface composition. It generates pages by world position; its material type determines the attributes it can store. DTA indexes source textures by layer ID. RVT assets, volumes, outputs, and sampling configuration are described in Epic's documentation.

Aspect DTA RVT
Content Source layer textures such as color and normals Material composition at a spatial position
Index GlobalId, SlotId, and slice UV World position and virtual texture page
Update Trigger Layer demand, uploads, slice replacement, or capacity changes Page requests or invalidation of cached regions
Cost Function Organizes accessible layers and texture resources Allows readers to reuse composited attributes

First-time page production and rebuilding after invalidation execute surface sampling and blending. Once a page is valid, readers can reuse those attributes. Benefits depend on reuse frequency, production cost, and residency. Current records do not include complete frame-time or hit-rate measurements.

7.2 RVT Page Production

Page production first needs scene primitives capable of writing to the RVT. The component lists the RVT assets, SceneProxy declares write support, and MeshBatch sets bRenderToVirtualTexture and the matching RuntimeVirtualTextureMaterialType. These register terrain with the page drawing path. The material's RVT output nodes then determine which attributes are written.

The RVT Volume must cover the actual XYZ bounds. This configuration preserves the acceptance volume's origin and scale and requests coarse L2–L4 page preloading during initialization. That establishes distant fallback coverage; it does not ensure all pages remain valid for arbitrary camera movement.

Page production also depends on source-texture residency. If a region produces an RVT page before its layers are ready, a later DTA upload and GlobalId mapping change do not automatically invalidate that cached result. The corresponding spatial region must be invalidated before its material composition runs again.

UpdatePatchLayerState() retains the previous patch readiness. On false→true, it calls Invalidate() with the patch's world bounding box. A switch controls this feature; its default behavior and fixed central window retain the reference implementation's limitations. Invalidation for arbitrary moving windows is not fully integrated.

Preloading and invalidation address different conditions. RequestPreload() requests pages for an extent and level to establish initial coverage. Invalidate marks existing content as inconsistent with its inputs. Invalidating every page every frame reduces reuse; never invalidating can retain stale layers. Runtime integration must select regions from actual input changes. Increasing VT pool capacity cannot replace correct invalidation ordering.

The readiness-driven invalidation process follows. CurrentConfiguredWindow() is the current source's fixed central window, not an implemented general camera-following window.

function UpdatePatchLayerState():
    if RVTVolume is absent or MarkVTDirtySwitch == Off:
        return
    for patch in CurrentConfiguredWindow():
        ready = DTACore.IsPatchReadyToRenderDTA(patch)
        previous = PatchReadyCache.get(patch, default = false)
        PatchReadyCache[patch] = ready
        if ready == previous or not ready:
            continue
        if patch == CenterPatch and CenterAlreadyInvalidated:
            continue
        RVTVolume.Invalidate(WorldBounds(patch))
        if patch == CenterPatch:
            CenterAlreadyInvalidated = true

This process handles only false→true transitions and retains the condition that the center is invalidated once. It does not cover every material-parameter change, layer replacement, or moving-camera input. Their page-rebuild requirements must be integrated according to the actual runtime material path.

7.3 Consistency Across Pages, Shadows, and the Main View

RVT page-view positions come from page mappings; shadow-view positions come from lights. If each uses its own position for CDLOD, the main image, page production, and shadows select different nodes and geomorphing states.

SceneProxy caches the main perspective view's XY position, which RVT and relevant orthographic views reuse. Before a main-view cache exists, the current view is used. This first-frame fallback and view classification still need validation in dynamic scenes.

The material separately has ProducerMainCamera, updated by UpdateProducerMainCamera() from a specified View Actor or PlayerCameraManager. The acceptance scene binds SceneCapture as its production reference. Game integration must identify the main camera and coordinate input changes with RVT invalidation. Updating material parameters alone does not ensure cached pages immediately recompute.

7.4 Multi-View Parameter Lifetime in UE5.8

GetDynamicMeshElements() can be called by the main view, multiple RVT pages, or different ViewFamilies within one frame. MeshBatch UserData points to batch parameters. Clearing the parameter array for every view would leave previously submitted batches with dangling pointers.

The implementation locks shared batch state and clears BatchParams only when the render-frame number changes. Parameters submitted by different views therefore remain valid within the frame, matching SingleFrame Uniform Buffer lifetime. The lock also protects shared state when different ViewFamilies collect meshes concurrently.

7.5 Shader Registration and Render Parameter Binding

At module startup, the plugin registers a virtual shader directory mapped to its Shaders directory, allowing the Vertex Factory to reference its shader file. The Runtime module loads at PostConfigInit so shader types and directory mappings exist before compilation needs them. A directory on disk alone does not register its virtual shader path.

The Vertex Factory supplies parameter bindings for vertex and pixel stages. The vertex stage binds constant, dynamic, instance, morph, and visibility Uniform Buffers; the pixel stage reads the constants it needs. FMeshBatchElement::UserData connects a submission's parameters to GetElementShaderBindings(), while PrimitiveUniformBuffer supplies terrain primitive transforms and related data. Depth and shadow paths need these bindings as well.

Ordinary and hole-aware Vertex Factories share a shader file and differ through CDLOD_DRAW_HOLE. The hole-aware variant evaluates per-triangle visibility; the ordinary variant uses a simplified path. Instance groups must match their variants. Passing an ordinary vertex layout to a variant requiring PrimId produces incorrect attribute interpretation even if the material compiles.

Render objects have three lifetimes. Shared grid buffers persist across frames; SceneProxy is created or replaced when component render state rebuilds; batch parameters serve the current frame's submissions. They may reference one another, but short-lived memory must not back data needed by a longer-lived object. The multi-view issue above occurs at the batch-parameter level.

7.6 Readiness of UObject, TextureResource, and RHI

A loaded UTexture2D does not imply its complete platform data, TextureResource, and RHI texture are ready. Upload code checks these stages separately. In the editor it waits for texture compilation, calls UpdateResource again, and waits for render commands. UpdateResource may replace TextureResource, so the pointer must be reacquired afterward rather than retaining the old object.

Synchronous waits simplify preview validation but increase startup time. Continuous streaming must separate blocking initialization operations rather than run them directly in every Tick. Resource-request initiation, source-mip availability, array-copy submission, and mapping publication need distinct states that advance over time, with defined failure and cancellation handling.

Component and Actor destruction must release corresponding resources. ShutdownShading() stops Tick, restores relevant material references, and releases resources; the resource layer distinguishes valid RHI state from late shutdown. Removing the visible Actor while retaining rooted texture wrappers or unreleased arrays accumulates resource use across previews. Lifetime records are needed to detect this; disappearance from the image is insufficient.

8. Independent Collision: Shared Heights, Different Topology

Visual LOD changes with the camera, but the physical surface under a character must remain stable. The collision component builds static physics geometry directly from collision heights and visibility rather than reading the morphed visual mesh.

8.1 Why 128 Samples Require 129 Vertices

A payload block contains 128×128 samples and advances its world origin by 128 × SampleSpacing. Using only those samples for collision gives 127 cells per axis. With the current 100 cm spacing, that leaves a 100 cm gap between blocks.

Collision construction reads shared rows, columns, and corners from neighboring blocks to produce 129×129 vertices and 128×128 cells. The origin step remains 12800 cm. Only the outermost full-map boundary repeats the last row or column. A missing internal neighbor is an explicit failure.

All blocks share one global height quantization step:

QuantZ = max(MaxAbsHeight / 32000, 0.01)
Hq = clamp(round(height_wu / QuantZ) + 32768, 0, 65535)

Both sides of a shared edge use the same source sample and quantization parameters, yielding equal heights and avoiding seams from independent per-block quantization.

Global coordinates address shared corners without separate directional height formulas. Local (128,128) in the current block reads (0,0) from the diagonally adjacent upper-right data block; (128,y) reads the X neighbor's first column, and (x,128) reads the Y neighbor's first row. Global addressing reduces all three cases to block-coordinate division and within-block remainder.

QuantZ is derived from the maximum absolute height across the collision payload. Independent per-block ranges would quantize the same shared height with different steps and produce different approximations. Global quantization makes edges agree before physics construction, and grid positions ensure horizontal continuity. If a required internal source block is missing, the function releases already-built collision and returns an error rather than exposing a partially initialized terrain as complete.

Collision heights are independent of encoded atlas pixels, allowing physics construction in a world without graphics presentation. Validation must also use independent references. Deriving collision expectations from the uploaded render atlas could put the same unit or stitching error into both the result and its reference.

8.2 Whole-Cell and Half-Cell Holes

Ordinary regions use Chaos::FHeightField. Fully invisible cells receive the special hole material value 255. Blocks entirely outside the valid bounds do not create physics Actors.

If only one triangle in a cell is visible, the current heightfield path cannot preserve that half-cell exactly. The block uses Chaos::FTriangleMeshImplicitObject and generates only visible triangles. Source topology is (00,11,10) and (00,01,11); collision construction adjusts winding to face normals upward.

Collision geometry selection
Figure 17 | Collision geometry selection. Ordinary regions use HeightField; half-cell holes use TriangleMesh. Reading shared edges from neighbors provides continuous block coverage.

Collision uses LOD0 triangle visibility so coarse masks do not remove nearby walkable surfaces. Of the current 4096 source blocks, 732 lie entirely outside the valid extent. The resulting 3364 collision blocks include 23 triangle-mesh blocks.

Both geometry types must participate in the same query and simulation paths. The code creates a static physics Actor, wraps HeightField or TriangleMesh in a Shape, sets owner, component, collision channels and responses, and registers it in PhysicsScene. Both SimpleCollision and ComplexCollision filtering flags are set so the specified ray and capsule-movement tests can use the appropriate surface.

Triangle-mesh vertices use the same quantized heights and horizontal spacing as heightfields. Within a cell, A, B, C, and D are the lower-left, lower-right, upper-left, and upper-right samples. The two-bit mask selects (A,B,D) or (A,D,C). Winding changes normal direction while retaining the same geometric half-cell as rendering. Substituting the other diagonal to address backfaces would make the rendering and collision holes disagree.

The original payload's physical holes and material indices also participate. A visibility-culled cell becomes a hole; a cell already marked as a physical hole is not filled just because both visual triangles are visible. Out-of-range physical material indices fall back to the default material, while the special 255 hole value remains unchanged. An unresolved material and an absent physical surface are separate conditions.

The 23 TriangleMesh blocks are the subset requiring half-cell handling in this dataset. More hole-dense data would change that proportion and may alter construction and query costs. No performance comparison has been recorded. The hybrid choice is supported by topology requirements and existing acceptance results, not evidence of optimal physics performance across all content.

The following pseudocode combines shared-edge sampling with hybrid collision construction. ReadSourceHeight() addresses neighboring blocks through global sample coordinates and repeats the final source sample only at the outermost full-map boundary.

function BuildCollision(payload, visibility, optionalWindow):
    quantZ = Max(MaxAbsCollisionHeight(payload) / 32000, 0.01)
    for block in ValidBlocksWithin(optionalWindow):
        masks = visibility.LOD0CellMasks(block)
        if All(masks == 0):
            continue

        for y in 0..128:
            for x in 0..128:
                gx = block.x * 128 + x
                gy = block.y * 128 + y
                height = ReadSourceHeight(gx, gy)
                if height is MissingInternalNeighbor:
                    DestroyBuiltCollision()
                    return Failed
                heights[y, x] = Quantize(height, quantZ)

        materials = CombinePhysicalMaterialsAndHoles(block, masks)
        if Any(mask == 1 or mask == 2 for mask in masks):
            vertices = DecodeQuantizedVertices(heights, quantZ, SampleSpacing)
            triangles = []
            triangleMaterials = []
            for each cell:
                if materials[cell] == Hole255:
                    continue
                A, B, C, D = CellVertexIndices(cell)
                if masks[cell] & 1:
                    triangles.append(A, B, D)
                    triangleMaterials.append(materials[cell])
                if masks[cell] & 2:
                    triangles.append(A, D, C)
                    triangleMaterials.append(materials[cell])
            geometry = TriangleMesh(vertices, triangles, triangleMaterials)
        else:
            geometry = HeightField(heights, materials, quantZ)
        RegisterStaticPhysicsActor(block, geometry)
    return SuccessIfAnyBlockBuilt()

For readability, the pseudocode fixes the current 128-cell configuration; the source derives dimensions from CellCount. It omits cleanup after physics Actor creation failure, component scale, and filtering parameters. Triangles uses the collision winding described above, and the two mask bits control their respective half-cells.

Local UE validation of a half-cell hole
Figure 18 | PIE ray validation of a half-cell hole. Green indicates a hit on triangle 0; red indicates a miss on triangle 1, matching visibility mask 1. The checkerboard material shows terrain topology.

8.3 Initializing a Collision Region

The terrain Actor provides InitTerrainCollision(), DestroyTerrainCollision(), and collision-block count queries. Before initialization, bCollisionUsePatchWindow and window bounds can restrict the construction region.

Specified-window initialization is supported. Continuous block creation and unloading as characters move are not yet implemented. Graphics preview does not initialize collision by default; gameplay must call the entry point explicitly in the game world.

9. Runtime Scheduling and Resource Boundaries

9.1 From Static Initialization to Continuous Updates

The current scene creates its reference image through synchronous initialization: load all heights, create full-map and cached weight textures, upload full-map DTA demand, establish RVT coverage, then create gameplay collision. This order defines module inputs, but continuous updates after camera movement still require separate integration.

The Clipmap Actor can read PlayerCameraManager automatically or use a manual viewpoint. StepView() switches to manual mode. The acceptance builder uses it to set a fixed camera, so moving the editor camera afterward does not replace that manual viewpoint automatically. The material production camera has a separate View Actor reference. Integration must explicitly connect these sources.

Continuous updates involve viewpoint reads, demand updates, resource loading, mapping publication, cache invalidation, and physical-region maintenance. This dependency order describes future integration; the corresponding unified scheduler is not implemented. Pages whose source layers are unfinished must use an explicit fallback or wait. After layer publication, dependent cached regions can rebuild.

The collision window follows character activity and need not match the main camera's render window. Distant third-person cameras, scene captures, and free cameras can move far from a character. Following only the camera could remove collision from the character's region. Multiplayer also needs a union of active-character regions or another capacity policy. Current single-window initialization does not cover that requirement.

9.2 Measuring Each Processing Path

CPU terrain costs include quadtree traversal, instance preparation, request processing, and cache-state updates. GPU costs include grid transformation, height sampling, direct material sampling, RVT page production, and page reads. Initialization also includes file reads, texture compilation, array copies, and physics object creation. These operations run on different threads and frame stages, so total frame rate alone cannot locate the bottleneck.

Performance records can pair camera trajectories with resource counts: selected nodes, ordinary and hole-aware instances, processed requests, uploaded pages, requested and resident DTA layers, RVT dirty bounds, and collision blocks. Stage timings along the same fixed trajectory distinguish steady-state cost, first-arrival peaks, and resource rebuilds.

Caches also need capacity and waiting-state measurements. With enough slots, page delay may come from a low processing budget or unavailable source textures. Insufficient slots may repeatedly evict content that will soon be needed again. Increasing capacity does not identify the cause of delay, and increasing the budget may concentrate uploads in one frame. Existing counters provide a starting point, but these measurements have not been completed here.

Graphics and physics can be analyzed separately, then validated together over the same active region. Distant-view optimization must not unload nearby character collision early, and physics initialization should not require every distant region's highest-resolution data to remain resident. Each path's resource extent must keep the active region covered by valid data.

Current Acceptance Results: all 67,108,864 source height samples match; positive rays 6628/6628; negative rays 252/252; fixed walk/run routes 40/40; maximum source-height error approximately 0.972 cm.

These results come from the existing 2026-10-07 acceptance run and cover finite samples and fixed routes. Appendix B lists full conditions and capability limits.

Appendix A: UE5.8 Project Integration

The terrain is integrated through project plugins and assets without engine-source modifications. Runtime responsibilities cover geometry and collision, weight caching, and texture resources with shading coordination.

Responsibility Main Objects Work
Visual geometry Terrain component and SceneProxy Load heights, build the quadtree, select LOD, and submit draws
Physics collision Collision component Build Chaos geometry from independent heights and visibility
Weight cache Clipmap core and Actor Manage windows, page mappings, budgets, and residency
Layer textures DTA core and resource manager Allocate slices, upload, expose mappings, and resize
Shading coordination Shading manager Connect geometry, weights, texture arrays, dynamic materials, and RVT

Scene initialization follows this dependency order:

  1. Create the terrain Actor, configure payload location, patch dimensions, and LOD parameters, then load heights and build SceneProxy.
  2. Create the clipmap, load weights, and submit page requests for the initial viewpoint.
  3. Create the RVT Volume with its asset and coverage bounds.
  4. Connect terrain, clipmap, material, and RVT; configure texture arrays and the production camera.
  5. Resolve the layer catalog, upload textures, and bind indirection and the dynamic material instance.
  6. After entering the game world, configure the collision window and initialize physics geometry.

Migration requires compatible UE5.8 modules and validation of payloads, texture assets, material switches, and RVT configuration. Preview records exist for the current project; deployment to arbitrary new projects has not been validated.

Appendix B: Validation Results and Implementation Scope

Evidence covers visual and physical behavior. Visual checks use fixed cameras and BaseColor to exclude lighting effects. Physics checks use TerrainWalker, CharacterMovement, and Chaos collision in a real PIE world. This section cites archived acceptance records from 2026-10-07. The local captures and two hole-ray checks on 2026-10-08 did not repeat the complete acceptance suite.

Test Archived Result Scope
Source heights 67,108,864 samples match independent extraction Input height consistency
Top-down coverage Raw mask differs by 3 edge pixels; no additions or omissions beyond contour tolerance Valid extent at one fixed top-down camera
Base color in the common region MAE approximately 0.378/255 with blur2 Blurred BaseColor in the shared region, not raw full-image pixel error
Close-view regression Maximum channel difference 1/255 from the preceding accepted version Version stability, not exact agreement with the source image
Positive rays 6628/6628 pass Ordinary ground, seams, and visible triangles
Negative rays 252/252 pass No hits outside bounds, in whole-cell holes, or on invisible half-triangles
Rays on both sides of half-cell holes 374/374 pass Included in the 6880 rays above
Actual walking and running 20 walking and 20 running routes pass, approximately 480.81 m total Capsule movement and seam continuity along fixed routes
Triangle-mesh seam routes 8/8 pass Included in the 40 routes above
Negative character fall tests 8/8 pass No invisible supporting surface at specified out-of-bounds locations and wide holes
Height error Maximum approximately 0.972 cm, tolerance 1.2 cm Tested rays relative to an independent source reference
Steep slope The 65.51° test point does not enter Walking Test character's 45° slope limit is effective

The test character is an existing ACharacter with capsule radius 34 cm and half-height 88 cm, walking at 3 m/s and running at 6 m/s. NullRHI disables graphics presentation while physics and movement execute. These results support acceptance of the terrain surface and specified character movement. They do not cover full animation, player input, climbing, swimming, or cave facilities.

Path Current State Remaining Work
Heights Full-map loading and resident height/normal atlas Regional requests, asynchronous reads, reclamation, and missing-page fallback
Weights Ring-cache core implemented; main material still uses full-map/coarse/source mips Complete integration of sampling with valid-page state
DTA Cache core and resizing implemented; scene initialization loads full-map demand Camera-driven demand, upload, and eviction scheduling
RVT Write registration, volume, production camera, and preloading configured Acceptance of moving inputs, invalidation timing, and multi-view consistency
Collision Independent surface, consistent holes and boundaries, optional window Character-driven window updates and block unloading
Performance No complete frame-time or residency measurements CPU selection, uploads, GPU page production, and peak memory measurements

These gaps define the remaining streaming work: request extent, residency, readiness, and page invalidation must be connected through the same data-use path. Budgeting and eviction inside an individual cache module do not make the current acceptance scene a complete large-world streaming implementation.

Why Test References Must Be Reconstructed Independently

Expected ray heights are generated independently from source collision heights, and hole expectations come from sidecar triangle bits. A sample must first be assigned to its triangle, then interpolated from that triangle's vertices for comparison with physics geometry. Bilinear interpolation across four cell corners generally differs from a fixed-diagonal triangle's height and can introduce apparent errors unrelated to the collision implementation.

Negative probes must establish both a valid test location and the expected absence of a surface. Whole-cell holes and invisible half-triangles must miss, while the visible half must still hit. A miss at the hole center alone cannot identify which half was removed. The final reference therefore covers both sides of individual triangles and retains positive probes on ordinary ground and shared edges.

Negative character tests also account for capsule dimensions. A hole narrower than the capsule can still provide valid support from its edges, which does not establish a failed hole implementation. Tests use holes wide enough for the capsule, start the fall above an independently computed source height, and check descent below the reference surface. Starting much higher and ending before the character reaches the surface can incorrectly report a pass.

Walk/run routes validate collision response over time. After actual landing, CharacterMovement drives the character across block boundaries and mixed-geometry seams. Rays check local geometry; routes check capsule movement. Together they cover local surfaces and continuous movement, which a top-down image or static hit cannot establish. Finite routes cannot exhaust every triangle, so the final report states coverage counts and capability limits.

Appendix C: Further Reading

When reviewing the implementation, check the following by module:

Area Review Focus
Payload and heights Units, encoding, block assembly, and error states
CDLOD Quadtree selection, instance data, geomorphing, and multi-view parameter lifetime
Weight clipmaps Ring centers, page mapping, request budgets, and residency
DTA Bidirectional mappings, group readiness, cancellation, and resizing
Materials and RVT Layer sampling, page production, and invalidation
Collision Shared edges, triangle visibility, physics queries, and character movement

Early seam-fix reports checked only local continuity. Later complete hole and bounds fixes addressed their unresolved failures. Physics conclusions here use the final walk/run report. Basic Lit captures remain previews and do not replace acceptance of final game lighting, normals, roughness, and post-processing.

The related procedural terrain implementation article discusses Recipe, Seed, layout, and height generation. This article starts with packaged data and explains rendering and physics geometry construction.

Related rendering references include this site's translations of tree animation in Assassin's Creed Shadows, Pixel Command and Deferred Lighting, and Path Tracing. Implementation and acceptance conclusions here are based on the terrain project records.

AI Collaboration Review

AI was used to retrieve conversations, cross-check source, organize data paths, and revise technical prose. Acceptance numbers come from existing reports. The visual work captured LOD and half-cell-hole views in UE and rechecked two PIE rays at that hole. It did not repeat the full-map ray or walk/run suite or add performance measurements.

Conversation review distinguishes cache-core capabilities from runtime integration and checks test results against the corresponding fixes. Reviewing InitShading(), material bindings, per-frame Tick, and the final collision report separated mechanisms, integration status, and acceptance scope.

Existing tests had also exposed an X/Y ordering error in the visibility reference. The reference was rebuilt from the render PrimId generation order, and probes were added on both sides of individual triangles. Test expectations can therefore contain indexing errors too. Agreement between a test and implementation still requires independent verification of the reference data.

Leave a Reply

Discover more from AI Native Game Development

Subscribe now to keep reading and get access to the full archive.

Continue reading