Open-World Population Orchestration: Identity Before the Entity, Crowds Tiered by Distance

Series The City That Runs Itself · Pedestrians · Scheduling Layer

Previous (series overview): Rebuilding a City That Runs Itself in UE5. This article picks up from it, going deep on the crowd / population part.

The traffic article covered how vehicle flow is scheduled. This one turns the lens to the crowds on the street—where they come from, where they go, when they appear, and when they are reclaimed. This is the pedestrian decision layer: population orchestration, stubs, three-tier representation, spawn budget, and crowd flow anchored on the lane graph. How they move (kinematics, animation, crowd presentation) belongs to the pedestrian Effects · Physics · Presentation article.

Intro: Crowds Are Orchestrated, Not a Pile of Independent AIs

As with vehicle flow, first dispel an intuition: that a city's crowd is "many NPCs each running its own AI." If every pedestrian actually ran a full perceive–decide–act loop, hundreds of them on screen at once would overwhelm the CPU.

The original (a mature, large open-world RPG) uses a population orchestration layer: between "the community data pre-placed in the level" and "the NPCs actually running at runtime," it interposes an entire orchestration—stubs, spawn services, three-tier representation, budgets, and bridges to the lane graph / save system / cross-system events. People on the street are not independently spawned; they are scheduled layer by layer by this orchestration according to the needs around the player.

This article follows one line—how stubs exist before entities → how the three tiers degrade by distance → how spawning is a request queue rather than instantaneous → how the budget caps the number of full entities → how crowds are anchored on the lane graph → how the stuck / out-of-range are cleaned up → how all of this is persisted—to take the pedestrian decision layer apart, and then to explain how to build a data-oriented population runtime on UE5.8, referencing the original (Mass optional, not required).

One discipline up front, and it is the soul of the whole population system: "spawning an NPC" is never an instantaneous operation, but an asynchronous chain spanning multiple stages and multiple tokens, one that may fail or time out. Treating it as "call it and it appears" is the number-one cause of population-system failure. This is the most concentrated expression, in the crowd context, of the phrase that runs through the whole series—acceptance ≠ completion.


1. Stub-First: Identity Precedes Entity Representation

① The original design

The cornerstone of the entire population system is stub-first. A "person's" first form of existence in the system is not a full NPC actor, but a lightweight stub—it has a stable entity ID, a position/orientation, and a body of persistent state, but none of the component pile a full actor carries (AI, animation, physics, perception…).

The key point: the stub exists before the entity, and attach/detach of the entity does not change the stub's identity reference. A person may first be a stub (far away, low-cost); when the player approaches, it is given an "attach" of a full entity (growing AI, animation, interactivity); when the player leaves, the entity is "dispose"d and it falls back to a stub. This can happen repeatedly, and the stub provides a stable identity layer that does not change with the attach/dispose of representation—so systems like quests, saves, and relationships can build associations on this stable identity, without binding to a volatile full actor. (Strictly speaking, a large world often has several ids—population id, persistence id, runtime entity handle; the "identity" here refers to that stable identity layer for upper layers to associate against, not an assertion that all systems use one id.)

Stubs are managed by a dedicated stub controller. There is a source-level detail here: the controller stores raw stub pointers, relying on a deletion callback (under a lock) to clear them—a performance trade-off (avoiding the overhead of smart pointers), but one that demands extreme care on the deletion path so as not to leave a dangling pointer.

Deep dive: why raw pointers—a dangerous but worthwhile trade-off

In a system with thousands of stubs, each churning every frame, the overhead of smart pointers (reference counting) is not trivial: every copy or pass atomically increments/decrements the count, with cache-line contention across threads. So the original chose raw pointers + deletion callback—when a stub is destroyed, a callback (under a lock) notifies everyone holding its raw pointer to "clear the corresponding pointer, this stub is gone." This eliminates all reference-counting overhead, but presses the safety burden onto the discipline of the deletion path: deletion must go through that callback, must be under the lock, must clear every reference; miss one and you get a dangling pointer plus a crash.

This is a classic "performance vs. safety" engineering trade-off, and one of the places to be most careful when porting. UE has TWeakObjectPtr (a weak reference that goes null when the object is gone)—far safer than a raw pointer, at the cost of a validity check on each dereference. For high-volume, high-frequency objects like stubs, use TWeakObjectPtr first for safety, and only once profiling confirms the dereference cost is a real bottleneck consider the original's raw-pointer + callback approach. Do not reach for raw pointers from the outset just to save a little overhead—safety first, optimize per measurement. This trade-off is worth writing into the porting notes.

Deep dive: the asynchronous creation of stubs—tokens, a pending table, and a "too-late, void it" guard

Stubs are not created by a synchronous new; reading the code reveals an asynchronous token mechanism. On a request to create a stub, the creator obtains a creation token and immediately pushes it into a "pending creation" table—at this instant no stub object exists yet, only a receipt of "I requested it, awaiting callback." When the stub is actually created, a callback fires; inside it, the token's status is checked for success, and only on success is the stub "taken over" and added to the container (the container never refuses—takeover is unconditional; the rejection logic lives earlier).

The most careful part is the handling of failure and races. The pending data holds a weak pointer to the traffic slot—because during stub creation, the slot it is meant to bind to may already have been deleted by the traffic system. So the callback carries a guard: even if stub creation succeeds, if the slot it is to bind to is found to have expired, this stub is voided too (along with reclaiming the slot and clearing the work-point token)—better to create one needlessly than to leave an orphan "bound to a slot that no longer exists." There is also a throttle constant: at most a fixed number of slot-request retries per frame, so that a pile of un-creatable stubs cannot stall the current frame.

This "token + pending table + void-if-late" is the standard skeleton of asynchronous spawning—more complex than a synchronous new, but nearly unavoidable in a streaming world: spawning spans frames, and the resources it depends on (slots, regions) may vanish mid-spawn, so the guard of "re-verify the premise on completion" is essentially mandatory. When porting, drive stub creation through an async handle + completion callback, and inside the callback re-verify dependencies (is the slot / lane still there?)—do not assume "it was there when I started, so it is there when I finish."

② Advantage over stock UE5.8

UE5.8's actor lifecycle is "spawn and you have a full actor." To achieve "stub before entity, entity attachable/detachable repeatedly, identity unchanged," you have to build a layer on top of the actor yourself. MassEntity has a similar "lightweight entity" idea (fragment composition), but it is a feature plugin and demands accommodating its ECS paradigm. The advantage of the original's stub-first approach is that identity and representation are thoroughly decoupled: identity (entity ID + persistent state) is stable, low-cost, always present; representation (the full actor) is attached/detached on demand, expensive, optional. This is the foundation of "a city that remembers every individual yet does not simulate every individual at once."

③ Porting to UE5.8

// Rebuild: stub-first (stub = stable identity layer; representation attachable/detachable)
struct FCyberEntityStub                    // lightweight, low-cost, always resident
{
    FCyberEntityId Id;                     // stable identity, unchanged across attach/dispose
    FTransform Transform;
    FCyberPersistentState State;           // persistent state (grudges, quest markers...)
    FCyberRepresentationHandle CurrentRepresentation;  // current representation handle (mapped to an actor internally; ACyberNpc not exposed)
};
// UCyberPopulationSubsystem owns the stub pool; approach promotes representation, retreat demotes it, identity reference stays put.
// Deliberately no TWeakObjectPtr<ACyberNpc> on the stub—that would leak "representation" into "identity," violating identity != representation.

Key points: identity (stub ID + persistent state) and representation (full actor) are decoupled, the entity is attachable/detachable repeatedly with the ID unchanged, the stub pool manages its own lifecycle. Do not use the bare UE model of "spawn = full actor."


2. Three-Tier Degradation: Full NPC / Stub / Distant Dot

① The original design

Crowds are split into three tiers of representation by distance/importance, cost decreasing and count increasing tier by tier:

  • Full NPC: near, interactive. A full actor with no components stripped—AI, animation, perception, can talk, can fight. Highest cost, fewest in number.
  • Lightweight stub / stripped actor: mid-distance. Has an ID, transform, persistent state, but a batch of components stripped by LOD (status effects, squad, visual perception, dismemberment, etc.); collision simplified to a query shape; and silent, no broadcast (it does not emit "I spawned"–type events—a performance optimization). The stub is independent of the entity: the entity can attach/dispose repeatedly while the stub stays put.
  • Distant dot: far away. Just a pure data point—it carries no full-entity semantics and takes part in no interaction logic—interpolated along the lane spline and drawn in a single GPU-instanced batch. It is primarily visual presentation (regardless of whether there is a lightweight proxy / pooled entry behind it, logically you cannot treat it as an interactive person). Lowest cost, most numerous (thousands).

Between the tiers is a bidirectional handoff: approaching promotes dot→stub→full step by step (attach); retreating demotes in reverse (dispose). One important discipline: a distant dot's "render data is valid" does not equal "there is an interactive person there"—a dot merely looks like a person; to actually interact, it must first be promoted to a stub/entity.

This three-tier scheme and the vehicle-flow three tiers of the traffic article are two instances of the same idea—both kinematic, both split into three tiers, both grown from a stub. People and vehicles sit on the same "cost is distance" curve.

Identity precedes representation: three tiers slide by cost / distance, the identity layer stays constant
Identity precedes representation: three tiers slide by cost / distance, the identity layer stays constant

Walkthrough: the life of one pedestrian (distant dot → full NPC → distant dot again)

Stringing together the three tiers, the handoff, the budget, and identity, here is what one pedestrian goes through as the player approaches: ① At the farthest range he is one dot within a patch on the plaza—pure position + orientation, drawn in an ISM batch, no stub, no ID, no logic whatsoever; ② the player approaches to a certain distance, and population orchestration decides to "land" stubs for this region—it creates a stub for him (now with a stable entity ID, position, and a body of persistent state), still lightweight, silent, components stripped by LOD; ③ the player gets closer, entering "interaction distance," and at this moment the full-entity budget still has room—the stub attaches a full NPC (growing AI, animation, perception, can talk and fight), the entity ID unchanged; ④ the player has an argument with him, even a fight—these are written into his stub's persistent state (he remembers the player); ⑤ the player leaves, he falls out of interaction distance, the full entity is disposed (releasing budget for new people in front of the player), falling back to a stub, but the stub and persistent state remain; ⑥ the player goes farther, the stub is reclaimed too, and he falls back to a dot on the plaza. Throughout, the entity ID never changes—so the next time the player returns, or on save/reload, the one encountered is still that "remembers having quarreled" him, whether he is at this moment a dot, a stub, or a full NPC. This chain runs through every mechanism in this article, and is the complete picture of "a city that remembers every individual yet never simulates every individual at once."

Deep dive: the promotion logic—not "approach and promote," but a silent contest of importance

"Approaching promotes" is the intuition, but reading the code reveals that promotion is not triggered directly by distance or time; rather, a multi-factor importance score runs a silent contest among all candidates—whoever scores highest occupies one of the limited full-entity slots. From the code structure, this score combines several kinds of factors (the exact weights and combination are implementation-specific; here only the dimensions are discussed):

  • Identity priority: those the plot hard-requires "always present" weigh highest, quest-related next, community residents after that, pure background crowd lowest—the more important the identity, the higher the base score.
  • Type: vehicles weigh markedly higher than ordinary characters (a vehicle is more conspicuous, the player more likely to interact with it).
  • Distance falloff: near is high, far is low (roughly a linear falloff on "remaining radius ratio").
  • Height difference: as the vertical gap from the player grows the weight drops, and drops sharply once large—people several floors up or down are not worth promoting (mostly out of sight and out of reach).
  • In-view / point-of-interest (POI) bonus: those the player is looking at, or within a plot hotspot, get the largest extra bonus.

The key is no broadcast, no state machine: it is not "a person at a certain distance emits an event to request promotion," but "every frame, compute the scores of all candidates and attach from highest down by slot count." This "pure-numeric silent contest" sidesteps the complexity of a state machine and naturally realizes "slot count fixed, contents flowing with player importance." When porting, do not write promotion/demotion as "distance-threshold triggered"; write it as "importance ranking + slots," so that under budget constraints the slots always go to those most deserving of promotion.

Deep dive: exactly which components are stripped / kept on promotion to a full entity

What "the stub strips components by LOD" strips specifically is, in the source, several explicit component blacklists. An entity promoted from the background crowd (the crowd tier) strips these categories of components: status effects, squad membership, visual perception, dismemberment, projectile spawning, object carrying—because a background pedestrian does not need to be under status effects, belongs to no squad, need not "see" things for perception, will not be dismembered, fires no bullets, carries no interactable objects. Stripping them saves memory and per-frame logic. There are also thriftier tiers ("simple" / "simple passenger") that strip more thoroughly (e.g. a passenger in a car even has its trigger-activation component removed—he just sits there, needing to respond to nothing).

One performance detail: the component classes to strip are looked up via one-time initialization + cached type pointers (not by string-name comparison each time), keeping the runtime fast. When porting, make "which components the stub / stripped actor strips" an explicit, tiered blacklist (one each for crowd / simple / occupant); do not gloss over it with a vague "lightweight mode"—strip wrong and you either leave useless overhead or strip away a capability it actually needs.

② Advantage over stock UE5.8

UE5.8 has no native "three-tier crowd representation + bidirectional handoff" system (MassCrowd has a similar LOD idea, but is a feature plugin, bound to ECS). The advantage of the original's approach is that each of the three tiers is a deliberately designed cost point: the full tier strips no components (it must interact), the stub tier strips components by LOD and stays silent, the dot tier is pure data. Which components to strip, at what distance to promote/demote, how the handoff preserves the ID—all are controllable, self-built logic.

In fairness, UE5.8 does provide plenty of mature facilities at the representation and batching layers: ISM/HISM to draw massive distant instances, Mass's fragment+processor for stateless batch work, SmartObject to attach interaction points, the animation budget allocator to control crowd-animation cost—these low-level capabilities are the engine's strength, and should be borrowed when self-building. What the engine genuinely lacks is the orchestration semantics that string these into "stub-first three tiers + bidirectional handoff + budgeted attach"—the engine provides the low-level facilities, the orchestration logic must be self-built. So when this article speaks of "self-building," it means the orchestration layer, not rebuilding ISM or the animation budgeter.

③ Porting to UE5.8

// Rebuild: three-tier crowd representation (cost is distance, bidirectional handoff)
enum class ECyberCrowdTier : uint8 { FullNpc, Stub, DistantDot };
// FullNpc   : full actor, no components stripped (interactive)
// Stub      : has ID/transform/persistent state, components stripped by LOD, silent (no broadcast)
// DistantDot: pure data point + ISM batch draw, no full-entity semantics, no interaction
// Promote = attach / demote = dispose, identity reference unchanged; a dot is visual only, not an interactive object.

Key points: each tier is a deliberate cost point, the stub tier strips components by LOD and stays silent, the dot has no entity. Isomorphic to the vehicle-flow three tiers.


3. Population Orchestration: From Community Data to People on the Street

① The original design

Where does the "content" of the three-tier representation come from? From community data pre-placed in the level—the designer's annotation of "what kind of people this region should have, how dense, on what schedule." The job of the population orchestration layer is to orchestrate this static community data into a dynamic crowd at runtime: by player position, time, and density parameters, decide which positions should have people right now, which people to spawn, and to which tier to attach them.

At the core of community data is a schedule (time-of-day table), whose granularity is worth seeing clearly. A community does not "spawn the same people at all hours"; it is cut by time-of-day (a time-of-day enum, daytime being the default, plus night and so on) into several spawn phases: each phase is bound to a time-of-day and records how many to spawn in this phase (quantity), at which markings / spot nodes to spawn them (marking / spot-node references), and whether to spawn in sequence (sequence—e.g. a group entering one after another rather than all appearing at once). So a routine like "morning-rush commuters, sparse late-night drunks, midday commercial-district throngs" is not hard-coded if-else, but the same community querying different spawn phases at different times of day—when the time comes, the orchestration layer lays out people by the set of phases for the current time-of-day. This is the data source of "the city has a routine, streets crowded by day and desolate by night."

This orchestration layer is a bridging hub—it connects: community data (content/routine), entity stubs (identity), the runtime entity-spawn service (building actors), traffic lanes (crowds sit on them), distant-crowd rendering, wanted-response spawning (see below), the save system (crowd-state I/O), and cross-system events (spawn/reclaim broadcast to interested systems). It does not spawn population itself, but coordinates this set of subsystems to orchestrate people on demand. One line to draw clearly: the orchestration layer only raises requirements ("this region should have this many people right now"); actually building the people still goes through the asynchronous pipeline of the next section—so "which people to spawn" here means "decide who should be spawned," not "they are there on the spot."

Population orchestration layer: a bridging hub, six-way decoupled, driven by the schedule
Population orchestration layer: a bridging hub, six-way decoupled, driven by the schedule

Deep dive: the prevention system—not "suppress spawning," but "wanted-response spawning"

There is a subsystem called prevention, a name all too easily read literally as "block/suppress the spawning of pedestrians"—but from its interface and call relationships, the direction is exactly the opposite. It behaves more like a "player-behavior response spawner": its interface is "request spawning a unit at some level," "cancel a batch by level," "query how many are currently spawned," and it carries a notion of a response level. Stringing these together, what it does is closer to—the player commits a violation (assault, gunfire, drawing attention), and by response level it spawns corresponding enforcement / hostile units to "stop" the player, the higher the level the stronger the units dispatched. That is, it dispatches more units into the city (to hunt the player), rather than fewer. (This is a characterization inferred from interface/data structure, not a line-by-line proof; but the literal reading "prevention = suppress spawning" is almost certainly wrong.)

It is also responsible for a related task—corpse cleanup: the killed in the city cannot pile up without limit, so a corpse counter monitors, deciding when to make a corpse disappear by distance, whether in view, and whether the total corpse count exceeds a threshold (clean immediately once the count is over the limit; clean once the player has turned and moved away). This spawn/despawn is likewise request queue + lock + per-frame processing (corpse updates even run only once every few frames to save cost)—again "acceptance ≠ completion."

It is called out specifically for two reasons: first, to correct the intuition trap of the name (prevention = wanted-response, not suppression); second, it is a key link in "the city reacting to player behavior": the player violates, and the street dispatches more units to hunt him; the player kills, and corpses are reclaimed by rule rather than piling up without limit. When porting, build this as a standalone "wanted-response spawn subsystem," not mixed in with ordinary population spawning, and not misled by the name into suppression logic.

prevention is really a wanted-response spawner: it shapes what gets spawned by player behavior
prevention is really a wanted-response spawner: it shapes what gets spawned by player behavior

Deep dive: reading corpse cleanup—three concentric rings, desperation deletion, and "don't rush to reuse a just-freed ID"

The corpse-cleanup code does "how to clean corpses gracefully under memory pressure" with great care, worth reading layer by layer. It runs only once every few frames (corpse cleanup is not urgent, saving cost). On a run it first collects all corpses, then processes them in three tiers by squared distance to the player:

  • Outermost ring (beyond the "off-screen cull distance"): delete immediately and forcibly—the player cannot see that far anyway, reclaim directly.
  • Middle ring (in the mid-distance band): soft delete (go through the normal deletability check, not forced).
  • Near ring (right around the player): do not delete yet, but record "the farthest corpse in view" and "the farthest corpse out of view," kept in reserve.

Most notable is desperation deletion: if the corpse count already exceeds the "immediate-cull threshold" and the total entity count has hit its ceiling (memory is genuinely critical), and this round deleted not a single one—the system will forcibly delete the "farthest out-of-view corpse"; and if there is not even one out of view, it falls back to deleting the farthest one in view (even with the player watching it vanish). This is a "gentle in the normal case, decisive in the extreme" grading—ordinarily never cleaning corpses before the player's eyes, crossing the visibility constraint only when memory is genuinely critical.

There is also an easily overlooked ID discipline: when an entity is deleted, its ID is inserted back at the front of the ID pool (not the tail). Why? So that a just-freed ID is not reused immediately—otherwise the script layer may have just recorded "No. 3 is that dead enemy," and the next frame No. 3 is handed to a newly spawned pedestrian, and the script mistakes one for another. "Don't rush to reuse a just-freed one" is a classic survival discipline of object pools. When porting this wanted / corpse-cleanup, the three rings + desperation deletion + front-of-pool ID recycling should all be carried over verbatim.

Deep dive: reading the schedule dispatcher—how multiple slots switch, and why only multiple slots register a callback

A community's "routine" lands in the code as a time-of-day dispatcher, whose slot-switching logic reads very compactly. When setting a set of slots, it first unregisters the old callback, updates the slot table, applies once immediately by the current game hour, then registers the new callback. And "which slot to pick by the current hour" has a subtlety: if there is only one slot, use it directly; if there are several, iterate the slot array back to front and take the first whose "start hour ≤ current hour"; if none is found (the current hour is earlier than every slot's start), wrap around to the last slot. This "search back-to-front + wrap-around" correctly handles slots that cross midnight (e.g. a "night slot starting at 22:00" must stay in effect into the small hours of the next day).

A small, refined optimization: a time callback is registered only when there is more than one slot—a single-slot community (the same all day) has no need to listen for time changes at all, so this subscription is dropped. When the callback fires it also deduplicates by frame number (the time system may notify several times within one frame; process only once), to prevent chatter. When porting, "community routine" should not be a per-frame query of "what to spawn now," but rather "switch triggered at slot boundaries + single-slot needs no subscription + intra-frame dedup"—the key to making "a city with a routine" both correct (correct across midnight) and efficient (no unnecessary subscriptions).

Deep dive: spawn/reclaim are events, not something others come to poll

A key design of population orchestration is that it broadcasts crowd spawn/reclaim as cross-system events, rather than letting interested systems each poll "who is here now." Why does this matter? Because an NPC's appearance/disappearance touches a whole range of other systems: the quest system needs to know "has the target NPC spawned," combat needs to know "is this enemy still there," dialogue needs to know "can it trigger," audio needs to know "should a voice source be added"… If each polled the population system "is that person there?" every frame, that would be N systems × one query per frame—wasteful, and prone to reading an intermediate state (that person half-attached).

Switch to event broadcast: when an NPC truly finishes attaching, or is truly reclaimed, the population system broadcasts one event, and subscribers respond each in turn. This way each change is processed once, and broadcast at the "completion" moment (attach fully ready / reclaimed cleanly); subscribers receive a consistent state, never meeting a half-initialized intermediate. This is the corollary of "acceptance ≠ completion": broadcast only on completion, not on acceptance—otherwise the quest system receives "NPC spawned" only to find it not yet fully attached, yielding bizarre null-pointer or half-initialized bugs. When porting, the population subsystem needs a clear event bus, and must emit events strictly at the completion moment.

② Advantage over stock UE5.8

UE5.8 has no "community data → runtime crowd" orchestration layer. The advantage of the original's approach is that it turns "how many people a city should have, where, and of what kind" from hard-coding into data-driven orchestration: community data is authored by content creators, and the orchestration layer lays it out into crowds in real time by player and rules. When porting, this orchestration layer must be self-built—it is the scheduling brain behind "the city's population feeling right."

③ Porting to UE5.8

// Rebuild: population orchestration layer (bridges community data / stubs / spawn / lanes / save / events)
// No single big Tick for all population—split into explicit phases, hooked into a custom World phase/schedule rather than a plain actor Tick.
class UCyberPopulationSubsystem : public UWorldSubsystem
{
    void TickPopulationPhase(float Dt);   // orchestrate: by the schedule, compute who should exist, whom to promote/demote (importance contest)
    void TickSpawnPhase(float Dt);        // spawn phase: drain the async request queue, throttle admissions
    void TickDespawnPhase(float Dt);      // reclaim phase: deferred deleter + orphan container + batch cleanup
    // Input: community schedule (time-of-day -> spawn phase: quantity/markings/spot/sequence) + player position + time
    // Coordinates: stub pool / entity-spawn service / traffic lanes / distant rendering / save / event broadcast
    // Wanted-response (prevention) is a standalone subsystem: spawn enforcement units by response level + corpse cleanup; keep it out of ordinary spawning
};

Key points: the orchestration layer is a coordination hub (it does not build people itself, it coordinates subsystems), driven by the community schedule (time-of-day → spawn phase: quantity/markings/spot/sequence), wanted-response (prevention) is a standalone "add people" subsystem (spawn enforcement units by wanted level + corpse cleanup); do not be misled by the name into suppression.


4. Spawning Is a Request Queue, Not an Instant Call

① The original design

This is the most important and most counter-intuitive point of the whole population system—"spawning a person" is an asynchronous request, not an instant call. The interfaces in the source (add entity, request spawn, enable/disable crowd) all return a signal of "request accepted / enqueued," not the completion of "entity now visible."

The "completion" of one spawn crosses several stages: stub-creation token → population registration queue → entity-spawn token → attach scheduling. Each step is asynchronous and may be delayed. So from "I request spawning an NPC" to "this NPC actually stands on the street, interactive" there lies a chain, any link of which may still be in progress, or may fail.

Why the indirection? Because a large world is streamed and the budget is finite: the position you requested may not have streamed in yet (the stub cannot be built), this frame's spawn budget may be used up (queue to the next frame), attach must wait for the entity-spawn service to be free. Making spawning a synchronous instant call would, first, stall (building a full actor is an expensive operation), and second, be unable to queue gracefully under streaming and budget constraints.

So the whole system must be built on "acceptance ≠ completion": requesting a spawn ≠ the person has appeared; enabling a crowd ≠ the crowd is already running. Any logic depending on "the person is already there" must await the real completion signal (the callback of the entity finishing attach), not the request's return.

Deep dive: from the source states, spawning abstracts into a four-baton relay, and any baton can break

The source has nothing explicitly labeled "four tokens," but stringing together the asynchronous states/receipts along the spawn path can be abstracted as a four-baton relay—seeing it this way clarifies why it must be asynchronous, and how each baton can break (the "baton" below is my abstract summary, not some named structure in the original):

  1. Stub-creation token: the request first needs a stub. But the stub must be built at some position, and that position may not have streamed in yet—this baton must then wait for streaming, or fail (invalid position).
  2. Population registration queue: the stub is built and must register into the population system's ordered table. But registration is enqueued, actually recorded only at the population system's processing beat—this baton is queuing for the beat.
  3. Entity-spawn token: the registered stub must attach a full entity (build the actor). But building a full actor is an expensive operation and is bound by the full-entity budget—when the budget is full, this baton queues for someone else to dispose.
  4. Attach scheduling: the entity is built and must attach onto the stub (mount components, AI, animation, restore persistent state). This baton may complete across frames; only when attach is done is "this person standing on the street, interactive" truly real.

A four-baton relay: any baton may be delayed (awaiting streaming / beat / budget / a frame boundary) or fail (invalid position / cancelled after budget exhaustion). So from "I requested a person" to "he is really there" is a chain that may break at any baton. This is why the system is full of graded states (Queued / StubCreated / EntitySpawning / Attached / Failed / TimedOut), and why logic depending on "the person is present" must await that final attach callback—not the intermediate return of any baton. Compressing these four batons into one "spawn success/failure" boolean is equivalent to assuming this relay never breaks midway—and it breaks often.

The four-baton spawn relay: acceptance ≠ completion, any baton can break
The four-baton spawn relay: acceptance ≠ completion, any baton can break

② Advantage over stock UE5.8

UE5.8's SpawnActor is synchronous and instant—the call returns and you have an actor. This is fine in a small scene, but insufficient for a population system with "large-world streaming + budgeting + three-tier handoff." The advantage of the original's async request queue is that it naturally fits streaming and budgeting: requests enqueue, and the system truly spawns only when streaming is ready and the budget allows, the graded state queryable throughout. When porting, you must build this async orchestration layer atop SpawnActor, not spawn synchronously and directly.

Worth stressing: UE's synchronous SpawnActor semantics actively induce developers to write a wrong population system—because it returns an actor immediately, one naturally writes "need one → SpawnActor → use it on return," then discovers, with hundreds on screen, per-frame spawn stutter, unhandled spawn failure when the position has not streamed in, and memory overflow for lack of a budget. The original sidesteps all these pits with an async queue, but UE's instant API hides them—only exposing them once scale arrives. So porting this layer is not "icing on the cake," it is mandatory to wrap an async request queue atop SpawnActor—demoting instant spawn to "the execution detail of the last step," fronted by "enqueue → throttle → streaming-readiness check → budget check → graded state." This is also why this article stresses "spawning is a request, not a call" repeatedly: it points straight at the UE API most easily misused.

③ Porting to UE5.8

// Rebuild: spawning = async request queue (acceptance != completion, across multiple stages)
FCyberSpawnHandle UCyberPopulationSubsystem::RequestSpawn(const FCyberSpawnDesc& Desc);
// Returns a handle (acceptance), not an actor (completion). Internal stages:
//   stub-creation token -> population registration queue -> entity-spawn token -> attach scheduling
enum class ECyberSpawnStatus : uint8 { Queued, StubCreated, EntitySpawning, Attached, Failed, TimedOut };
// Logic depending on "the person is present" awaits the Attached callback; do not rely on RequestSpawn's return value.

Key points: spawning is an async request, returning a handle not an actor, completion spans multiple tokens/stages, logic depending on "the person is present" awaits the attach callback. This is the foundational discipline of the whole system.

Deep dive: reading spawn throttling and "spawn takes priority over delete"

The spawn queue has two scheduling disciplines only noticed on reading the code. First, per-frame spawning is hard-throttled—at most a fixed handful of stubs are actually spawned in a frame. Why? Because building stubs and attaching entities is expensive; even a hundred spawn requests flooding in within one frame can only be processed step by step; without throttling, the moment the player enters a new region and a large batch lands at once, there is an instantaneous stall. So spawning is a "steady trickle": requests may arrive in bursts, their landing throttled at an even rate.

Second, "spawn takes priority over delete" is coded explicitly. When processing the delete queue, it first checks whether this to-be-deleted entity is still in the spawn queue, not yet spawned—if so, it removes it straight from the spawn queue and returns the ID to the pool, not spawning it at all (since it is about to be deleted, why spawn then delete, adding one expensive operation needlessly). Further: whenever a new spawn request arrives, it clears the entire delete queue—the meaning being "since we are spawning again, defer those pending deletes." This is an "avoid wasted work" scheduling strategy: spawn and delete are a pair of expensive operations; cancel out what can be cancelled, defer what can be deferred. When porting, spawning needs per-frame throttling, and "spawn vs delete" must be able to cancel out (a pending delete not yet spawned is cancelled outright, an incoming spawn defers deletion), rather than mechanically "request → spawn, mark delete → delete."

Incidentally, a symmetric ID-pool design: allocation pops from the back of the pool, recycling inserts at the front—out one end, in the other, guaranteeing that a just-recycled ID takes a full round through the pool before being handed out again. This is the same discipline as the wanted system's front-of-pool ID recycling in the previous section: the more recently an ID was used, the longer it should wait before safe reuse.


5. Spawn Budget: When a Full Entity Can Be Attached

① The original design

The costliest of the three tiers is the "full entity" (attaching a full NPC). So a budget constrains it: how many full entities may be attached at once. Stubs can be many (low cost), but the number attached into full entities has a ceiling—exceed it and you must queue, awaiting some other entity to dispose and free budget.

This budget is the gate for "a whole city of crowd without stutter": the thousands of distant dots take no budget (pure data), mid-distance stubs take a very small budget (lightweight, an independent pool), and only the truly attached full NPCs nearby occupy the main budget. The player turns, the people that should be attached in view change, and within budget the system disposes those behind and attaches those in front—budget constant, contents flowing with the player.

The stub pool itself also has an independent budget, separate from the "full-entity" budget. This makes "many stubs but few full entities" the natural state—most individuals exist at low cost as stubs, and only those the player is looking at, potentially interacting with, are promoted.

Importance contest + tiered budget gate: promotion is decided by ranking + slots
Importance contest + tiered budget gate: promotion is decided by ranking + slots

② Advantage over stock UE5.8

UE5.8 has no "full-entity attach budget + independent stub-pool budget." The advantage of the reference implementation is that it uses a tiered budget to pin the cost of "a whole city of people" to a constant value: no matter how many people this region theoretically should have, the number of full entities attached at once has a hard ceiling, and exceeding it queues. This is the fundamental gate keeping frame rate from exploding with population scale.

Deep dive: reading the budget—two budget lines, read from a table by quality tier, and "if it can't attach, fall back to staging"

In the code, the budget is actually two lines: an attached-entity budget and an unattached-entity budget, each with a current usage and a ceiling. These two ceilings are not hard-coded but read from a config table (CSV) by quality tier—the low-quality tier has its own, smaller attached ceiling. So "how many full NPCs a city has at once" is configurable and scales with quality: on a low-spec machine it automatically attaches a few fewer, frame rate first.

The spawn algorithm is careful about "not wasting": it sorts the spawn queue by importance and tries in turn to attach each into the "attached budget"; if it cannot attach (budget full), it is not discarded outright but converted to an "unattached entity" for staging—the entity is already built, just not yet attached to a full representation, placed in the unattached budget to queue until the attached budget frees a slot. This way the expense of building the entity is not wasted. Only a genuine spawn failure (invalid position and the like) is cancelled immediately, removed from the queue, not retried.

These three layers—attached budget (expensive, quality-dependent) + unattached staging (built, awaiting attach) + fail-and-discard (no retry)—balance "waste as little as possible under budget, yet do not retry endlessly." When porting, do not hard-code the budget as a single number; make it "read from config by quality tier + two levels of attached/unattached + stage rather than drop when it can't attach."

③ Porting to UE5.8

// Rebuild: tiered budget (full-entity budget + independent stub-pool budget)
// Full-entity attach count has a hard ceiling -> exceed it and queue, awaiting a dispose to free budget
// Stub pool has an independent budget (lightweight), dots take no budget (pure data)
// Player moves -> within budget, dispose those behind / attach those in front; budget constant, contents flowing.

Key points: full-entity attach has a hard budget (frame-rate gate), the stub pool's budget is independent, dots take no budget. Budget constant, contents flowing with the player.


6. Crowds Sit on the Lane Graph: People Flow on "Lanes" Too

① The original design

A key fact stitching the traffic article to this one: crowds also sit on the lane graph. The lane graph the traffic article described is not for vehicles only—a sidewalk is also a "lane," and the way people flow on it is isomorphic to how vehicles flow on carriageways: both take slots on the lane, both integrate arc length along the lane, both are subject to occupancy constraints.

Population orchestration shares lane events and the slot mechanism with the traffic system through a people–vehicle coordinator. Many of the people walking the street are on "traffic work-points"—anchors annotated on the lane for "what a person should do here" (wait for a bus, window-shop, cross the road); crowds are assigned to these work-points on spawn.

There is a source-level detail here, embodying "acceptance ≠ completion" and "graded processing": in the crowd's response to lane events, it ignores the "lane blocked" event (unless in monitor state), and defers the cleanup of dead-ends/congestion to the delete phase to do uniformly—not hastily deleting people the instant a block is received, but accumulating to the delete phase for batch processing. This is again the "event-driven + batch" scheduling paradigm.

Deep dive: reading lane density—how "how many people this lane should have" lands as spawn/delete

How does the crowd know how many people a lane should have, deleting when over and adding when under? Reading the code, via a per-lane-segment density ledger. Each lane is cut into several segments (fragments), and each segment records a "deficit density"—target headcount minus actual headcount. When this number is positive, the segment is under-peopled and should be topped up (spawn requests are directed here); when negative, it is over-peopled and should be trimmed (the "a lane has too many people" from the earlier cleanup section is triggered by reading this negative value). On deletion it also sorts by "deficit density," the most "over-supplied" lane segments trimmed first.

The benefit of this "per-segment deficit density, top up on positive, trim on negative" is that spawn and delete share one ledger—not two independent logics each guessing, but one density target, topping up or trimming as needed. The density target itself comes from community data (how crowded that lane should be by day, how empty by night), so "the density of the street" runs from the data all the way to "how many people each segment should add or delete." When porting, the crowd's spawn/delete should not each write its own trigger conditions but share one "per-lane-segment target vs. actual" density ledger: positive throws a spawn, negative triggers a delete, priority sorted by surplus—so density is controllable and spawn/delete do not conflict.

People and vehicles share one lane graph: pedestrians sit on edges, topped up / deleted by segment density
People and vehicles share one lane graph: pedestrians sit on edges, topped up / deleted by segment density

Deep dive: how distant-crowd points find the lane they belong on

There is a dedicated distant-crowd lane finder that assigns to a distant crowd point "which lane it should flow along." A distant crowd is a patch of points moving along routes (next section), but "which route" is not handed out arbitrarily—the finder must locate a suitable sidewalk lane around the crowd point and bind the point to it, so they appear to walk on the sidewalk rather than pass through walls or float in the middle of the carriageway.

This and the traffic vehicle's "spawn query" (the sub-segmented free-space search from the traffic article) are two versions of the same class of spatial problem: a traffic vehicle must find a long-enough gap on the carriageway to place a car; a distant crowd point must find a home on a sidewalk lane. They share the "lane graph + spatial index (BVH)" infrastructure—again confirming the advantage of people and vehicles sharing one graph. When porting, the distant crowd's lane lookup directly reuses the traffic article's lane graph + BVH; no separate spatial structure need be built for people.

Work-point assignment also happens here: after a crowd point is assigned to a lane, it may be directed to a nearby work-point (wait for a bus, look at a shop window, cross the road), so it does not merely "walk along the lane" but "walk to the work-point, do a spell of what it should do, then move on." This makes the crowd look purposeful, rather than forever drifting along the sidewalk at a constant speed.

② Advantage over stock UE5.8

In UE5.8, people (MassCrowd) and vehicles (MassTraffic) are two plugins with two representations; sharing lanes means bridging them yourself. The advantage of the original's "people and vehicles share one lane graph and one slot mechanism + work-points" is unification—people and vehicles flow in the same coordinate frame, and interaction (a car yielding to a pedestrian at a crossing) has a common basis. When porting, if you self-build (not using Mass), let people and vehicles share this lane graph from the very start, avoiding the fracture of bridging two plugins.

③ Porting to UE5.8

// Rebuild: crowds sit on the lane graph (shared with vehicles, anchored by work-points)
// Sidewalk = a class of lane in the lane graph; a person takes a slot and walks by arc length, isomorphic to the traffic article.
// UCyberCrowdTrafficCoordinator: shared lane events + slots + work-points.
// Lane-blocked event: ignored (except in monitor state), dead-end cleanup deferred to the delete phase for batch processing.

Key points: crowds sit on the lane graph (shared with vehicles), work-points anchor crowd behavior, lane-blocked cleanup is deferred to the delete phase for batch processing (no need to rush the instant one is received).


7. Cleanup: Stuck, Out of Range, Spawn Failure

① The original design

As crowds spawn continuously, they must be reclaimed continuously. The original's delete path handles several categories of case, all batch-processed uniformly in the delete phase rather than deleted piecemeal at any time:

  • Out of range: a person has walked out of the active region around the player and should be reclaimed.
  • Stuck: a person is stuck on geometry and can no longer move; clean it up and replace it.
  • Spawn error: something went wrong during spawning (invalid position, a half-built thing after budget exhaustion); cancel it cleanly.
  • Lane dead-end / congestion: the lane the person is on became a dead end; the people on it should be evacuated or reclaimed.

All these deletions go through a deletion callback, and the stub controller holds raw pointers—so deletion must clear under the lock, in order, leaving nothing dangling. Concentrating deletion into a batched "delete phase" rather than deleting anytime is to avoid concurrently modifying the crowd collection while iterating it (a classic iterator-invalidation / concurrency bug).

Deep dive: reading the deferred deleter—brief fade-out, FIFO, and "demote before delete"

Reading the delete path reveals that when a stub is judged deletable, most of the time it is not deleted at once but goes through a deferred deleter. The flow is deliberate: first set non-resident, then lower priority to the crowd tier, then trigger fade-out, and finally push it along with a timestamp into a FIFO queue. Each frame it checks the head: current time − enqueue time ≥ a brief fade-out duration before it truly deletes.

Three details worth noting: First, "demote" before delete—setting non-resident and lowering priority is to prevent it from being promoted back up by the "importance contest" during the fade-out (a person on the way out should not be attached again). Second, FIFO guarantees delete order = request order, memory release is predictable, aiding performance analysis. Third, only those "deletable immediately" (e.g. already out of view, utterly useless) go through synchronous delete; the rest all take this fade-out—so you almost never see a person vanish abruptly, always dissolving away (this also matches the dissolve show/hide of the traffic-effects article).

Deletion has many entrances but one exit: whether the reason is "reached destination," "stuck in traffic," "died," "out of range," or "a lane has too many people," all funnel finally into one delete function—which uniformly hands the stub to the orphan container, removes the slot from the traffic system, and decrements the in-range count. Several of its paths have precise thresholds: the traffic-stuck must be "out of view and far enough from the player" to be deleted (a nearby stuck one is kept, lest it vanish before the player's eyes); the out-of-range, before deletion, first requests a dot stand-in from the distant system (the entity is removed, but that far-off point remains, visual unbroken); the over-crowded lane picks a few at random to delete, and only those that have crossed the "last visible time" threshold (just-seen ones are not deleted). When porting, this "many entrances → single-exit delete function + each path's view/distance/time thresholds" should be carried over—it is the guarantee that "reclamation never happens abruptly before the player's eyes."

Deep dive: reading the orphan container—where do detached stubs go

The delete path repeatedly features an "orphan container"—stubs detached from various places are not destroyed directly but first "handed to the orphan container." This is a way station, and reading its implementation there are two points of care. First, it stores sorted by entity ID, inserting with a binary search (lower_bound)—because orphans may be many, sorting makes "look up by ID whether an orphan is present" logarithmic rather than a linear scan. Second, it lives and dies with the "event-bridge object": when a stub enters the orphan container, an event-bridge object is created for it; on deletion, that activator is deleted along with it. The event-bridge object is the hook that lets the script layer capture stub events (spawn complete, entered view…)—so although an orphan stub is logically "retired," the script layer can still sense it through its trigger.

Why such a way station rather than delete-on-detach? Because between "detach" and "truly destroy" there is still work to do—fade-out, notification, awaiting a safe moment. The orphan container is where this "retired but not yet destroyed" buffer period belongs. This and the earlier deferred deleter are a paired two layers: the deferred deleter handles "time the fade-out then delete," the orphan container handles "sorted storage + trigger hooks", together forming the complete "graceful stub exit" flow. When porting, "stub retirement" needs such a transit buffer; do not delete on detach—that would leave fade-out, script notification, and safe timing nowhere to be carried out.

Graceful exit: deferred deleter + orphan container + corpse three-ring cleanup
Graceful exit: deferred deleter + orphan container + corpse three-ring cleanup

② Advantage over stock UE5.8

UE's actor destruction is instant, but "multiple reclaim reasons (out-of-range / stuck / spawn-failure / dead-end) concentrated into batch processing in the delete phase + raw pointers cleared under the lock" is a self-built scheduling discipline. The advantage of the original's approach is that it gathers the dangerous business of "crowds churning constantly" into one safe batch point—avoiding the hardest-to-find bugs like concurrent collection modification and dangling pointers.

③ Porting to UE5.8

// Rebuild: crowd cleanup (multiple reasons concentrated into batch processing in the delete phase)
enum class ECyberDespawnReason : uint8 { OutOfRange, Stuck, SpawnError, LaneDeadEnd };
// Accumulate to-be-deleted stubs -> uniform batch processing in the delete phase (clear raw pointers under the lock, in order, nothing dangling)
// Do not delete concurrently while iterating the crowd (iterator invalidation / concurrency bug).

Key points: multiple reclaim reasons concentrated into batch processing in the delete phase, cleared under lock protection, do not concurrently modify the collection while iterating.


8. Distant Crowds: Visible ≠ There Is a Person

① The original design

The farthest of the three tiers—the distant crowd—deserves a separate note on its conceptual discipline. A distant crowd is a patch of points interpolated along lanes and GPU-instanced for rendering. What it provides is the availability of visual data, not "a swarm of living crowd entities."

That is: the milling crowd on the far plaza is valid in rendering (visibly there are people), but logically there is no entity there—no individual NPCs making decisions, only a patch of points moving along routes, drawn in a batch. As the player approaches, the needed ones are promoted to stubs/entities.

Deep dive: reading the distant-point update—when a dot must "respawn"

Distant points are not static; they update every frame, and the most interesting part of the update logic is when to respawn a dot elsewhere. Reading the code, the respawn trigger is a set of "or" conditions: its lane index became invalid, it is too far from the player, it has been stuck for more than a certain number of frames, this lane should shed points, or it entered the distance of the regular crowd and should be upgraded. There is also a counter-intuitive guard: if a dot is within the player's view frustum yet too close (close enough that it should soon upgrade into a real entity), it is likewise force-respawned—avoiding "a point that ought to become a real person still floating there as a mere point." On respawn it tries at most a fixed number of times, each time picking a probability-weighted random lane, resetting its speed, position, and stuck counter.

That "respawn after a few stuck frames" counter is crucial: the distant tier creates the sense of vehicle/pedestrian flow by "points flowing along lanes at constant speed," and once a point stalls because of congestion ahead, it breaks the illusion of "flowing"—so it is simply respawned elsewhere rather than stuck alongside. The distant tier does not solve congestion, it bypasses it: stuck, it starts over somewhere else. This is a different logic from a nearby real entity's "stuck → go to cleanup deletion"—near must be truly handled, far need only "look like it is moving."

Underpinning this is a constant lane network: at init, all lanes are preprocessed into a batch of "in-degree and out-degree both ≤ 1" straight-path chains (a road one can walk straight down, with no forks or merges), and distant points interpolate along these chains. Construction uses cycle detection—if tracing a chain loops back on itself, it is marked invalid (rather than deleted, kept for debugging). Why preprocess into constant straight-path chains? Because distant points need no real pathfinding and no decisions at junctions; they only "walk straight down one road," and a straight-path chain provides exactly this lowest-cost "road." When porting, distant crowds/flows take "preprocessed straight-path chains + points interpolating along chains + respawn on stuck + ISM batch draw"; give distant points no real pathfinding or real AI. This discipline is identical to the traffic article's dot: valid render data ≠ there is an interactive object. Confusing the two—thinking every point in a distant crowd is a person you can talk to—is the conceptual error streaming population most easily commits, leading to "I clearly see someone there, but walk over and there is nobody," or the reverse, and various anomalies.

② Advantage over stock UE5.8

UE's ISM/HISM can draw massive instances—free at the low level. The advantage of this design is not in rendering but in the conceptual division that makes clear a distant crowd is "visual data," not "entity completion." When porting, draw distant crowds with ISM, and logically insist "they are not entities."

A deeper advantage is that "constant straight-path chains + respawn on stuck" motion model: it makes distant crowds "look like they flow with purpose" yet cost almost no CPU—no pathfinding, no decisions, no avoidance, only "interpolate at constant speed along one preprocessed straight road, and after a few stuck frames start over on another." If the UE side rashly runs even a little AI on distant points (however light), thousands of points still accumulate into considerable overhead; the value of the original's approach is exactly pressing the distant point's "intelligence" to zero, keeping only "looks like it is moving." This boundary must be held firmly: the distant tier solves "there is visually a patch of flowing people/vehicles," never "where each is headed and how it avoids obstacles"—those are matters after promotion to a stub/entity. When porting, any real logic given to a distant point is wasted overhead on the very tier that should be lowest-cost.

③ Porting to UE5.8

// Rebuild: distant crowd = visual data (ISM batch draw), not entities
// Points interpolated along lanes -> UInstancedStaticMeshComponent batch rendering
// Discipline: valid visual data != there is an interactive person; promote to a stub/entity only when needed.

Key points: a distant crowd is visual data, not entities; draw with ISM, and logically insist "visible ≠ there is a person."


9. Persistence: The City Remembers Everyone

① The original design

The population system must save and restore with streaming. Because identity hangs on the stub (entity ID + persistent state), the save mainly stores the stub-layer state: who is where, and what each one's persistent markers (quest progress, relationship with the player, alive-or-dead) are. The full entity is a volatile representation, rebuildable from the stub at any time; the stub is the identity to be saved.

The effect: the people a player has killed in a region, helped, or the events triggered—leave and return, save and reload, and they are still remembered. Because these hang not on the volatile full actor but on the persistent stub ID + state.

Deep dive: reading "create vs restore"—the two paths must be strictly symmetric

Reading the stub-construction code, one sees it splits into two modes: create and restore (reload / streaming load-in). These two paths must be strictly symmetric, or persistent state will duplicate or be lost. Create runs "enable persistence → create base-data persistent state → create each component's persistent state"; restore runs "skip initial persistence setup, directly access the existing persistent state → callback a 'restored' notification for each component's persistent state." Both paths have a precondition check: before create, the component persistent state must be confirmed not to exist (else duplicate creation); before restore, it must be confirmed to exist (else nothing to restore)—fail the check and it asserts outright.

What best shows "how important symmetry is" is the destructor: on stub destruction it hard-asserts that its base-data persistent state has been nulled—that is, a stub may be destroyed only through the proper deletion flow (which first handles the persistent state properly), never deleted offhand. This assertion is the last sentry against "someone bypassing the deletion flow and leaving half-cleaned persistent state."

Why insist on "symmetry" so hard? Because persistence bugs are among the hardest to find—saved-but-not-restored, or restored-then-saved-again, the symptom often "some NPC's state is subtly wrong after reload," reproducing only across a save/reload cycle. The original walls off this class of bug at compile/assert time with "two symmetric create/restore paths + precondition existence checks + destructor assertion." When porting, a stub's (and any persistent object's) "create" and "restore" must be written as two strictly mirrored paths, each with an existence check—do not let "restore" reuse "create"'s path with patches on top; that is a breeding ground for persistence bugs. This shares the philosophy of the vehicle article's "per-vehicle persistence": save identity and persistent state (stub / per-vehicle state), not the volatile representation (full actor / physics frame state).

Deep dive: saveable vs non-saveable IDs—not everyone is worth putting in the save

An easily overlooked but important distinction: not every stub goes into the save. The original distinguishes entity IDs as "saveable / non-saveable"—important, plot-related NPCs the player has interacted with use saveable IDs and are kept in the save across reloads; while pure background pedestrians (soon reclaimed, the sort the player will never remember) use non-saveable IDs and do not enter the save—reclaimed is reclaimed, a fresh batch randomly respawned after reload will do, no need to spend save space on "where some nameless pedestrian is at this moment."

This distinction is the key to keeping the save from bloating. A city has thousands of stubs at once; if every one went into the save, the save would grow absurdly large and slow to read/write. So the system persists only "the few worth remembering"—those the player interacted with, plot-marked, quest-related—while the rest of the background crowd is regenerable and non-persistent. Judging which ID a person should use is itself an orchestration-layer decision (is this NPC important? might the player remember him?). When porting, this "saveable / non-saveable" split should be built into the stub design early—else either the save bloats, or those who should be remembered are not.

② Advantage over stock UE5.8

UE has a save framework, but the identity/representation-separated save design of "stub-layer persistence + full entity rebuilt from the stub" is self-built. Its advantage is that the save need only store low-cost, stable stub state, without serializing a mass of volatile full actors—both saving space and avoiding the "NPC state scrambled after reload" bug.

③ Porting to UE5.8

// Rebuild: population persistence (save stub identity + persistent state, not the volatile entity)
// Save = stub layer (EntityId + FCyberPersistentState)
// Full entities are volatile, rebuilt from the stub on streaming/reload
// The city "remembering everyone" rests on the stable stub ID + persistent state, not on full actors.

Key points: save stub identity + persistent state, rebuild the full entity from the stub. Same philosophy as per-vehicle persistence.


Appendix: A Code-Level Walkthrough—One Pedestrian's Life from Spawn to Reclaim

Stringing the code paths read into one line, see what a pedestrian actually goes through in the code (each step corresponds to an earlier deep dive):

  1. The community decides he should exist: the schedule dispatcher, by the current time-of-day, selects a spawn phase, which says "this region should have this many people of this kind."
  2. Issue an async spawn request: the request enters the queue, takes an ID (popped from the pool's back), and returns a handle—there is no person yet. Per-frame throttling admits only a few.
  3. Asynchronously create a stub: the creation token enters the pending table; the callback verifies "is the slot to bind still there," and if so takes it over into a stub and adds it to the container; the ID is locked throughout.
  4. An importance contest decides whether to promote: together with all surrounding candidates, he competes every frame, by a multi-factor importance score (identity priority + type + distance + height difference + in-view/POI bonus), for the limited full-entity slots.
  5. Promoted to a full NPC: a slot is free and he made the cut → the stub attaches a full entity, stripping by blacklist the components he does not need (status effects/squad/perception, etc.), and restoring his memory from persistent state (his grudge with the player remains).
  6. Alive, affected by the player: the player's interactions with him are written into his stub's persistent state.
  7. The player leaves, he is demoted: the full entity disposes and frees a slot, falling back to a stub—but via the deferred deleter: first demote, briefly fade out, enter the FIFO, and before deletion pass through the orphan container (sorted + trigger hooks).
  8. Farther, degraded to a dot: even the stub is reclaimed, and he becomes one distant point, flowing by interpolation along a constant straight-path chain; after a few stuck frames he respawns elsewhere.
  9. If he dies: the corpse counter cleans by the three concentric rings, and the ID is inserted back at the front of the pool (not rushed into reuse, lest "that dead person" in a script be mistaken for another).

Throughout, his identity (ID + persistent state) is stable while his representation (dot / stub / full NPC) changes repeatedly—this is exactly what "a city that remembers every individual yet never simulates every individual at once" really looks like at the code level: a string of state transitions carefully orchestrated among tokens, queues, the importance formula, fade-out timing, and the orphan container.

One character's lifecycle: representation moves up and down, identity is always present
One character’s lifecycle: representation moves up and down, identity is always present

10. Iron Laws Read from the Code: The Discipline Behind a Pile of Constants

Extracting the "only-known-by-reading-the-code" details of the earlier sections and viewing them together, one finds the population system's stability rests on a set of disciplines each seemingly trivial yet each walling off a class of bug. When porting, these are worth more than the architecture diagram:

On "not trusting the present" (defense under streaming churn)

  1. Stubs trade raw pointers for performance, but deletion must go through the callback, under the lock, cleared clean. Worth it only at high volume and frequency; on the UE side use a weak reference for safety first, optimize per measurement.
  2. On async-creation completion, re-verify dependencies. During stub creation the slot it is to bind to may already be deleted—if the completion callback finds the slot expired, void this stub, leaving no orphan.
  3. "Update" and "delete" may happen in the same frame, across systems. Before a critical write, add a guard for "is it about to be deleted / does it have a parent."

On "reclamation must be graceful" (never abrupt before the player's eyes) 4. Deletion is delayed by a brief fade-out, via FIFO, demoted before delete (set non-resident + lower priority, lest it be promoted back during the fade-out). 5. Many-entrance deletion funnels into a single-exit function, each path carrying view/distance/last-visible-time thresholds. 6. Retired stubs transit through the "orphan container" (sorted by ID + trigger hooks); do not delete on detach.

On "slots and IDs" (no waste, no cross-numbering under budget) 7. Promotion is a silent contest of a multi-factor importance score (identity priority + type + distance + height difference + in-view/POI bonus), not distance-threshold triggered, no events emitted. 8. The full-entity budget is read from a table by quality tier; if it cannot attach, convert to "unattached staging" rather than drop; on failure discard, no retry. 9. Spawning is throttled per frame (a fixed handful), and "spawn vs delete" cancel out where possible. 10. A recycled ID is inserted at the front of the pool—a just-freed one is not rushed into reuse, preventing script mis-association.

On "decisiveness when memory is critical" 11. Corpse cleanup in three concentric rings (outer delete immediately / middle soft delete / near keep); only in the "desperation" moment—total over the ceiling and not one deleted—does it cross the visibility constraint to delete one before the player's eyes.

On "persistence symmetry" 12. Create / restore must be strictly mirrored, each with an existence check, the destructor asserting persistent state cleared—walling off the hardest-to-find reload bugs like "saved-not-restored / restored-and-saved-again."

The common undertone of these disciplines is the same as the vehicle-scheduling article's ten: in a streaming, parallel, ever-invalidating environment, do not trust any "what I hold now will still be there, still valid, the next moment." That a city can "remember every individual yet not simulate every individual at once" rests not on some clever architecture, but on doing these dozen-odd disciplines without missing one. One more on locks: every critical section in the population system is extremely short (guarding only queue/table insert-remove, no I/O or complex computation inside the lock)—spawn/delete requests often come from other threads (script, input), and short critical sections + per-subsystem locks are the premise of "thread-safe yet jitter-free."


11. Migration Trade-offs: Reference the Original, Self-Build a Data-Oriented Population Runtime

To close. This article's porting stance: prefer referencing the original and self-building a data-oriented population runtime; Mass is an optional upgrade, not a requirement.

Mechanism Use Don't use Why
Identity Stub-first (ID + persistent state, independent of the entity) ❌ spawn = full actor Identity/representation decoupling
Three tiers Full NPC / stub / dot, bidirectional handoff ❌ single representation Cost is distance
Orchestration Community data → orchestration layer → crowd — Data-driven; prevention = wanted-response spawning (standalone subsystem)
Spawning Async request queue (acceptance ≠ completion, multi-token) ❌ synchronous SpawnActor Fits streaming + budget
Budget Full-entity attach hard budget + independent stub-pool budget — Frame-rate gate
Anchoring Crowds sit on the lane graph (shared with vehicles + work-points) ❌ bridging two Mass plugins Unified coordinate frame
Cleanup Multiple reasons batched in the delete phase ❌ piecemeal deletion anytime Avoid concurrent collection modification
Distant Visual data (ISM), not entities ❌ entities for distant points Visible ≠ a person
Persistence Stub identity + persistent state ❌ save volatile full actors The city remembers everyone
Substrate Data-oriented self-build (SOA stubs + budgeted subsystem + BVH + ISM) ❌ Mass mandatory Higher semantic fidelity + control
Brain Rules (default) → LLM (crowd director/personality, optional) — LLM swaps decisions, execution unchanged

Why prefer self-build over adopting Mass directly: Mass (MassEntity + MassCrowd) can save the ECS plumbing and tap MassCrowd/MassTraffic's ready-made capabilities, but at the cost of accommodating its paradigm, and this set of stub-first / single-threaded draining / visibility-priority / budgeted semantics would still have to be self-supplied in large part even with Mass. Referencing the original to self-build a data-oriented population runtime (an SOA stub array + a budgeted UWorldSubsystem + BVH spatial queries + ISM distant points) yields a higher degree of semantic fidelity and control. But "self-build" does not mean "reject Mass wholesale"—adopting Mass locally at certain layers is in fact apt: e.g. the distant crowd, or those stateless lightweight proxies, run their batch updates well with MassEntity's fragment + processor, which is its strength; while the core semantics of "identity/representation decoupling, budgeted attach, stub-first lifecycle" stay in the self-built layer's control. So the more accurate stance is: self-build the core, borrow Mass at the edges as needed, rather than an either-or.

A final word on the place of LLMs, lest it seem abrupt: this article is about the population "runtime"—identity, representation, budget, asynchronous lifecycle; this is the foundation, unrelated to LLMs and not to be replaced by them. If an LLM is introduced in future, it fits better at the content and decision layers (a crowd's personality, a district's atmosphere, a few NPCs' reactions at a moment), rather than taking over the runtime scheduling of "who to spawn, who to promote, who to reclaim."

So what exactly does the LLM handle? Isomorphic to the traffic director: the slow layer uses an LLM offline to generate the crowd's "personality and script" (what routines and reaction tendencies people in different districts have), solidified into community-data parameters; the mid layer promotes only a few NPCs near the player and plot-related into LLM agents for runtime decisions. The three-tier "cost is distance" carries straight over to the LLM—distant people run not even rule-based AI, let alone call an LLM. And the premise of all this is that the runtime foundation has already firmly handled "where people come from, where they go, when they are reclaimed"; the LLM is merely the "brain" growing on this foundation, not the foundation itself.


AI Collaboration Retrospective

  • What the AI did: this article is based on a source-level deep read of the original's population module (stub-first with stable entity IDs, three-tier representation with bidirectional handoff, the population orchestration layer bridging community data/stubs/spawn/lanes/save/events, the request-queue semantics of spawning and the multi-token completion chain, the full-entity attach budget and independent stub-pool budget, crowds sitting on the lane graph with work-points, multi-reason reclamation batched in the delete phase, the stub controller's raw pointers and deletion callback, distant crowds as visual data rather than entities, stub-layer persistence), bringing back each real mechanism and mapping it to UE5.
  • How the human filled in: ① the stance is set by the human—following "prefer referencing the original and self-building, Mass optional not mandatory" (the user had made this clear in the main article's crowd chapter); ② anonymization—the original's internal class names / line numbers are kept only in research notes; the article uses generic engineering terms.
  • How it was verified: every conclusion traces back to the population-module deep read (e.g. "the spawn interface returns an acceptance signal rather than entity completion," "the stub controller holds raw pointers relying on a deletion callback," "distant crowds are visual-data availability" all come from the deep-read dossier).
  • One self-discipline (converged through a round of AI review): this article distinguishes three layers—mechanisms explicitly present in the source (read directly from interfaces/data structures/assertions), models inferred from structure (e.g. "four-baton relay," "the dimensions of the importance score," marked with "can be abstracted as…" / "from the interface it is closer to…"), and my UE5 migration design (the self-built UCyber* part). An elegant model is not written as "a proven runtime mechanism"; concrete constants (counts/durations/distances) are described in the text only as "there is such a valve/threshold," without hard numbers (to avoid, combined, fingerprinting a specific work).
  • This article's boundary: it is a development plan / rebuild blueprint, explaining thoroughly "how crowds are orchestrated and scheduled, and how to reproduce it on UE5"; how these people move (kinematics, animation, crowd presentation) belongs to the pedestrian Effects · Physics · Presentation article.

Leave a Reply

Discover more from AI Native Game Development

Subscribe now to keep reading and get access to the full archive.

Continue reading