Rule 11.8
Chapter 11, Targets and the oracle Test
A kernel function’s call (Rule 8.21) runs on the first tier of configure({ prefer }) that can run it, WebGPU, WebGL2 and the CPU tier being the default order: WebGPU when the function lowers (Rule 8.22) and there is a device, WebGL2 when every loop writes one array of 4-byte elements (f32, i32, u32) at exactly i and every scalar parameter is a number or a vector, each loop one fragment program into an R32UI target that starts out holding the array, so an iteration that writes nothing leaves its element, and the CPU tier always; a list of one tier makes that tier required, and a call no tier on the list can run throws an Error naming each tier’s reason.
Every tier computes a reduction in the tree order (Rule 7.2).
Kernel calls run one after another in the order they were made, whether or not the caller awaits them.
resident(value) wraps a host value (Rule 8.21) once in a Resident, which holds its own copy: a call on WebGPU uploads it the first time and binds the same buffer at every call after, reading nothing back, and await r.read() waits for the calls made before it and returns a new copy of what it holds; r.write(value) replaces what it holds and r.destroy() releases its buffer, each after the calls made before it, or at once when none waits, a use after destroy() being a TypeError; a call that fails while nobody awaits it leaves its error on each Resident it writes, which read() throws. One Resident must not be two parameters of one call. The same Resident is a buffer of the program runtime too, bound by any draw or dispatch of a runtime on the same device, and configure({ runtime }) puts every call on that runtime’s device (Rule 11.11); a program runtime’s Texture passed where a call takes an image is bound as it is, and a draw that falls to WebGL2 refuses it with a TypeError.
A @compute entry’s call (Rule 8.24) runs in the same order as kernel calls, and on the first tier of configure({ prefer }) that can run it: WebGPU when there is a device, and the CPU tier when the entry reaches no barrier and no texture; WebGL2 has no compute stage and never runs one, and a call no tier on the list can run throws, a TypeError naming what the entry reaches when the CPU tier refused it, an Error naming each tier’s reason otherwise. A storage array binding with no size may be a Resident, bound on WebGPU as its device buffer and read back by nobody, and a call whose written bindings are all resident has a second signature, returning void; a Resident for any other binding, or as two bindings of one call, is a TypeError.
A draw’s tiers are fixed: WebGPU, then WebGL2, then the CPU tier, where a canvas keeps the tier of its first draw; a draw may read a Resident for a storage array with no size, and then runs in the order of kernel and entry calls, seeing what the calls before it wrote, with the handle bound as its device buffer on WebGPU and its copy brought up to date on the CPU tier; a Resident for any other binding is a TypeError.
The rule text and the parts under it are the compiler's own English, as the design document writes them.
Rationale
a chain of calls over the same data (a particle step, a relaxation) should pay one upload and one read, not one of each per call: #252 measured the conversion of a million 32-byte structs at about 220 ms in and 1.3 s out, which residency pays once. The order of the tiers is the caller’s to set, and a caller that needs the GPU says so and gets an error that names why rather than a slow answer (#252’s open question 6). The calls keep their order because a queued call has no await to order it by.
Derives from
change 0013 in changes/ (“Tiers”, “resident(array)”); change 0016 (“Tiers”); docs/dx.md; Rule 11.7.
How it is verified
Checked by a test. A test, a gate script or a CI workflow names this rule, and the traceability check fails when a file listed below stops naming it.
Where the rule says the compiler enforces it:
resident, configure and the call queue in src/core/resident.ts, which callKernel in src/core/host-kernel.ts, callCompute in src/core/host-compute.ts and callDraw in src/core/host-draw.ts read, the WebGL2 programs from lowerKernelGl in src/core/passes/kernel-lower.ts, drawn by src/core/host-kernel-gl.ts, and the host view’s two signatures from hostFace in src/compiler/ts/host-face.ts; pinned by src/compiler/ts/host-kernel.test.ts (the view under tsc, a queued call read back, an error kept for read(), each tier’s reason, one handle as two parameters), by src/compiler/ts/host-entry.test.ts (an entry’s view under tsc, an entry and a kernel function queued on one handle, each refusal, each tier’s reason), by src/compiler/ts/host-draw.test.ts (a draw between two queued kernel calls on one handle) and by the import journey, which renders into and reduces over resident arrays on WebGPU, chains two entry calls through one, draws one a kernel function filled on WebGPU and the CPU tier, and runs two maps with WebGL2 required, in Chromium; the compile gate links each kernel’s WebGL2 program on a WebGL2 context; the handle as a program runtime’s binding, write, destroy and configure({ runtime }) by src/runtime/runtime.test.ts.
The files that verify it at commit 26de7be8, each at the first line that names the rule:
-
src/compiler/ts/host-draw.test.ts:333 -
src/compiler/ts/host-entry.test.ts:177 -
src/compiler/ts/host-face.ts:765 -
src/compiler/ts/host-kernel.test.ts:279 -
src/core/host-compute.ts:37 -
src/core/host-draw.ts:25 -
src/core/host-kernel-gl.ts:1 -
src/core/host-kernel.ts:104 -
src/core/passes/kernel-lower.ts:1062 -
src/core/resident.ts:2 -
src/runtime/runtime.test.ts:5
Explained in
The sections of the surface document that explain this rule, at commit 26de7be8:
See also
Source
The rule at commit 26de7be8:
docs/language-design.md:1170(the design document)reqs/rules/RULE-1108.md(its traceability item)