Rendered from AUTHORING.md at commit d894fc0. The package is imported here by its 0.1.0 name, typeshade.
fp64
GPUs have no f64; the DSL emulates it as an unevaluated two-f32 pair (hi + lo,
“double-float”/df64 — DSFUN90 → NVIDIA CUDA dsadd/dsmul → Thall lineage), giving
~48 significand bits at f32 exponent range. The authoring surface is IDENTICAL to
f32 — only the declared type differs:
import { f64T, f64, toF64, toF32, splitF64 } from 'typeshade'
const U = uniformStruct( 'U', { group: 0, binding: 0, as: 'u' }, { origin: f64T, // one vec2<f32> slot — host packs splitF64(value) },)
const k = fn( 'k', { x: f64T, s: f32T }, (p) => toF32(sqrt(p.x.add(U.field.origin).mul(p.s))), // .add/.mul/sqrt — unchanged syntax)const m = module({ funcs: [k], uses: [U] }) // nothing fp64-specific to declareThe pre-emit fp64Lower pass rewrites every f64 into vec2<f32> + injected
df64_* emulation calls (WGSL, GLSL, and — natively, as JS numbers — the CPU oracle
all agree). What to know:
- Conversions.
f32 → f64widens implicitly in arithmetic (exact) or explicitly viatoF64(x)/ a bare number literal (split losslessly at build time — JS numbers ARE f64). Narrowing is ONLY explicit:toF32(x)(= hi + lo, precision-losing). Mixing f64 with ints/bools is an author-timeSD0004. - Supported ops.
+ − × ÷, all comparisons (lexicographic),neg,abs,min,max,sqrt,mix(f32 interpolant),floor,fract,sin,cos. Anything else on an f64 operand fails loud at emit (SD0041) — narrow explicitly first.%and bitwise are rejected at author time. - Transcendentals (
sin/cos). A luma.gl port: 3-stage argument reduction (mod 2π → quadrant → π/16 index) + tabled angle-addition + a short Taylor on the tiny residual. Two things to know. (1) Accuracy is lower than the arithmetic: the Taylor truncation floors relative error at ~2⁻³⁶ for the transcendental itself, and it degrades with argument magnitude through the reduction (the inherent large-argument precision loss — still far past f32, whose sine of a ≳2²⁴ argument is pure noise: the f32 argument has already lost the sub-ulp phase). (2) The df64_mul caveat applies. sin/cos are built on the df64 multiply, and on Apple/Metal that multiply collapses under default fast-math in a way that is not robustly guardable in-shader (see thedf64_mulnote incore/fp64/df64-lib.tsand the fp64 blog Part 7/8). So sin/cos are correct on backends where the multiply survives (verified on Blackwell and on a Windows/D3D12 Turing path) but inherit the same fragility on Apple/Metal — the on-device gate asserts this device-conditionally, it does not claim universal correctness. Also onvecN<f64>(per-lane). - The guard texture (auto-injected). Every f64-arithmetic module gets a
texture_2d<f32>binding named_fp64injected automatically (deterministically at group 0, first free binding). The host must bind a 1×1 texture whose texel reads exactly 1.0 (RGBA8 white / R32F 1.0) — WGSL §15.7.5 permits reassociation and Metal defaults to fast-math, and without a runtime-opaqueonethreaded through the error-compensation terms a downstream compiler can legally fold df64 back to f32 precision (the luma.gl CODE_ELIMINATION_WORKAROUND lineage). It is a TEXTURE rather than a uniform because some drivers specialize pipelines on observed uniform values and hot-swap re-optimized variants that fold the terms anyway (seen in the field on Windows/NVIDIA); no compiler treats texel values as constants. To pin the slot to an engine’s fixed bind-group layout, declarefp64Guard({ group, binding })inuses:; a conflicting_fp64declaration isSD0042. - Layout & packing. An f64 uniform field / vertex attribute occupies a plain
vec2<f32>slot (size 8, align 8); pack it withsplitF64(x)(hi, lo). An f64 VARYING is rejected (SD0044) — interpolating hi/lo pairs is numerically wrong. - Names. The
df64_fn prefix andDF64VecNstruct names are reserved for the injected emulation (SD0043). - Cost. Each op is several-to-10× an f32 op — opt in per VALUE, not per shader.
Vectors — vec2f64T / vec3f64T / vec4f64T
vecNf64(x, y, …) builds an emulated-double vector; components (v.x), swizzles
(v.zyx), componentwise + − × (with f64/f32/number broadcast), ÷, neg,
the componentwise builtins abs/min/max/mix (scalar f32 interpolant)/floor/
fract/sin/cos/normalize, and the reductions dot/length/distance (→ f64)
all use the unchanged surface. A vec64 lowers to
struct DF64VecN { hi: vecN<f32>, lo: vecN<f32> } — componentwise arithmetic runs
the same EFTs on whole hi/lo planes (one twoSum for all lanes); the builtins with
per-lane branching (abs, min, …) and normalize compose the verified SCALAR
df64 fns lane by lane inside one df64_vN_* helper body, and
dot/length/distance accumulate through the SCALAR df64 chain
(extended-precision accumulation is the point); sin/cos compose the scalar
df64_sin/df64_cos per lane. Everything else on a vec64 (exp, clamp, …) is
SD0041 — narrow per lane (toF32(v.x)) first.
A vec64 uniform field occupies its struct layout (n=2: 16 B, n=3/4: 32 B under
std140); a vec64 vertex ATTRIBUTE is rejected (SD0041) — pass hi/lo as two
vecN<f32> @locations (the existing DSFUN lane convention) and rebuild lanes with
f64FromParts.
See examples/fp64-deep-zoom.ts for the full picture (f32 collapse vs f64 stripes).