fp64
On this page

Rendered from AUTHORING.md at commit 26de7be8. The package is imported here by its 0.2.0 name, typeshade.

fp64

After this page you can declare a double-precision value in a shader, know which operations it supports, and bind the guard texture the emulation needs.

A GPU has no 64-bit float type. In this package f64 is emulated: one value is a pair of f32 numbers, a high part and a low part, and the number is their unevaluated sum. That gives about 48 bits of significand, the fraction bits that decide how many digits a value keeps, over the ordinary f32 exponent range. A value near 1e8, a position in world units for one, where one f32 step is already 8 units, still resolves detail far below one unit. The authoring surface is the same as f32, and only the declared type differs. Before emit, the fp64Lower pass rewrites every f64 into a vec2<f32> and injects the emulation functions the shader now calls. WGSL, GLSL and the CPU oracle agree on the results.

The 48 bits assume the GPU rounds f32 +, - and * to nearest, ties to even. Every shipping GPU does, but neither specification promises it: GLSL ES 3.00 leaves the rounding mode undefined and lets a subnormal flush to zero, and WGSL fixes no rounding mode. The figure rests on that hardware practice; §39 of docs/use-typeshade-surface.md has the detail.

Declaring an f64 value

Write f64T where you would write f32T: in a uniform field, in an fn parameter, in a local. The arithmetic methods, the comparisons and the builtins keep their spelling, so a body converted from f32 usually changes only in its signature.

import { f32T, f64T, fn, module, sqrt, toF32, uniformStruct } from 'typeshade'
const U = uniformStruct(
'U',
{ group: 0, binding: 0, as: 'u' },
{
origin: f64T, // one vec2<f32> slot: the host packs splitF64(value)
},
)
const k = fn('k', { x: f64T, s: f32T }, (p) => toF32(sqrt(p.x.add(U.field.origin).mul(p.s))))
const m = module({ funcs: [k], uses: [U] })

Nothing on the module announces fp64. The lowering pass finds the f64 types itself.

Conversions

An f32 widens to f64 implicitly inside arithmetic, and the widen is exact, so x.add(dx) for an f64 x and an f32 dx needs no conversion and emits what x.add(toF64(dx)) emits. toF64(x), and x.f64(), are the explicit spellings of that same widen, for a value that has to be f64 before it meets any arithmetic. A JavaScript number is already a double, so f64(1e-9), or the bare 1e-9 in f64 arithmetic, splits the literal into its pair at build time with nothing lost. f64FromParts(hi, lo) assembles a value from two f32 halves that arrive already split, a vertex attribute pair for instance. Narrowing is always explicit: toF32(x) adds the two halves back together and gives an f32, losing the extra precision. Mixing an f64 with an integer or a boolean is an author-time SD0004.

import { f32T, f64, f64FromParts, fn, toF32, toF64 } from 'typeshade'
const g = fn('g', { a: f32T, hi: f32T, lo: f32T }, (p) => {
const widened = toF64(p.a) // exact
const lit = f64(1e-9) // split at build time
const rebuilt = f64FromParts(p.hi, p.lo) // two lanes that were split by the host
return toF32(widened.add(lit).add(rebuilt)) // the explicit narrow
})

The operations f64 supports

The four arithmetic operations, every comparison, neg, abs, min, max, sqrt, mix with an f32 interpolant, floor, fract, sin and cos accept f64 operands. Anything else on an f64 operand fails at emit with SD0041, so narrow the value with toF32 first and finish that part of the computation in f32. % and the bitwise operators are refused earlier, at author time. sin and cos are less accurate than the arithmetic around them. Relative error for the transcendental itself floors at about 2^-36, and it grows with the size of the argument through the range reduction.

import { f64, f64T, floor, fn, min, toF32 } from 'typeshade'
// A triangle wave on a coordinate that has outgrown f32.
const stripe = fn('stripe', { x: f64T }, (p) => {
const y = p.x.mul(0.5)
const f = y.sub(floor(y))
return toF32(min(f, f64(1).sub(f)))
})

The guard texture

A shader compiler may reassociate float arithmetic, and reassociation deletes exactly the small correction terms this emulation is built on. Every emitted helper therefore threads a value the compiler cannot see through those terms: a 1 read from a texture. Any module whose f64 arithmetic calls one of those helpers gets a texture_2d<f32> binding named _fp64 injected for it, at group 0 and the first free binding, and it appears in reflect() as an ordinary 2D texture. A module that only compares, widens, narrows, negates or scales f64 values by a power of two calls none of them and gets no binding, so a host reads reflect() rather than binding _fp64 to every pipeline. The host binds a 1 by 1 texture whose texel reads exactly 1.0, white RGBA8 or R32F holding 1.0. The value lives in a texture because some drivers specialize a pipeline on the uniform values they observe and re-optimize it, which folds the correction terms away again; no compiler treats a texel as a constant.

Each function that does f64 arithmetic reads the texel once, at the top of its body, and passes it to the helpers as a parameter, so a loop reads it before it starts rather than on every iteration. On GLSL the texture is declared uniform highp sampler2D _fp64;: GLSL ES 3.00 gives sampler2D a default precision of lowp in both stages, and a texel fetch returns its sampler’s precision.

Two things to do at the host. If the bind group layout is fixed, pin the slot with fp64Guard({ group, binding }) in the module’s uses. And on Apple GPUs the guard is not enough, because the platform’s default fast math collapses the float lowering anyway. Pass whatever device signals you have to recommendFp64Flavor and hand the result to the emit as fp64Flavor; on those devices it selects an integer lowering, which needs no guard binding at all.

// `k` and `U` are the function and the uniform struct from the first example.
import { emitModule, fp64Guard, module, recommendFp64Flavor } from 'typeshade'
const pinned = module({ funcs: [k], uses: [U, fp64Guard({ group: 0, binding: 3 })] })
const flavor = recommendFp64Flavor({ userAgent: navigator.userAgent })
const wgsl = emitModule(pinned, { fp64Flavor: flavor })

The emulation owns two name spaces in the emitted shader. A function name starting with df64_, or a struct name starting with DF64Vec or DF64Mat, fails emit with SD0043.

Packing and layout

An f64 uniform field or vertex attribute occupies one plain vec2<f32> slot, size 8 and align 8. Pack it on the host with splitF64(x), which returns the [hi, lo] pair in the order the shader reads it. A value interpolated from the vertex stage to the fragment stage, a varying, cannot be f64: that is SD0044, because interpolating a hi/lo pair across a triangle is numerically wrong. Interpolate an f32 quantity instead, or carry the value in a uniform and do the f64 arithmetic in the fragment stage.

import { splitF64 } from 'typeshade'
const [hi, lo] = splitF64(-8_234_567.890123)
new Float32Array(uniformBuffer, 0, 2).set([hi, lo])

Vectors

vec2f64T, vec3f64T and vec4f64T are the vector types, built with vec2f64(x, y) and its siblings. Components, swizzles, componentwise arithmetic with f64, f32 and number broadcast, the componentwise builtins abs, min, max, mix, floor, fract, sin, cos and normalize, and the reductions dot, length and distance, which give back an f64, all read the same as their f32 counterparts. Anything outside that list is SD0041 again, so narrow one component at a time with toF32(v.x). A vector lowers to a struct holding a hi plane and a lo plane, so componentwise work runs once for the whole vector.

import { dot, f64T, fn, toF32, uniformStruct, vec2f64, vec2f64T } from 'typeshade'
const P = uniformStruct('P', { group: 0, binding: 1, as: 'p' }, { center: vec2f64T })
const d = fn('d', { x: f64T, y: f64T }, (p) => {
const q = vec2f64(p.x, p.y).sub(P.field.center)
return toF32(dot(q, q)) // the subtraction and the dot both run in f64
})

A vector uniform field takes its struct layout: 16 bytes for two components, 32 bytes for three or four under std140. A vector vertex attribute is rejected. Pass hi and lo as two vecN<f32> locations and rebuild each lane with f64FromParts.

What it costs

One f64 operation is several f32 operations, ten or more for a multiply or an add. So opt in per value: hold the coordinates that need the range in f64, narrow as soon as the difference is small enough for f32, and let the rest of the shader run at f32 speed.

import { f64T, fn, toF32 } from 'typeshade'
// The subtraction needs the range; the shading after it does not.
const shade = fn('shade', { world: f64T, camera: f64T }, (p) =>
toF32(p.world.sub(p.camera)).mul(0.5).add(0.5),
)

Two f64 multiplies cost less, with nothing to write differently. x.mul(x) is a square, about 30% cheaper than a general multiply, as long as computing x has no side effect, such as a call that writes a storage binding. A multiply or a divide by a power-of-two literal, 2.0, 0.5 or -4.0, scales the two f32 words and nothing else, which is exact barring overflow and underflow: two f32 operations, none for 1.0, and a sign change for -1.0. A scale that can grow the value (by more than 1) applies to a value computed at run time; a constant keeps the general multiply, so that WGSL, which evaluates constants while it creates the shader and refuses one that overflows, never has to.

One example in the gallery, examples/fp64-deep-zoom.ts, runs one formula on both types side by side, and shows the f32 half collapsing to a flat field while the f64 half keeps its stripes.

Edit this page Report a problem