GPU Accelerated JavaScript
A kernel may also be given as source text. This is the only form available where the engine does not retain function source — React Native's Hermes, for instance, returns "function name(a0, a1) { [bytecode] }" from Function.prototype.toString().
Void
Compile a whole multi-kernel computation into one callable plan. The orchestration function runs once, at build time (first call), with opaque handles for arguments; the kernel calls it makes are recorded and replayed on later calls with intermediates kept resident. Calling the pipeline always returns a Promise.
Void
why this kernel's work was degraded to the cpu backend, when it was (#868)
Void
webasm sits last, one step above the cpu fallback: any working GL backend outranks it, so auto modes only reach it where no GL context exists
Void
The GPU.js library class which manages the GPU context for the creating kernels
GPU
Creates an instance of GPU.
| Name | Type | Description | |
|---|---|---|---|
| settings |
IGPUSettings
|
|
Optional |
Void
mode 'async' only: the adapter probe's settled answer (true/false), or null while it is in flight. Started at construction so that by kernel creation -- usually at least a task later in real applications -- the backend for a graphical kernel can be decided BEFORE its canvas is exposed, since a canvas is permanently committed to its first context type and can never be swapped between backends afterwards.
Void
| Name | Type | Description | |
|---|---|---|---|
| source |
Function
String
object
|
|
|
| settings |
IGPUKernelSettings
|
|
Optional |
IKernelRunShortcut
callable function to run
| Name | Type | Description | |
|---|---|---|---|
| reasons |
Array.<IReason>
|
||
| args |
IArguments
|
||
| _kernel |
Kernel
|
| Name | Type | Description | |
|---|---|---|---|
| fn |
Function
|
|
|
| settings |
IPipelineSettings
|
|
Optional |
IPipelineRunShortcut
callable pipeline
Create a super kernel which executes sub kernels and saves their output to be used with the next sub kernel. This can be useful if we want to save the output on one kernel, and then use it as an input to another kernel. Machine Learning
| Name | Type | Description | |
|---|---|---|---|
| subKernels |
Object
Array
|
|
|
| rootKernel |
Function
|
|
const megaKernel = gpu.createKernelMap({
addResult: function add(a, b) {
return a[this.thread.x] + b[this.thread.x];
},
multiplyResult: function multiply(a, b) {
return a[this.thread.x] * b[this.thread.x];
},
}, function(a, b, c) {
return multiply(add(a, b), c);
});
megaKernel(a, b, c);
Note: You can also define subKernels as an array of functions.
> [add, multiply]
Function
callable kernel function
Combine different kernels into one super Kernel, useful to perform multiple operations inside one kernel without the penalty of data transfer between cpu and gpu.
The number of kernel functions sent to this method can be variable. You can send in one, two, etc.
| Name | Type | Description | |
|---|---|---|---|
| subKernels |
Function
|
|
|
| rootKernel |
Function
|
|
combineKernels(add, multiply, function(a,b,c){
return add(multiply(a,b), c)
})
Function
Callable kernel function
| Name | Type | Description | |
|---|---|---|---|
| source |
Function
String
|
|
|
| settings |
IFunctionSettings
|
Optional |
GPU
returns itself
| Name | Type | Description | |
|---|---|---|---|
| name |
String
|
|
|
| source |
String
|
|
|
| settings |
object
|
Optional |
GPU
returns itself
Inject a string just before translated kernel functions
| Name | Type | Description | |
|---|---|---|---|
| source |
String
|
GPU
Makes kernels easier for mortals (including me)
| Name | Type | Description | |
|---|---|---|---|
| kernel |
function()
Pipeline compilation (docs/design/pipeline-compilation.md): the orchestration function runs ONCE, at build time, against opaque handles; every kernel call made while the trace is open is recorded into a static plan, and later pipeline calls execute the plan without re-entering user code. JS loops in the orchestration therefore unroll at trace time, and closure-captured plain values freeze into the plan the same way constants do.
Void
The class exists for instanceof and for its name in errors; all state lives in the trace's WeakMap so the frozen instance has no own properties for the Proxy get trap to conflict with.
Void
Consulted by kernelRunShortcut on every call; non-null only while an orchestration function is being traced, which is always synchronous, so a module-level slot cannot see two traces at once.
Void
Trace-time state: records kernel calls as plan steps and mints the opaque handles that stand in for values the orchestration never gets to see.
Void
distinct kernel run-shortcuts, in first-use order; steps refer to them by index so the ping-pong loop shape compiles to ONE kernel entry
Void
| Name | Type | Description | |
|---|---|---|---|
| meta |
Object
|
|
Proxy.<PipelineHandle>
Entry point from kernelRunShortcut while a trace is open: validate the kernel, bind the arguments, and answer with a fresh step-output handle instead of running anything.
| Name | Type | Description | |
|---|---|---|---|
| shortcut |
IKernelRunShortcut
|
||
| args |
IArguments
|
Proxy.<PipelineHandle>
Object
argBinding per the plan IR; non-handles snapshot here, which is the moment closure-captured mutables freeze
Call-time sampling: mutable JS values copy before the call promise can
yield, so const p = pipeline(buf); buf[0] = 9; computes on the value buf
held at the call. Handles never reach this function -- bindValue checks
the WeakMap first -- so property access here cannot trip a handle trap.
Void
Static liveness over the unrolled DAG, then greedy slot reuse: a step may
write a buffer only when the previous occupant's last reader ran strictly
earlier -- a reader AT the writing step still needs the old contents while
the new ones are produced, which is exactly what forces u = sweep(u, q)
in a loop onto two alternating buffers. Slots are only shared between
steps of identical output shape so the fused executor can lay them out as
fixed regions.
| Name | Type | Description | |
|---|---|---|---|
| steps |
Array
|
|
|
| resultBindings |
Array
|
Array
buffers
| Name | Type | Description | |
|---|---|---|---|
| trace |
PipelineTrace
|
||
| returned |
|
Object
results descriptor {kind, entries: [{key?, binding}]}
| Name | Type | Description | |
|---|---|---|---|
| gpu |
GPU
|
||
| fn |
Function
|
|
|
| settings |
IPipelineSettings
|
Optional |
Void
executor identity probe for tests and later phases: 'generic' executes step-by-step through the normal kernel machinery on every backend; 'fused-sync' is the webasm executor running every step over one shared wasm memory; 'fused-threaded' is that executor with pool workers walking the whole plan on an Atomics barrier; 'fused-encoder' is the webgpu executor recording every step into one command encoder over persistent storage buffers
Void
why the fused executor declined this plan; null while fused (or before the first call decides)
Void
undefined: not yet attempted for this plan; false: attempted and declined (generic runs); otherwise the compiled fused executor
Void
concurrent calls to one pipeline serialize on this tail, the same contract as threaded webasm kernels
Void
The threaded executor rejects asynchronously (worker death, stalled barrier, destroy mid-run); any such failure leaves its barrier state unusable, so the executor is dropped and the next call compiles a fresh one. Fallback decisions stay synchronous — the signature check throws before dispatch — so a FusionFallback can never surface here.
| Name | Type | Description | |
|---|---|---|---|
| result |
|
Void
Runs the orchestration function once with handles for arguments; the recorded steps become the plan. Math.random is barred for the duration because a trace-time draw would freeze into every later call.
Object
plan IR
The generic executor's writer for one (kernel, seat signature, output slot) triple. One MUTABLE, STATICALLY-TYPED clone per triple reproduces the hand-rolled two-kernel ping-pong mechanically: each clone owns one output texture/array for the life of the plan (steady state allocates nothing per step) and sees one argument-type signature (no per-call dynamicArguments re-typing -- the forced re-typing was most of a 7x loss even after the texture churn was gone). Static liveness (assignBuffers) is what makes mutability safe: no step ever reads a slot while that slot's writer renders. Argument drift across CALLS is the clones' own switch machinery's business, as for any kernel.
| Name | Type | Description | |
|---|---|---|---|
| plan |
Object
|
||
| step |
Object
|
IKernelRunShortcut
Attempts the backend's fused executor for the current plan against this call's sampled arguments: the single-encoder lowering on webgpu (async — kernel builds await the device), the wasm-memory lowering everywhere else. Anything the fused compile cannot take degrades to the generic executor with the reason recorded, its usual degradation contract.
| Name | Type | Description | |
|---|---|---|---|
| args |
Array
|
|
Promise
The plan runs on private instances configured for pipeline use --
pipeline: true, immutable: true -- so intermediates stay resident
(textures on GL, fresh arrays on cpu) and the user's kernel settings
are never observably touched. Kernels stay shared between pipelines and
direct use through their own shortcuts.
| Name | Type | Description | |
|---|---|---|---|
| shortcut |
IKernelRunShortcut
|
|
IKernelRunShortcut
private clone
Array pipeline arguments upload ONCE per call on backends where an upload costs (GL textures, webgpu buffers): a lazy per-arg identity kernel parks the value device-side and every consuming step binds the handle -- feeding the raw array to a 200-step plan re-uploaded it 200 times, which was most of the remaining gap to hand-rolled ping-pong. cpu/webasm consume arrays natively, so there the raw value is optimal.
Void
Eager uploads are only sound where the upload call is SYNCHRONOUS (the GL family): the texture materializes before user code can run again. webgpu uploads return promises, so its generic path keeps copies.
Void
String
'LE' or 'BE' depending on system architecture Credit: https://gist.github.com/TooTallNate/4750953
| Name | Type | Description | |
|---|---|---|---|
| funcObj |
Function
|
|
Boolean
TRUE if the object is a JS function
| Name | Type | Description | |
|---|---|---|---|
| fn |
String
|
|
Boolean
TRUE if the string passes basic validation
| Name | Type | Description | |
|---|---|---|---|
| funcStr |
String
|
|
String
Function name string (if found)
| Name | Type | Description | |
|---|---|---|---|
| fn |
String
|
|
Array.<String>
Array representing all the parameter names
| Name | Type | Description | |
|---|---|---|---|
| obj |
Object
|
|
Object
Array
Cloned object
| Name | Type | Description | |
|---|---|---|---|
| array |
Object
|
|
Boolean
true if is array or Array-like object
| Name | Type | Description | |
|---|---|---|---|
| length |
Number
|
TextureDimensions
A texture takes up four
| Name | Type | Description | |
|---|---|---|---|
| dimensions |
OutputDimensions
|
||
| bitRatio |
Number
|
TextureDimensions
| Name | Type | Description | |
|---|---|---|---|
| dimensions | |||
| bitRatio |
TextureDimensions
| Name | Type | Description | |
|---|---|---|---|
| x |
Array
String
Texture
Input
|
|
|
| pad |
Boolean
|
|
Optional |
OutputDimensions
Puts a nested 2d array into a one-dimensional target array
| Name | Type | Description | |
|---|---|---|---|
| array |
Array
|
||
| target |
Float32Array
Float64Array
|
Void
Puts a nested 3d array into a one-dimensional target array
| Name | Type | Description | |
|---|---|---|---|
| array |
Array
|
||
| target |
Float32Array
Float64Array
|
Void
Puts a nested 4d array into a one-dimensional target array
| Name | Type | Description | |
|---|---|---|---|
| array |
Array
|
||
| target |
Float32Array
Float64Array
|
Void
Puts a nested 1d, 2d, or 3d array into a one-dimensional target array
| Name | Type | Description | |
|---|---|---|---|
| array |
Float32Array
Uint16Array
Uint8Array
|
||
| target |
Float32Array
|
Void
| Name | Type | Description | |
|---|---|---|---|
| array |
Array.<Number>
|
|
|
| part |
Number
|
|
Array.<Number>
An array of smaller chunks
A number as a GLSL float literal. Integer-valued numbers at 1e21 and beyond stringify in exponential form, which is already a valid GLSL float literal — appending .0 to it is not (#864).
Void
| Name | Type | Description | |
|---|---|---|---|
| lines |
Array
|
|
String
Single combined String, separated by \n
| Name | Type | Description | |
|---|---|---|---|
| source |
String
|
||
| settings |
Object
|
String
A visual debug utility
| Name | Type | Description | |
|---|---|---|---|
| gpu |
GPU
|
||
| rgba | |||
| width | |||
| height |
Array.<Object>
| Name | Type | Description | |
|---|---|---|---|
| kernel |
Kernel
|
||
| FunctionNode |
FunctionNode
|
||
| extraNodeOptions |
object
|
Optional |
FunctionBuilder
| Name | Type | Description | |
|---|---|---|---|
| settings |
IFunctionBuilderSettings
|
Optional |
Void
| Name | Type | Description | |
|---|---|---|---|
| functionNode |
FunctionNode
|
|
Void
| Name | Type | Description | |
|---|---|---|---|
| functionName |
String
|
|
|
| retList |
Array.<String>
|
|
Optional |
Array.<String>
Returning list of function names that is traced. Including itself.
https://github.com/gpujs/gpu.js/issues/207 if dependent function is already in the list, because a function depends on it, and because it has already been traced, we know that we must move the dependent function to the end of the the retList.
Void
https://github.com/gpujs/gpu.js/issues/207 if dependent function is already in the list, because a function depends on it, and because it has already been traced, we know that we must move the dependent function to the end of the the retList.
Void
| Name | Type | Description | |
|---|---|---|---|
| functionName |
String
|
|
String
The full string, of all the various functions. Trace optimized if functionName given
| Name | Type | Description | |
|---|---|---|---|
| functionName |
String
|
|
Optional |
Array
The full string, of all the various functions. Trace optimized if functionName given
| Name | Type | Description | |
|---|---|---|---|
| functionList |
Array.<String>
|
|
String
The string, of all the various functions. Trace optimized if functionName given
| Name | Type | Description | |
|---|---|---|---|
| functionList |
Array.<String>
|
|
Array
Prototypes of all functions converted
| Name | Type | Description | |
|---|---|---|---|
| functionName |
String
|
|
String
settings - The string, of all the various functions. Trace optimized if functionName given
| Name | Type | Description | |
|---|---|---|---|
| functionName |
String
|
FunctionNode
| Name | Type | Description | |
|---|---|---|---|
| functionName |
string
|
||
| argumentName |
string
|
number
| Name | Type | Description | |
|---|---|---|---|
| functionName |
string
|
||
| argumentName |
string
|
||
| calleeFunctionName |
string
|
||
| argumentIndex |
number
|
number
| Name | Type | Description | |
|---|---|---|---|
| source |
string
object
|
||
| settings |
IFunctionSettings
|
Optional |
Void
| Name | Type | Description | |
|---|---|---|---|
| name |
String
|
boolean
| Name | Type | Description | |
|---|---|---|---|
| ast |
Object
|
|
String
the function namespace call, unrolled
Whether this backend needs for (a, b; ...) inits hoisted to statements
before the loop -- WGSL cannot express the comma; GLSL and JS take it
natively.
Boolean
| Name | Type | Description | |
|---|---|---|---|
| ast |
Object
|
|
String
Type of the parameter
Generally used to lookup the value type returned from a member expressions
| Name | Type | Description | |
|---|---|---|---|
| type |
String
|
String
Recursively looks up type for ast expression until it's found
| Name | Type | Description | |
|---|---|---|---|
| ast |
String
| Name | Type | Description | |
|---|---|---|---|
| ast | |||
| dependencies | |||
| isNotSafe |
Array
| Name | Type | Description | |
|---|---|---|---|
| ast |
Object
|
|
|
| retArr |
Array
|
|
Array
the parsed string array
| Name | Type | Description | |
|---|---|---|---|
| error |
string
|
|
|
| ast |
Object
|
|
Void
| Name | Type | Description | |
|---|---|---|---|
| ast |
Object
|
||
| retArr |
Array.<String>
|
Array.<String>
| Name | Type | Description | |
|---|---|---|---|
| ast |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| esNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| eNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| brNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| crNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| iVarDecNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| uNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| uNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| logNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| ast |
IFunctionNodeMemberExpressionDetails
De-minification: minifiers (esbuild, terser) fold statements into
expressions -- if (c) { x = 1; } becomes c && (x = 1), statement
sequences become comma expressions, if/else becomes a ternary of
assignments. In statement position the folded expression's VALUE is
discarded, so unfolding back into statements is always
semantics-preserving, no side-effect analysis required. Runs on the parsed
AST before FunctionTracer records anything, so every backend sees plain
statements; on webgl it also runs before the FXC hoisting normalization,
which only understands statement shapes.
Synthetic nodes are stamped with unique start/end: astKey and the literal-type cache are keyed by position. The 0x20000000 base is disjoint from real acorn offsets and from the 0x40000000 base the webgl hoisting machinery stamps its own synthetic nodes with.
Void
A single-statement position (an unbraced loop body or if branch) that unfolds into several statements needs a block around them.
Void
Loop simplification for minified for-headers. for (i = 0, j = 0; test; i++, j++) cannot be expressed on every backend (WGSL takes one statement
per clause), so a comma INIT hoists to statements before the loop and a
comma UPDATE moves to the end of the body -- with a copy ahead of every
continue that belongs to this loop, preserving per-iteration timing.
Returns the statements to place before the loop. A labeled continue makes
the update rewrite unsafe, so such a loop is left exactly as written.
Void
A deep copy with fresh synthetic positions on every node: the same update lands both at the body's end and ahead of each continue, and position-keyed caches must see distinct nodes.
Void
Puts a copy of prefix ahead of every continue belonging to this loop.
Nested loops keep their own continues. Returns the rewritten block, or
null when a labeled continue makes the rewrite unsafe.
Void
Recursively scans AST for declarations and functions, and add them to their respective context
| Name | Type | Description | |
|---|---|---|---|
| ast |
Void
| Name | Type | Description | |
|---|---|---|---|
| source |
string
IKernelJSON
|
||
| settings |
Void
Supplied by GPU.createKernel; swaps in a kernel compiled for the arguments this one was handed. Declared here rather than on the GL kernel so mergeSettings carries it onto every backend -- the cpu kernel needs it for the same argument-type changes.
Void
Maximum loops when using argument values to prevent infinity
Void
Makes the kernel return a Promise of its result on every backend. Backends with a genuinely non-blocking readback (webgl2 fences, webgpu natively) use it; the rest resolve their synchronous result, so the calling contract is uniform either way.
Void
Make GPU use single precision or unsigned. Acceptable values: 'single' or 'unsigned'
Void
Seed for Math.random() so kernel runs are reproducible; null seeds from Math.random()
Void
Reasons this kernel cannot serve the call it was handed, collected for the caller's switch; null when it can.
Void
| Name | Type | Description | |
|---|---|---|---|
| settings |
IDirectKernelSettings
IJSONSettings
|
Void
Float32Array
Array.<Float32Array>
Array.<Array.<Float32Array>>
Result The final output of the program, as float, and as Textures for reuse.
| Name | Type | Description | |
|---|---|---|---|
| settings |
IDirectKernelSettings
|
{string[]};
| Name | Type | Description | |
|---|---|---|---|
| source |
KernelFunction
string
IGPUFunction
|
||
| settings |
IFunctionSettings
|
Optional |
Kernel
| Name | Type | Description | |
|---|---|---|---|
| name |
string
|
||
| source |
string
|
||
| settings |
IGPUFunctionSettings
|
Optional |
Void
| Name | Type | Description | |
|---|---|---|---|
| args |
IArguments
|
|
Void
| Name | Type | Description | |
|---|---|---|---|
| output |
Array
Object
|
Array.<number>
| Name | Type | Description | |
|---|---|---|---|
| output |
Array
Object
|
|
this
| Name | Type | Description | |
|---|---|---|---|
| flag |
Boolean
|
|
this
| Name | Type | Description | |
|---|---|---|---|
| max |
number
|
|
this
| Name | Type | Description | |
|---|---|---|---|
| constantTypes |
IKernelValueTypes
|
this
| Name | Type | Description | |
|---|---|---|---|
| functions |
Array.<IFunction>
Array.<KernelFunction>
|
this
| Name | Type | Description | |
|---|---|---|---|
| nativeFunctions |
Array.<IGPUNativeFunction>
|
this
| Name | Type | Description | |
|---|---|---|---|
| injectedNative |
String
|
this
Set writing to texture on/off
| Name | Type | Description | |
|---|---|---|---|
| flag |
this
Set Promise-returning mode on/off
| Name | Type | Description | |
|---|---|---|---|
| flag |
Boolean
|
this
Set precision to 'unsigned' or 'single'
| Name | Type | Description | |
|---|---|---|---|
| flag |
String
|
'unsigned' or 'single' |
this
Set a seed for Math.random(), so kernel runs are reproducible
| Name | Type | Description | |
|---|---|---|---|
| seed |
Number
|
this
| Name | Type | Description | |
|---|---|---|---|
| context |
WebGLRenderingContext
|
|
Void
| Name | Type | Description | |
|---|---|---|---|
| argumentTypes |
IKernelValueTypes
Array.<GPUVariableType>
|
this
| Name | Type | Description | |
|---|---|---|---|
| args |
IArguments
|
||
| reason |
String
|
|
Optional |
Void
| Name | Type | Description | |
|---|---|---|---|
| subKernel |
ISubKernel
|
|
Void
| Name | Type | Description | |
|---|---|---|---|
| removeCanvasReferences |
Boolean
|
remove any associated canvas references |
Optional |
Void
bit storage ratio of source to target 'buffer', i.e. if 8bit array -> 32bit tex = 4
| Name | Type | Description | |
|---|---|---|---|
| value |
number
| Name | Type | Description | |
|---|---|---|---|
| flip |
Boolean
|
Optional |
Uint8ClampedArray
| Name | Type | Description | |
|---|---|---|---|
| kernel |
Kernel
|
||
| args |
IArguments
|
GPUVariableType[]
| Name | Type | Description | |
|---|---|---|---|
| kernel |
Kernel
|
||
| argumentTypes |
Array.<GPUVariableType>
|
Void
| Name | Type | Description | |
|---|---|---|---|
| source |
String
Function
|
||
| settings |
IFunctionSettings
|
Optional |
IGPUFunction
| Name | Type | Description | |
|---|---|---|---|
| value |
KernelVariable
|
||
| settings |
IKernelValueSettings
|
Void
mulberry32, a fast counter-based PRNG with good distribution
| Name | Type | Description | |
|---|---|---|---|
| seed |
Number
|
Function
| Name | Type | Description | |
|---|---|---|---|
| gl |
WebGLRenderingContext
|
||
| options |
IGLWiretapOptions
|
Optional |
GLWiretapProxy
| Name | Type | Description | |
|---|---|---|---|
| extension | |||
| options |
IGLExtensionWiretapOptions
|
| Name | Type | Description | |
|---|---|---|---|
| ast |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| ast |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| ast |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| idtNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| forNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the parsed webgl string
| Name | Type | Description | |
|---|---|---|---|
| whileNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the parsed javascript string
| Name | Type | Description | |
|---|---|---|---|
| doWhileNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the parsed webgl string
| Name | Type | Description | |
|---|---|---|---|
| assNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| bNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| varDecNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| ifNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| tNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| mNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| ast |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| arrNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| Kernel |
GLKernel
|
||
| args |
Array.<KernelVariable>
|
||
| originKernel |
Kernel
|
||
| setupContextString |
string
|
Optional | |
| destroyContextString |
string
|
Optional |
string
| Name | Type | Description | |
|---|---|---|---|
| argument |
KernelVariable
|
||
| kernelValues |
Array.<KernelValue>
|
||
| values |
Array.<KernelVariable>
|
||
| context | |||
| uploadedValues |
Array.<KernelVariable>
|
string
| Name | Type | Description | |
|---|---|---|---|
| fix |
Boolean
|
|
Void
| Name | Type | Description | |
|---|---|---|---|
| flag |
String
|
|
Void
| Name | Type | Description | |
|---|---|---|---|
| flag |
Boolean
|
|
Void
A highly readable very forgiving micro-parser for a glsl function that gets argument types
| Name | Type | Description | |
|---|---|---|---|
| source |
String
|
[object Object]
Picks a render strategy for the now finally parsed kernel
| Name | Type | Description | |
|---|---|---|---|
| args |
KernelOutput
| Name | Type | Description | |
|---|---|---|---|
| flip |
Boolean
|
Optional |
Uint8ClampedArray
Promise.<Uint8ClampedArray>
a Promise under the async contract, so await kernel.getPixels() is portable across
every backend including webgpu, where the readback is genuinely async
| Name | Type | Description | |
|---|---|---|---|
| kernelValue |
WebGLKernelValue
|
||
| arg |
GLTexture
|
Void
A ternary with an integer consequent but a float alternate promotes to float (the WGSL node's rule); the type system must agree with what exprConditional emits or the enclosing expression converts wrongly.
Void
FunctionBuilder drives tracing through toString(); for this backend that is the analysis pass — no text exists, the return value is always ''.
Void
Bytecode pass. assembler carries the module builder, the kernel's
baked memory layout and the shared global indices; it changes per size
signature, so this may run repeatedly on one node.
| Name | Type | Description | |
|---|---|---|---|
| assembler |
Object
|
Void
Converts the wasm value on the stack top between categories. bool is an i32 constrained to 0/1, so bool→i32 is free and i32→bool renormalizes.
Void
Emits ast guaranteed to leave want on the stack, choosing the cast
path by the type system's verdict, like the WGSL node's per-case
castValue/castLiteral dispatches.
Void
JS truthiness for conditions: comparisons pass through, numbers test against zero. Guards short-circuit code from ever seeing a raw f32.
Void
Array(n) locals live as n consecutive f32 locals — wasm has no aggregate values outside memory, and these never escape the function.
Void
The root kernel stores into the output region at data_index (a shared global the run loop advances) and returns — the wasm-level return makes JS early returns exact at any nesting depth.
Void
The WGSL node's exact safe-loop criteria: literal-init single declarator, safe test and init, both test and update present. Shared with the SIMD walk so the LOOP_MAX cap fires identically on both paths.
Void
The switch lowers to an if/else chain on a discriminant local, exactly like the WGSL node: a case-terminating break is consumed, empty cases fall through by OR-ing their tests into the next case, a non-final default moves to the chain's end.
Void
Fallthrough-empty-case grouping shared between the scalar and SIMD lowering: empty cases OR their tests into the next non-empty case, a non-final default moves to the chain's end.
Void
Emits ast, leaving exactly one value on the stack; returns its wasm
category: 'f32' | 'i32' | 'bool' (i32 0/1) | 'void'.
Void
Short-circuit is load-bearing, not an optimization: the right side may
guard an out-of-range memory read (x > 0 && a[x - 1] > 0).
Void
Math.* calls compute in f32; native wasm opcodes where they exist, imports only for what the module actually uses (the kernel scans usedMathImports after analysis). All arguments go through the float ladder, matching the WGSL node's math-call casting.
Void
Flat row-major load, index = x + sizeX * (y + sizeY * z) — the same formula as the GL path's get32 and web-gpu's get_user_X, with missing y/z as zero. Dims and the region offset are baked; the kernel rebuilds per size signature.
Void
Clamps the i32 index on the stack top into [0, max]. i32 has no native min/max, so two selects.
Void
Bytecode pass for kernel_simd. Root kernel only; may run once per size
signature like emitFunction.
Void
Emission-time variance analysis, run once per node and cached. Fixpoint over the local set: a local is varying when it is ever assigned a lane-varying value OR assigned at all under lane-varying control (divergent branch, varying-trip loop, ternary branch, short-circuit right side). A loop is varying-controlled when its test is varying or a break/continue reaches it from under a varying condition.
Void
Rebuilds vCur from a saved base after a construct. Retire masks are
monotone accumulators within their scope, so base & ~each is exact at
any later point; break/continue always target the innermost loop and the
analysis forces any loop with a masked exit to be a varying loop, so
only the top varying entry's masks apply.
Void
Stores the stack top into a v128 local; under a live branch mask the inactive lanes keep their previous value. bitselect copies exact bit patterns, so predication never perturbs IEEE results.
Void
Leaves an i32x4 lane mask (all-ones/all-zeros) for ast as a condition;
lane truth matches the scalar emitCondition exactly.
Void
A break/continue/return retires every lane that reached it, so the rest of its block is dead on both paths; the walk stops emitting there, and termination never leaks past the construct boundary.
Void
Stores the quad's output. componentCount 1 is 4 consecutive f32 — one v128 store (load+blend+store when a mask is live). componentCount n is lane-strided, so components store scalar with a per-lane select.
Void
Emits ast in the vector walk. Uniform expressions go through the
scalar walk untouched (one shared value; splatted only where a varying
context needs it) and return scalar categories; varying expressions
return 'vf32' | 'vi32' | 'vbool' (i32x4 lane mask) | 'void'.
Void
i32x4 shifts take ONE scalar count for all lanes; a lane-varying count lane-scalarizes through the scalar opcode (same mod-32 masking).
Void
Predicated logic evaluates BOTH operands as masks (per-lane skipping cannot exist); side effects in the right operand still predicate correctly because vCur narrows to the left verdict while it runs, and the gather clamps keep formerly short-circuit-guarded reads from trapping.
Void
Helpers stay scalar; a varying call lane-scalarizes: per lane, set that lane's thread.x and PCG state, extract the lane's arguments, call, and rebuild the result vector. Same function bodies and imports as the scalar path, so every lane is bit-identical to its scalar run.
Void
Lane-varying gather: flat row-major index in i32x4 (the scalar formula lane-wise), clamped into the region so lanes a divergent branch turned off cannot trap, then 4 scalar loads + lane inserts — v128 has no gather. In-bounds lanes are untouched by the clamp.
Void
Conservative thread-dependence for the SIMD phase: thread.x is the lane axis (thread.y/z are uniform across an x-row), Math.random is per-cell, user helper calls may read thread state internally, tainted locals propagate in walk order and never clear.
Void
SIMD quads must not cross an x-row (thread.y/z are uniform per quad): rows a multiple of 4 wide vectorize in one span, otherwise each row gets a vector span plus a scalar epilogue for its remainder cells. Shared with the pipeline executor, which drives per-step instances directly.
String
the path taken, for _lastRunPath
The analysis pass: FunctionBuilder's trace runs each node's toString(), which for this backend resolves types and collects math-import/random usage without emitting a byte. Returns false for a return type this backend cannot store, so build() can degrade to cpu.
Void
[ args | constants | output ], each record 16-byte aligned, flat f32
(scalars are one 4-byte slot, Integer/Boolean viewed as i32). Dims come
from the actual argument values, so the layout is per size signature.
Void
Threaded only under the async contract, only when threads exist, and only when the output is big enough (4096 cells) that splitting beats the postMessage round trip.
Void
Sharedness is part of the cache key: a wasm memory import declares shared or not at compile time, so the same size signature needs a distinct module when the async contract routes it to the pool.
Void
The bytecode pass plus the run(start, end, seed) driver. The driver derives thread ids from the flat cell index with baked output dims (x fastest: x + sizeX * (y + sizeY * z), the storage order every backend shares) and seeds the PCG state per cell.
Void
run_simd(start, end, seed): 4 consecutive x cells per step. The caller guarantees (end - start) % 4 == 0 AND that no quad crosses an x-row, so thread.y/z are uniform per quad and thread.x is base + [0,1,2,3].
Void
The scalar pcg_random lane-wise on an i32x4 state. Uniform-count shifts vectorize; the RXS shift count is per-lane, so that one step runs through the scalar opcodes per lane. The mask parameter predicates the state advance: a draw evaluated for an inactive lane (the untaken side of a divergent branch) must not advance that lane's stream, or every later draw in reconverged code desynchronizes from the scalar run.
Void
PCG (permuted congruential, RXS-M-XS output) — the web-gpu kernel's pcg_random verbatim in i32 ops: bit-exact across platforms, top 24 bits scale into [0, 1) at full f32 mantissa resolution.
Void
Frees everything an entry pins. The Memory itself has no explicit free, but dropping every reference (including the workers' — their instantiations hold the shared buffer) is the most a library can do to let it die young (#870). A shared entry defers until the threaded tail settles so an in-flight dispatch keeps what it captured.
Void
The pool path, only ever reached under the async contract. Two timing
constraints shape it: arguments must be sampled at CALL time (the sync
contract's semantics), but the shared args region may still be feeding
an in-flight run — so arguments flatten into staging copies now and are
copied into wasm memory only when this run's turn on the memory comes
up (_threadedTail). The seed is also drawn now so an unseeded kernel
reseeds per call, not per settlement order.
The split follows the contract: min(pool, ceil(cells/4096)) contiguous ranges, each start aligned down to a multiple of 4 so every worker can enter run_simd; the last worker absorbs the tail.
Void
The output region is tightly packed, so scalar returns reuse the memory-optimized erectors; Array(n) returns are stride n, shaped locally — the web-gpu kernel's exact conventions.
Void
Unsigned LEB128. Values are coerced through >>> so i32 bit patterns passed as negative JS numbers encode as their u32 counterpart.
Void
Signed LEB128 for i32 immediates. The |0 coercion pins the value into i32 range so constants supplied as u32 bit patterns (0x9E3779B9-style hash multipliers) encode to the same 32 bits.
Void
Fixed-width 5-byte ULEB128, written into an existing buffer. Used for call-target patch slots whose value is unknown when the body is emitted.
Void
Minimal UTF-8 encoder so the builder stays dependency-free in both Node and the browser bundle (no Buffer, no assumed TextEncoder).
Void
Target is a NAME (import or defined function); the index is patched in at toBytes() so declaration order never matters.
Void
| Name | Type | Description | |
|---|---|---|---|
| lanes |
ArrayLike.<number>
|
16 bytes, little-endian lane order |
Void
One imported memory (env.memory) so a single compiled module can bind either a plain or a shared WebAssembly.Memory. Shared memories require a maximum by spec.
Void
| Name | Type | Description | |
|---|---|---|---|
| name |
String
|
||
| signature |
Object
|
||
| signature.params |
Array.<String>
|
Optional | |
| signature.results |
Array.<String>
|
Optional | |
| signature.locals |
Array.<String>
|
Optional |
WasmFunctionEmitter
body emitter; more locals via addLocal()
The worker body. A template string with NO closure captures: it must
survive being evaluated from a blob URL (browser) or eval: true (Node),
where nothing from this module's scope exists. Everything a task needs
arrives by message: the compiled WebAssembly.Module and the shared
WebAssembly.Memory structured-clone once per (worker, entry), then
{start, end, seed} per task — results are written straight into the
shared memory, so acks carry no data.
The run dispatch mirrors the kernel's sync path: quads must not cross an x-row, so run_simd is used only when the row width is a multiple of 4 (a 4-aligned range start then lands every quad inside one row); other shapes take the scalar export, which is bit-identical by the SIMD contract.
Pipeline entries ('pipelineSetup'/'pipelineRun') execute a WHOLE fused plan per task: every step module is instantiated over the plan's shared memory once at setup, then one run message walks all steps with an Atomics barrier between them — a generation counter in the shared memory, so step boundaries cost no postMessage round trip. Waits are sliced to 100ms so a barrier that can never fill (a peer died) is escapable: the main thread sets the abort word and notifies the generation word, and every check of either releases the worker to ack and go idle.
Void
Event-loop handle accounting, per worker: ref'd while it has a setup OR a task in flight, unref'd when idle. The setup window matters -- a pending Promise does not hold Node's event loop, so a pool unref'd during module setup lets the process exit silently mid-dispatch. Browser workers have no ref/unref and need none.
Void
One setup message per (worker, entry) — concurrent tasks for the same entry share the in-flight ready wait rather than re-sending the module. Kernel entries and pipeline entries share this bookkeeping (a pool is owned by exactly one kernel or one pipeline executor, and pipeline ids are string-prefixed, so the id spaces cannot collide); only the setup message shape differs.
Void
| Name | Type | Description | |
|---|---|---|---|
| entry |
Object
|
kernel module entry: {id, module, memory, mathImports, sizeX} |
|
| tasks |
Array
|
contiguous {start, end, seed} ranges, one per worker |
Promise.<undefined>
resolves when every range has been computed into the entry's shared memory
One task per worker for a WHOLE fused plan: the barrier between steps lives in the entry's shared memory, so this is the only postMessage round trip a pipeline call makes. Every worker in [0, workerCount) must receive its task — the barrier fills only at workerCount arrivals — and a worker that dies rejects its task through the pool's usual machinery, which is the caller's signal to set the entry's abort word.
| Name | Type | Description | |
|---|---|---|---|
| entry |
Object
|
pipeline entry: {id, pipeline, memory, modules, moduleMathImports, steps, countIndex, genIndex, abortIndex, workerCount, workerRanges} |
|
| run |
Object
|
per-call inputs: {baseGen, seeds} |
Promise.<undefined>
resolves when every worker has acked its walk of the plan
Drops an entry's instantiation from every live worker: the worker-side instances are what keep an evicted entry's shared memory alive (#870). The caller guarantees no task for this entry is still in flight.
Void
Fused pipeline execution (docs/design/pipeline-compilation.md): every plan
step compiles to a wasm module over ONE shared memory laid out
[ barrier control | pipeline args | literals | constants | plan buffers ]
(the control words exist only on the threaded path), with each module's
input/output offsets baked against that layout. Per call there is one
flattenTo per pipeline argument and one readback per result, however many
steps the plan unrolls to.
Sync path ('fused-sync'): steps run back-to-back on the calling thread.
Threaded path ('fused-threaded'): when wasm threads exist and the plan is big enough, pool workers execute the WHOLE plan — each worker owns a contiguous cell-range slice of every step and advances step-to-step on an Atomics barrier (generation counter in the shared memory). The main thread dispatches once per call and then waits only for the final generation, so step boundaries cost no main-thread round trip. Threads unavailable or the plan too small falls back to the sync path; anything the backend cannot take at all degrades to the generic executor as usual.
Void
The degradation signal, per the backend's usual contract: the pipeline
catches it and runs the generic executor with this reason. recompilable
marks argument size/type drift a fresh fused compile can absorb.
Void
| Name | Type | Description | |
|---|---|---|---|
| pipeline |
Pipeline
|
||
| plan |
Object
|
|
|
| args |
Array
|
|
WebAssemblyPipelineExecutor
kernels created for second and later type signatures of one plan kernel (the plan clone carries the first); destroyed with the executor
Void
A program is a plan kernel prepared for one argument-type signature: type inference and bytecode translation ran, but no module was instantiated — modules are per step-offset assignment, built below.
Void
Stand-ins with the exact types and dims each binding will have at run time, for setupArguments/computeLayout: sampled values stand for themselves, a step output becomes an Input over its producer's dims.
Void
The analysis half of WebAssemblyKernel.build() without instantiation: modules are assembled against the shared layout instead. Inference is reset first — on a fused recompile the same kernel must re-infer for the new signature, not keep the old one.
Void
The layout baked argument sizes and scalar types; a call that drifts from them throws recompilable so the pipeline compiles a fresh fused plan for the new signature, the way the kernel itself re-instantiates per size signature.
Void
| Name | Type | Description | |
|---|---|---|---|
| args |
Array
|
|
results shaped per the plan; synchronous on the sync path (the pipeline's tail promise provides the async contract), a Promise on the threaded path
One pool dispatch for the whole plan; the workers walk every step over the already-written args and meet at the memory-resident barrier, so the only thing left to await here is the final generation. The pipeline tail serializes calls, which is what makes resetting the generation counter safe: no worker touches the control words between a run's final barrier and its next run message.
Void
Resolves when the generation counter reaches target, rejects on abort
or when the counter stalls past sanityTimeoutMs. Atomics.waitAsync
where the host has it (woken by the workers' notify and by _abort),
short-slice polling otherwise — either way the main thread never blocks.
Void
Releases every wait on the run: workers poll the abort word at each barrier (and inside their sliced Atomics.wait), the main thread checks it on every generation wake. First cause wins; the executor is dead afterwards — the barrier count is indeterminate.
Void
Entry point for Pipeline.destroy() while a run may be in flight: the sync path cannot be mid-run (it never yields), so only the threaded path has anything to interrupt.
Void
The one readback: slice copies results out of wasm memory only here. On the threaded path the barrier's final generation happened-before this read, so the workers' stores are visible.
Void
| Name | Type | Description | |
|---|---|---|---|
| ast |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| ast |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| ast |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| ast |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| ast |
Object
|
||
| retArr |
Array
|
Array.<String>
| Name | Type | Description | |
|---|---|---|---|
| ast |
Object
|
||
| retArr |
Array
|
Array.<String>
| Name | Type | Description | |
|---|---|---|---|
| ast |
Object
|
||
| retArr |
Array
|
Array.<String>
| Name | Type | Description | |
|---|---|---|---|
| ast |
Object
|
||
| retArr |
Array
|
Array.<String>
| Name | Type | Description | |
|---|---|---|---|
| idtNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| forNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the parsed webgl string
| Name | Type | Description | |
|---|---|---|---|
| whileNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the parsed webgl string
| Name | Type | Description | |
|---|---|---|---|
| doWhileNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the parsed webgl string
| Name | Type | Description | |
|---|---|---|---|
| assNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| bNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| ast |
Object
|
|
Void
| Name | Type | Description | |
|---|---|---|---|
| block |
Object
|
|
Void
| Name | Type | Description | |
|---|---|---|---|
| statement |
Object
|
|
|
| key |
String
|
|
Void
| Name | Type | Description | |
|---|---|---|---|
| statement |
Object
|
|
Object
a replacement BlockStatement, or null when the header is clean and the loop should be left as written
| Name | Type | Description | |
|---|---|---|---|
| varDecNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| ifNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| consequent |
Array
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| tNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| mNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| ast |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| arrNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| node |
Object
|
|
Boolean
| Name | Type | Description | |
|---|---|---|---|
| statement |
Object
|
|
Boolean
| Name | Type | Description | |
|---|---|---|---|
| node |
Object
|
|
|
| name |
String
|
|
Boolean
| Name | Type | Description | |
|---|---|---|---|
| statement |
Object
|
|
Boolean
| Name | Type | Description | |
|---|---|---|---|
| statement |
Object
|
|
Boolean
| Name | Type | Description | |
|---|---|---|---|
| textureCache |
Array.<WebGLTexture>
|
|
|
| programUniformLocationCache |
Object.<string, WebGLUniformLocation>
|
|
|
| framebuffer |
WebGLFramebuffer
|
|
|
| buffer |
WebGLBuffer
|
|
|
| program |
WebGLProgram
|
|
|
| functionBuilder |
FunctionBuilder
|
|
|
| pipeline |
Boolean
|
|
|
| endianness |
string
|
|
|
| argumentTypes |
Array.<string>
|
|
|
| compiledFragmentShader |
string
|
|
|
| compiledVertexShader |
string
|
|
Void
| Name | Type | Description | |
|---|---|---|---|
| type | |||
| dynamic | |||
| precision | |||
| value |
KernelValue
| Name | Type | Description | |
|---|---|---|---|
| source |
String
IKernelJSON
|
||
| settings |
IDirectKernelSettings
|
Void
| Name | Type | Description | |
|---|---|---|---|
| settings |
IDirectKernelSettings
|
Array.<string>
| Name | Type | Description | |
|---|---|---|---|
| args |
Array
|
|
Object
An object containing the Shader Artifacts(CONSTANTS, HEADER, KERNEL, etc.)
| Name | Type | Description | |
|---|---|---|---|
| args |
Array
|
|
Object
An object containing the Shader Artifacts(CONSTANTS, HEADER, KERNEL, etc.)
| Name | Type | Description | |
|---|---|---|---|
| args |
Array
|
|
String
result
| Name | Type | Description | |
|---|---|---|---|
| src |
String
|
|
|
| map |
Object
|
|
Void
| Name | Type | Description | |
|---|---|---|---|
| args |
Array
|
|
string
Fragment Shader string
| Name | Type | Description | |
|---|---|---|---|
| args |
Array
IArguments
|
|
string
Vertical Shader string
| Name | Type | Description | |
|---|---|---|---|
| idtNode |
Object
|
|
|
| retArr |
Array
|
|
Array
the append retArr
| Name | Type | Description | |
|---|---|---|---|
| args |
Array
|
|
String
result
| Name | Type | Description | |
|---|---|---|---|
| settings |
Object
|
||
| settings.buffer |
GPUBuffer
|
||
| settings.output |
Array.<number>
|
logical dims, e.g. [512, 512] |
|
| settings.componentCount |
number
|
1 for scalar returns, 2/3/4 for Array(n) |
Optional |
| settings.context |
Object
|
the |
|
| settings.kernel |
Object
|
owning WebGPUKernel, used for readback + shaping |
Void
Promise.<Float32Array|Array>
values shaped exactly as a non-pipeline run of the producing kernel would resolve them
Destroys the shared device. gpu.destroy() does NOT call this — the
device outlives any one GPU instance; this is the page-level teardown.
Promise.<undefined>
WGSL rejects float literals that overflow f32 (GLSL forgave them), and JS toString of an integral double has no decimal point. Every float literal goes through here so the emitted text is always a committed, in-range WGSL float.
| Name | Type | Description | |
|---|---|---|---|
| value |
number
|
String
User function names collide with WGSL keywords and builtins where GLSL names did not; mangle rather than reject. Unconditionally: WGSL reserves over sixty words that are legal JavaScript function names (filter, get, set, type, self, ...), and a curated list drifts out of date with the spec -- #861 found 64 missing. The fn_ prefix removes the class the way user_ already does for variables; the registry side (FunctionBuilder, type inference) keys on original names and never sees this.
| Name | Type | Description | |
|---|---|---|---|
| name |
String
|
String
A WebGPUBufferResult argument reads like a flat storage array; the base typeLookupMap does not know the type, so value lookup is resolved here.
Void
WGSL has no ternary. Pure-value case lowers to select(false, true, cond)
— both sides evaluate eagerly, which is observationally safe in the
side-effect-free kernel language. Void (minified) case lowers to if/else.
Void
Root kernel: emits only body statements — the kernel class assembles the
@compute entry, guard and data_index around them. Non-root: a full
fn name(args) -> type { ... } declaration.
Void
** upconverts to pow(f32, f32); bitwise operators are native in WGSL
(GLSL ES 1.00 needed helper functions) and need only operand casts.
Void
kernels created for second and later type signatures of one plan kernel (the plan clone carries the first); destroyed with the executor
Void
Stand-ins with the exact types and dims each binding will have at run time, for the kernel build: sampled values stand for themselves, a step output becomes an Input over its producer's dims.
Void
The layout baked argument sizes and scalar types; a call that drifts from them throws recompilable so the pipeline compiles a fresh fused plan for the new signature, the way the kernel itself rebuilds per size signature.
Void
| Name | Type | Description | |
|---|---|---|---|
| args |
Array
|
|
Promise.<*>
results shaped per the plan. Argument drift throws a recompilable FusionFallback synchronously, before anything is encoded; async failures (device loss) reject, which drops the executor via the pipeline's _guardAsync so the next call compiles fresh.
The shared device is a module singleton that outlives any one GPU instance; page-level teardown goes through WebGPUContext.destroy().
Void
Everything device-independent — validation, JS→WGSL translation, module assembly — happens synchronously here, so unsupported constructs throw at the first call rather than rejecting; the device round trip continues in _buildAsync and run() chains on its promise.
Void
Params struct, matching the WGSL struct member for member:
outputX/outputY/outputZ/_pad0, one vec4
Void
| Name | Type | Description | |
|---|---|---|---|
| threadDim |
Array.<number>
|
[object Object]
Oversized bindings must throw here: past the device limit, bind-group validation fails asynchronously, the submit is dropped, and the zero- initialized staging buffer would resolve a fully-shaped all-zeros result — silent wrong data instead of an error.
Void
The mutation-sensitive step: array arguments are flattened into fresh Float32Arrays synchronously inside run(), before any await, preserving the GL path's snapshot semantics.
Void
Steady state is synchronous through queue.submit — writeBuffer copies its data at call time and one queue executes submits in order, so overlapping un-awaited calls cannot race; only the readback awaits, on a staging buffer of its own.
Void
A staging buffer is unusable from mapAsync until unmap, so overlapping un-awaited calls each need their own; sequential awaited calls reuse one buffer forever. Capped so a burst cannot ratchet memory.
Void
Readback is tightly packed, so scalar returns reuse the memory-optimized erectors; Array(n) returns are stride n (not the GL texel stride 4), shaped locally.
Void
Readback for a pipeline handle: copy → map → shape, same conventions as a non-pipeline run of the producing kernel.
| Name | Type | Description | |
|---|---|---|---|
| handle |
WebGPUBufferResult
|
Promise.<Float32Array|Array>
| Name | Type | Description | |
|---|---|---|---|
| flip |
Boolean
|
Optional |
Promise.<Uint8ClampedArray>
Puts this value's texture back on its unit. Texture units are context state shared by every kernel, and each kernel numbers its own from zero, so between runs another kernel's textures sit on them (#862). The data in this texture is intact -- only the binding needs to come back.
Void
bit storage ratio of source to target 'buffer', i.e. if 8bit array -> 32bit tex = 4
| Name | Type | Description | |
|---|---|---|---|
| value |
number
| Name | Type | Description | |
|---|---|---|---|
| value |
KernelVariable
|
||
| settings |
IWebGLKernelValueSettings
|
Void
Re-establishes whatever this value put on shared context state before a run. Scalar uniforms live on the program and need nothing.
Void
Used for when we want a string output of our kernel, so we can still input values to the kernel
Void
| Name | Type | Description | |
|---|---|---|---|
| inputTexture |
GLTextureMemoryOptimized
|
Void