webpost.ing

How webpaint.ing draws: the engine and its pipeline on the graphics card

By releases ·

webpaint.ing is the drawing app that sits beside this site. This article is for programmers and graphics people, and it follows one question from the pen to the screen: what runs on the graphics card, in what order, and why each piece is written the way it is. A companion article, Materials that wear and flow, goes inside the brushes that behave like physical things: chalk, graphite, watercolor, smudge and paper. Everything here was read from the code as it stands today, and numbers are the code's numbers. Where something is only planned, it says so.

Two notes on the format. The site's post format has no inline math, so quantities inside sentences are written in words and every formula stands on its own line. And the figures are of three kinds, always named in the caption: screenshots of the running app, charts computed by a script from the same constants and formulas the shaders use, and diagrams of boxes and arrows drawn for this article.

1. Two promises that shape every shader

The first promise is that a drawing is exactly its list of operations. Every mark is an op, a small record of what was done. A stroke op holds its points as whole numbers: positions in sixteenths of a pixel, pressure from 0 to 255, each point after the first stored as a difference from the one before. A fill op holds the runs of pixels it filled. Nothing changes the drawing except applying an op, and opening a saved drawing, Undo and a shared drawing all replay ops through the same door. So replaying must give the very pixels that drawing gave, on the same machine and, for the parts that matter, on any machine.

That promise decides how each shader is written. Where a result feeds later results (a worn stick, a smudged picture, a flooded region, a paper's grain, a dither pattern, paint spreading across wet paper), the arithmetic is done in whole numbers, because a float sum depends on the order of its terms and different cards round differently. Where a result is a pure function evaluated fresh from its inputs every time (the blend of one layer with the screen under it, a color adjustment, the soft edge of a stamp), floats are used, and the output is rounded to a byte once at the end. The article says which is which as it goes.

A few devices appear again and again. Pixels live on the card as bytes, and shaders receive them as fractions of 255. Rounding the fraction times 255 gives the byte back exactly, so a shader that wants whole numbers rounds on the way in and divides by 255 on the way out. A square root is taken as a float guess and then corrected, so the result is the exact whole number:

\small \operatorname{isqrt}(n) = \max \{\, s \in \mathbb{Z} : s^2 \le n \,\}

And a shift of a negative number would round toward minus infinity on one machine and toward zero on another if left to chance, so positions in sixteenths have a multiple of 16 added before a shift, which makes the shift a floor.

The second promise is a rule of the hardware: a shader cannot read the picture it is writing. Three habits follow. A step that must build on its own last result reads picture A and writes picture B, then the roles swap. A job that needs a finished picture is queued in one frame and runs in the next, because a picture drawn in a frame is not read in that frame. And an answer asked of the card, such as pixels to read back, arrives frames later on WebGPU, so a question is asked, drawing goes on, and the answer is collected when it comes.

2. What runs where

A graphics card can run two kinds of program. A drawing shader paints shapes and pictures into a picture. A compute shader runs over a grid of numbers with nothing drawn, thousands of copies side by side. The third kind of work is ordinary code on the processor. Nearly everything in the app is the first kind: stamps and the pieces that join them, laying a stroke on its layer, showing each layer and mixing it with the ones below, fills, gradients, dither, papers, the soft shadow under a cut piece, the selection's moving dashes, onion skin, the colors of the color wheel, and the 3D reference model.

Compute shaders are fewer: there are four uses. The flood behind Fill and Select by color has two passes. The enlarger that makes an export 2, 4 or 8 times larger is one pass, and runs on the page's own WebGPU device. The Chalk and Graphite simulation has two passes. The Smudge and Blend tool has three. Ordinary code still does the placing of brush stamps, the shape of a selection as a list of runs, and saving.

Three rows of boxes and arrows. Chalk: steps and stick buffers feed WEAR, which adds into a pigment buffer, which COVER turns into a coverage picture. Smudge: PULL copies the layer into a work picture, then STEP writes a scratch picture and APPLY copies it back. Fill: WITHIN fills a state picture, then 16 SWEEP passes alternate between two state pictures and mark a flags picture.

Diagram: the passes (orange), buffers (blue) and pictures (green) of the three compute families.

Several of the choices below come from limits of this engine's WebGPU driver, which are said here once. Reading a buffer back returns zeros, so every check of a compute result reads a picture instead, and the app never reads a buffer to learn a result. A group-shared array is refused when the shader is converted from SPIR-V to WGSL, so none is used; where a group has to agree on something, as the chalk simulation does on its lowest column, it uses an atomic in a buffer and barriers. A picture asked for in the same frame as the dispatch that fills it can come back as zeros, and the data of a card comes back frames late, which is why a read is asked for a frame after the work. And a shader that read an unsigned-integer texture with a nearest sampler produced an invalid pipeline on the main device in an early test, with the cause not found; the app's shaders read pixels as 8-bit float textures and round them back to bytes.

3. A stroke, from pen to layer

The stamper turns the pen's points into stamps: a position, a size and an alpha. The default spacing is 0.12 of the brush size, and the tool's settings clamp it to between 0.02 and 4. While the pen is down, each stamp is drawn into a scratch picture as one textured square, and the shader decides each pixel of the square by its distance from the middle. With d running from 0 at the center to 1 at the edge, hardness h (0.8 by default), and a one-pixel smoothing width w, the stamp's alpha is a smoothstep:

\small \begin{aligned} a(d) &= 1 - \operatorname{smoothstep}(e_0, 1, d) \\ e_0 &= \min(h,\ 1 - w) \end{aligned}
\small \begin{aligned} \operatorname{smoothstep}(e_0, e_1, x) &= 3t^2 - 2t^3 \\ t &= \operatorname{clamp}\!\left(\tfrac{x - e_0}{e_1 - e_0},\ 0,\ 1\right) \end{aligned}

Without antialiasing each pixel is wholly in or wholly out of the disc, decided at the pixel's center, and a stamp under three pixels across is its whole square, because a disc decided at pixel centers can miss every pixel and leave gaps in a line. The Pencil brush multiplies the alpha by a grain: each sheet pixel has a fixed height from 0 to 1, a 16-bit integer hash of its position, and the alpha is scaled by a mix between 1 and a smoothstep of that height from 0.15 to 0.85. At the Pencil's full grain of 0.75, a pixel in a hollow of the paper takes about a quarter of what a high point takes.

Stamps go into a transparent scratch picture with the engine's ordinary mix blend. Into nothing, that leaves premultiplied pixels: each color channel already multiplied by alpha. Everything on the card is held that way, for reasons the next section gives. Later stamps are laid over earlier ones, so the scratch picture holds the union of the stroke, and the stroke's opacity is applied once, when the finished scratch picture is laid on its layer. That is why a stroke at half opacity does not get darker where it crosses itself.

Joining the stamps into pieces

Stamps in a row leave a faintly scalloped edge, and on a large brush you can see it. So when a stroke lands, it is drawn again as pieces. Each pair of neighboring stamps is joined by the shape a disc sweeps as it slides from one to the next, growing or shrinking to the next size, which is a capsule with ends of different radius. The sheet is cut into square tiles of about four brush widths, rounded up in steps of 64 pixels, between 64 and 256. Each tile is one drawing job, given up to 64 pieces that touch it, and for every pixel the job takes the strongest cover of any piece, not the sum. For a pixel p and a piece from a to b with radii and alphas that vary along it:

\small \begin{aligned} t &= \operatorname{clamp}\!\left(\tfrac{(p - a)\cdot(b - a)}{\lvert b - a\rvert^2},\ 0,\ 1\right) \\ r &= \max\big(\operatorname{mix}(r_a, r_b, t),\ 0.5\big) \\ d &= \frac{\lvert p - (a + t(b - a))\rvert}{r} \end{aligned}
\small e = \operatorname{clamp}\big((1 - d)\,r + \tfrac12,\ 0,\ 1\big)
\small \text{cover}(p) = \max_{\text{pieces}} L(d)\; e\; \operatorname{mix}(\alpha_a, \alpha_b, t)

Here e smooths the edge over one pixel, and a pixel's cover is the largest, over all the pieces that reach it, of L times e times the piece's alpha. Taking the maximum is what makes joins invisible: where two pieces overlap at a stamp, the pixel is covered once. Tiles do not overlap, so nothing is added between jobs either, and a tile touched by more than 64 pieces gets a second job, where covers do add, as they do between two strokes. The loop stops early at a pixel that is already 0.999 covered, which is most of what makes a wide brush quick to land.

The function L is the brush's softness, and it is read from a table of 32 entries worked out once per stroke. A soft brush gets its softness from many overlapping discs, each adding a little, so the table holds what an endless even row of stamps adds up to at each distance from the line. Stamps laid over each other leave through the product of what each leaves, so with the stamp step 2s radii apart for a spacing s:

\small L(d) = 1 - \prod_{j=-J}^{J} \big(1 - a(d_j)\big)
\small \begin{aligned} d_j &= \sqrt{d^2 + (2sj)^2} \\ J &= \min\!\big(\lceil \tfrac{1}{2s} \rceil + 1,\ 60\big) \end{aligned}

where a is the single-stamp alpha above, with s at least 0.02 and d at most 0.999 so a hard brush is whole right up to its edge, and the shader smooths the edge itself. Where exactly the stamps sit along the line shifts the sum a little, and that shift is the scallop. The table holds the fullest case, a stamp exactly abreast of the pixel, which is the smooth outline the scallops touch from inside.

A thick black line drawn twice at three times its size, so each pixel is a square. The upper line is made of stamps and has a faintly scalloped top and bottom edge. The lower line is made of one piece and its edges are smooth.

Computed from the stamp and piece formulas above, for a brush 64 pixels wide, hardness 0.8 and spacing 0.3, the widest spacing that is joined. Upper: stamps. Lower: one piece.

A chart of cover against distance from the line as a share of the radius. A single stamp falls from 1 at 0.8 to 0 at 1. A row of stamps at spacing 0.12 and hardness 0.8 holds 1 a little further out and falls steeply to 0. A row at hardness 0.4 falls gradually from about 0.5.

Chart, computed from the table formula: what a row of stamps adds up to, for three settings, against one stamp. Spacing 0.3 gives nearly one stamp's edge.

The stamps stay the single source. Pieces are made from the stamp list alone, and the stamp list is made from the op's integer points, so replaying a stroke gives the same stamps and the same pieces. Only strokes for which joining is safe get pieces: those that say so themselves, the round brush with soft edges, every stamp of even strength, stamps no more than 0.3 of the size apart, and no color jitter, because joined pieces are one color. Everything else stays stamps. The eraser has its own pair of shaders that use the same cover and lay it on differently: they multiply what is on the layer by one minus the cover, color and alpha alike, which is true erasing on premultiplied pixels.

4. Showing layers

Premultiplied color

A pixel holds C, its color times alpha, and alpha a. Laying a source over a backdrop is then a single line, with no division:

\small \begin{aligned} C &= C_s + (1 - a_s)\,C_b \\ a &= a_s + (1 - a_s)\,a_b \end{aligned}

The reason to store color this way is what happens when pixels are averaged or interpolated, which happens constantly here: zooming out averages many sheet pixels into one screen pixel, and smudging mixes neighbors. An average of premultiplied pixels is correct as it stands, and a transparent pixel contributes nothing to the color, so paint thinning into empty paper keeps its color and does not darken or grow a fringe at its edge. A mix of valid premultiplied pixels with the same weights in every channel stays valid, because color never exceeds alpha.

One body, two shaders

A layer is shown by a body of shader code that reads the layer's picture at a sheet pixel, lays in the stroke being drawn at the layer's own place in the stack, and, while an Image menu dialog is open, previews its adjustment inside the selection's box. The stroke's provisional end, the stretch from the ink to the pen redrawn every frame, is a second texture joined to the stroke first, so where they overlap they count once. All of it passes through one function, so the stroke looks the same while the pen is down and after it is part of the layer.

How a screen pixel is made from sheet pixels depends on the view. At 100% zoom and above, a screen pixel shows the one sheet pixel it falls in, so a line does not look soft until the pen lifts and sharp after. Zoomed out, a screen pixel covers a square of the sheet, n pixels on a side, and shows the average of everything under it, each sheet pixel weighted by how much of it lies inside the square:

\small \begin{aligned} c &= \frac{1}{n^2} \sum_{i,j} w_x(i)\, w_y(j)\; c_{ij} \\ w_x(i) &= \min(x_{hi},\, i + 1) - \max(x_{lo},\, i) \end{aligned}

Neighboring screen pixels' squares meet edge to edge, so every sheet pixel is shown once, shared out among the screen pixels it lies under, which keeps a thin line evenly dark along its length. Before this, a screen pixel blended only the four sheet pixels nearest its center, so most of a thin line's pixels were never looked at and a smooth line looked broken. The cost does not grow as the view shrinks: a sheet shown at half size has a quarter of the screen pixels, each reading four times as many, which is about one read of each sheet pixel a frame. The loop is capped at 64 sheet pixels across, which is reached below about 1.6% zoom.

When the view is turned by anything but a right angle, the same average is taken over a square that lies askew across the sheet. Each nearby sheet pixel counts by its overlap with the square, taken in the square's own two directions u and v, with n the number of sheet pixels the screen pixel spans:

\small \begin{aligned} \text{share}(\delta, n) &= \operatorname{clamp}\big(\tfrac{n + 1}{2} - \lvert\delta\rvert, \\ &\qquad 0,\ \min(n, 1)\big) \\ \text{weight} &= \text{share}(\delta_u, n)\cdot \text{share}(\delta_v, n) \end{aligned}

The square's shares of one sheet pixel across neighboring screen pixels add up to exactly one. It costs up to twice the reads of the upright average, and capped at 96 sheet pixels across. Between 100% and about 141% zoom a turned picture is slightly soft, as any turned picture is, and from there up whole pixels are shown again.

Blend modes, and why a blended layer reads the screen

Mixing a layer with what is under it needs the thing under it. For a Normal layer the engine's ordinary blend does the laying, so the shader just returns the layer's color times its opacity. For any other mode the shader has to know the backdrop, which is whatever the layers below have already put on the screen, so it samples the screen under itself, mixes by the mode, and writes the finished pixel with blending switched off, so nothing is laid on top a second time.

Reading the screen makes the engine copy it for every layer that does, and a Normal layer should not pay for that, which is why there are two shaders sharing one body and not one shader with a switch. The mixing itself is in one shared include, so the screen, the export (a picture of the sheet as these shaders draw it), and the color picker, which lays layers one after another on a one-pixel surface, cannot disagree. With backdrop alpha and color a_b and C_b, source a_s and C_s (both premultiplied, the layer's opacity already in the source), and lowercase c the same colors with alpha divided out:

\small \begin{aligned} C &= (1 - a_s)\,C_b + (1 - a_b)\,C_s \\ &\quad + a_s\, a_b\, B(c_b, c_s) \\ a &= a_s + a_b - a_s a_b \end{aligned}

Put B equal to c_s, which is Normal, and the first formula reduces to the over operator above, since (1 - a_b) C_s + a_b C_s = C_s. Over nothing the result is the layer and under nothing it is the backdrop, and the shader returns the backdrop at once where the layer has no alpha. The function B is applied to each channel of the unpremultiplied colors, and the twelve modes are the W3C Compositing and Blending ones: Normal, Multiply, Screen, Overlay, Add, Darken, Lighten, Color dodge, Color burn, Soft light, Hard light and Difference.

\small \begin{aligned} \text{multiply} &= c_b c_s \\ \text{screen} &= c_b + c_s - c_b c_s \\ \text{add} &= \min(c_b + c_s,\ 1) \\ \text{difference} &= \lvert c_b - c_s \rvert \\ \text{dodge} &= \min\!\big(1,\ \tfrac{c_b}{1 - c_s}\big) \\ \text{burn} &= 1 - \min\!\big(1,\ \tfrac{1 - c_b}{c_s}\big) \end{aligned}
\small \text{hard light} = \begin{cases} 2\, c_b c_s & c_s \le \tfrac12 \\ c_b + s - c_b s & \text{otherwise} \end{cases}
\small s = 2c_s - 1

Overlay is hard light with the backdrop and the source swapped. Soft light is piecewise, and uses a helper D of the backdrop cb, which is the cubic 16 cb cubed minus 12 cb squared plus 4 cb up to a backdrop of 0.25 and the square root of cb above that:

\small \begin{aligned} c_s \le \tfrac12:\ & c_b - (1 - 2c_s)\, c_b (1 - c_b) \\ c_s > \tfrac12:\ & c_b + (2c_s - 1)\big(D(c_b) - c_b\big) \end{aligned}
A chart of the result of each of seven blend functions against the backdrop value, for a source value of 0.6. Multiply is a straight line to 0.6. Screen rises from 0.6 to 1. Color dodge reaches 1 early. Difference falls to 0 at 0.6 and rises again.

Chart, computed from the shared blend formulas: the result of each function against the backdrop, for a source value of 0.6.

Five small pictures of the same two strokes, a yellow horizontal one on a Normal layer and a blue vertical one on a layer above it, with the upper layer in modes Normal, Multiply, Screen, Add and Difference. In Normal the blue crosses over the yellow. In Multiply the crossing is dark green. In Screen and Add the blue disappears from the white sheet and only a pale patch shows at the crossing. In Difference the blue becomes orange.

Screenshot of the running app: the same two strokes with the upper layer set to Normal, Multiply, Screen, Add and Difference. Over the white sheet, Screen and Add give white, which is why the blue vanishes there.

The screenshot shows the reason for reading the screen. The backdrop of the upper layer includes the white sheet, so Screen over white is white and the blue disappears, while Difference of blue against white is its complement, orange. These are floats on purpose. The arithmetic is also written in plain code, and the tests compare that with the formulas. The card is checked in a real browser by an independent version of the formulas written for the check, at every mode, where the strokes cross and where only the upper one lies over the sheet, and again on a picture exported through the real Export dialog, with a tolerance of four units of 255 for the 8-bit steps on the card, in the screen copy and in the file.

Image adjustments

The Image menu's four adjustments run in the same shader whether they are being previewed or made, so what is previewed

Links