sensorWhite hung its nine counts of window off the plane's largest sample. A
hot pixel sits hundreds of counts above the level the sensor stops at, so on a
frame that carries one the window held the stray alone, found no pile in it, and
handed the develop the factor two instead of the frame's own level.
Measured on a Panasonic DMC-LX10 RW2: one sample at 15993 and one at 14665 over
a pile of 9,594,544 at 13855. The frame's level is 1.75x the fall-back, so the
develop opened 0.81 of a stop bright, clipped the sky the sensor had held to
flat, and the fit to the camera's preview could only pull the exposure back
after the highlight detail was gone. Against the camera's own JPEG of the shot,
mean |dL| 23.48 and dRGB +8.57,+0.90,+4.90 became 19.14 and +5.32,+1.89,+2.04,
and the level itself 7908 (gain 8.2872) became 13874 (gain 4.7236) against the
13855 the plane piled at — 0.14%, 0.002 of a stop.
The level is now the highest count the plane piled at, off a whole-plane
histogram: `floor` counts is a pile, and the same cliff rule as before still has
to hold over the count below it, so a smooth bright sky is left alone and a
frame that has not clipped still falls back on the factor two.
The same stray was in four of the ten bodies to hand. A Nikon _GDN0447.NEF read
the fall-back 8190 where its plane piles at 13806 (gain 8.0018 -> 4.6022, 0.79
of a stop), a Fujifilm RAF 29696 where the documented level is 30993, and a Sony
ARW 31742 against the 2.002x the body's ratio was measured at. Six were
untouched, and a Canon CR2 develops byte-identical through the change — the
window moves only on a frame whose largest sample is not its highest pile.
scripts/white-level-check.mjs keeps the LX10 shape: a plane piled at 30995 with
a stray above it has to answer 30995.
The AUTO chip measured the photo and wrote one knob, EV. It now writes the four
the measurement actually names, off the same binned ramp (ui/Histogram.tsx),
which is why it is one chip and not four: the means it needs are all in the
histogram the exposure answer already reads.
EV the mean luma, unchanged
HIGHLIGHT the top 1% (p99 > 0.9 pulls back), to -5 of the ruler at most
SHADOW the bottom 1% (p01 < 0.02 opens up), to +5
TEMPERATURE/TINT the gain that puts the three channel means on each other,
green as the anchor: gray-world on linearised means, then the
closest of 76 temperatures x 21 tints under the renderer's own
kelvinToRGB, so the pair cannot drift from what the ruler applies.
The ends rather than the average is what keeps a small blown window from
dragging the whole frame: a specular in the corner wants HIGHLIGHT, not a
flatter picture everywhere. Both ends stop at half the ruler, so the frame is
corrected and a hand can still finish the move; the knobs then report the
numbers AUTO chose, the way the EV knob does.
Each reading is a pure function of the ramp, so pressing AUTO twice lands on the
same recipe by construction, and a frame with a dead channel leaves the WB ruler
where it is rather than inventing a cast. ponytail: one linear ramp per end and
no scene analysis; add a curve, or weight by how much of the frame is clipped,
when AUTO starts overshooting a scene with a genuine specular in it.
scripts/auto-tone-check.mjs holds the three readings: the percentile walk, the
thresholds that leave a knob alone, and the scan landing back on the gain it was
asked for.
LibRaw's half-size demosaic was on. The Ricoh GR's own DNG (D0004128.DNG)
developed to 3010x2012 while the JPEG written beside it in the same second is
6000x4000, and the Fuji's RAF to 3008x2007 against its own 6000x4000 -- the
quarter was the flag, not the file. With `halfSize: false` the same develop
returns 6020x4024 and it is the sensor's frame on every body tried:
D0004128.DNG 6020x4024 IMGP6916.DNG 6028x4024
DSCF1701.RAF 6016x4014 _DSC0009.ARW 6024x4024
AFXT2721.RAF 6246x4170 Nikon-D850 NEF 6216x4136
_GDN0447.NEF 4284x2844 P1010607.RW2 3472x3472
5G4A9396.CR2 2880x1920
Nine files, 27s to 155s a develop on one core. Checked through the app
itself, not only through LibRaw: photo-dims 6020x4024 on the DNG against
6000x4000 on the JPEG, both err none.
The colour it opens with is now fitted per file to the preview the camera wrote
into it (previewMatch.ts): a 3x3 over a block grid of the develop against the
same grid of that preview, then one cubic a channel for what the 3x3 leaves.
The offline per-body table this replaces (cameraMatch.ts) stopped matching the
moment the path under it changed -- its rows no longer summed to 1 once the
highlight knee landed ahead of it -- and a body with a row opened with a cast
one without did not. The file's own preview does not age.
The white level the gain carries is the frame's own plateau rather than
`maximum` (sensorWhite.ts), a factor of 1.89 to 2.00 out; without it every
frame opened a stop bright and a body that sat lower (X-Trans, 1.892) never
reached the highlight desaturation at all.
The desaturation gate reads the gain-lifted levels as well as the sensor's,
which is the whole of the magenta: on a body whose cam_mul lifts red and blue
(the GR's [2.64, 1, 1.73]) a blown sky crosses the white level at 0.38 of the
raw range in red while green crosses at 1.0, so a gate read on the sensor's
levels alone stayed shut across it. Measured in the app against the camera's
own JPEG, mean dRGB over a 16x16 block grid: +1.20, -5.95, -6.11 with the
sensor's clip alone, +0.21, +0.24, +0.47 with both, mean |dL| 21.5 against
10.3. The same grid on the Fuji comes back balanced (+4.7, +5.0, +3.6) and best
aligned at offset 0,0.
-HL is recovery and +HL is a lift, so they are different moves now: recovery is
the doc's soft knee in linear light over the top half, which is the only term
in the tone shader that is not a shift and the only one that can put detail
back into a blown sky rather than merely darken it.
The four checks pin the develop down where it can only run in a browser:
raw-develop-check, preview-match-check, white-level-check, highlight-knee-check.
An RGBA_F32 image with an sRGB tag comes back off the GPU backend sampled on a
1/255 grid; the same shader on a raster surface returns the floats untouched.
The plane is raw/65535, so the shadows the black level is there to keep sit at
1e-3 and quantise to zero -- a 3010x2012 develop landed 41189 pixels under luma
2 with the dark end speckled blue/yellow, against none on the raster surface.
A half is uploaded as float, so the plane stays exact either way.
Rejects the earlier guess that the render target's colour space was to blame:
gpu+rt-srgb and gpu+img-untagged came back byte-identical to gpu.
scripts/half-check.mjs checks the conversion: the named encodings, and no plane
value in a 14-bit sensor's range moving more than 4.8e-4 relative.
The GPU export now runs the same upscaler in half precision. The chip is handed
2.34MB of weights instead of 4.88MB, and where its shaders can multiply in fp16
it does twice the work per pass.
`realesr-fp16.py` is the conversion, run on what `realesr-gpu.py` already
wrote (the PReLU-rewritten model), never instead of it. onnxconverter-common's
`keep_io_types` needed two of its own mistakes put right:
- It rewrites the consumers of the graph input but misses the one that never
goes through the network. This model adds a Resize of the ORIGINAL photo to
the upsampler's output, that Resize reads the graph input directly, and the
runtime refuses a graph whose final Add mixes fp32 and fp16. The consumer is
rewired onto the cast that `keep_io_types` should have sent it through.
- It also half-precisions Resize's `scales` — ONNX defines that input as
float32 whatever the rest of the graph does, and a runtime that opens the
file at all rejects the whole graph: "Type 'tensor(float16)' of input
parameter (/Constant_output_0) of operator (Resize) is invalid", on the GPU
as much as on the processor. The script widens it back and asserts it did.
The tensor the app builds stays float32 and the model's two Cast nodes are its
own edge, so nothing in superRes.ts or App.tsx has to know which copy it got:
205 nodes, 101 fp16 weights, io still float.
`openSession` asks for the model only where the adapter advertises
`shader-f16` — a provider without it emulates the type on the same file at the
same speed, so the smaller download would be the only thing gained. The order
is fp16 on the GPU, fp32 on the GPU, fp32 on the processor, each attempt
falling through on its own failure.
Measured on the rebuilt container (BASE=http://localhost:8090):
- fp16 vs fp32 on a 128x128 tile, same graph: max abs diff 0.0025 (0.65/255),
mean 0.00028, psnr 71.0dB.
- sr-f16-chooser.cjs 4 PASS / 0 FAIL: on a forged adapter advertising
`shader-f16`, the fp16 file is the FIRST model asked for; on one whose device
refuses, the fp32 file is fetched for the processor and the 4K export still
lands (7,555,377 bytes, 19.6s), no console errors.
- superres-test.cjs 32 PASS / 0 FAIL, sr-crop-export.cjs 0 FAIL,
web-smoke.cjs 0 FAIL, sr-model-probe.cjs 0 FAIL.
- npx tsc --noEmit clean.
ponytail: the speed of the fp16 path is NOT measured — this container has no
WebGPU adapter (not even lavapipe/swiftshader, headed through xvfb), so every
export here runs the wasm fallback. sr-model-probe.cjs on a machine with a GPU
is what would show it.
Also worth noting for the next person: in a browser with no working adapter,
the runtime builds the device BEFORE it fetches the model, so no probe in a
GPU-less container can observe which model was chosen — a stub whose device
throws leaves the network silent. The chooser probe forges a device good enough
to be accepted for exactly that reason.
The WebGPU execution provider has no PReLU kernel. The model is 34 convolutions
with a PReLU after every one of them, so an export that took the GPU path was
split 33 times: each activation came off the chip to be activated on the
processor and went straight back, a 64-channel map in both directions, per tile.
A machine with a good graphics chip was not exporting any faster for having it.
PReLU(x) is exactly Relu(x) - slope * Relu(-x), and Relu, Neg, Mul and Sub the
provider does implement, so scripts/realesr-gpu.py writes the 33 activations out
as those four and drops the slopes nobody reads any more. The model file is the
output of that script, not the file as published.
One 256x256 tile through the model before and after, on a WebGPU session: the
runtime no longer reports nodes left off the preferred provider (it did, once,
before) and the processor path answers bit for bit what it answered before. The
warning itself cannot be switched off from here - env.logLevel is read when the
runtime module initialises, before any of this runs - so the graph was fixed
rather than the lines hidden.
The server still never sees a photo, so the model has to run in the page.
Real-ESRGAN x4v3 ships as a 4.9MB ONNX in public/models and is loaded
lazily on the first export that actually needs it; the wasm runtime is
copied next to CanvasKit at build time and stays lazily fetched, cached
for 30 days. Vite is told onnxruntime-web is external-wasm so no 28MB
asset lands in the bundle.
UNCHANGED keeps the old path and the tier cap; 2K/4K/custom upscale only
when the request is larger than the photo being edited, otherwise they
resize down. Guests keep UNCHANGED and 2K. Tiling is 256px with an 8px
overlap, so memory follows the target size rather than four times it.
`docker/` now holds the whole web build — frontend (Vite + React + CanvasKit),
backend (Fastify + SQLite) and the compose file — so the folder can be moved to
another machine and run without the React Native project:
cd docker && cp .env.example .env && docker compose up -d --build
Only `${WEB_PORT:-8090}` is published; nginx serves the SPA and proxies /api to
the `api` container over Docker's DNS. Photos never reach the server.
The shared render code is vendored into `docker/frontend/shared/` and aliased to
a CanvasKit shim, so the app's own frameUtils/toneShader/jpegDpi run unchanged.
Fix the all-black render on GPU surfaces: `MakeWebGLCanvasSurface` creates a
separate WebGL context per call, and a texture from one context cannot be
sampled by a surface on another — so any pass that drew a snapshot onto a second
surface (output sharpen, screen sharpen, polaroid/wallframe cards) came out
solid black, while the raster fallback was correct. Use one shared
GrDirectContext + MakeRenderTarget instead.
Verified in headless Chromium against the running stack: 12MP JPEG in, preview
mean=120.5 sd=60.5, export 2048x1536 mean=107.2 sd=62.1, JFIF density 300/300,
EXIF present, no console errors; health/signup/login/me/recipes all 2xx through
the nginx proxy.