The GPU export now runs the same upscaler in half precision. The chip is handed
2.34MB of weights instead of 4.88MB, and where its shaders can multiply in fp16
it does twice the work per pass.
`realesr-fp16.py` is the conversion, run on what `realesr-gpu.py` already
wrote (the PReLU-rewritten model), never instead of it. onnxconverter-common's
`keep_io_types` needed two of its own mistakes put right:
- It rewrites the consumers of the graph input but misses the one that never
goes through the network. This model adds a Resize of the ORIGINAL photo to
the upsampler's output, that Resize reads the graph input directly, and the
runtime refuses a graph whose final Add mixes fp32 and fp16. The consumer is
rewired onto the cast that `keep_io_types` should have sent it through.
- It also half-precisions Resize's `scales` — ONNX defines that input as
float32 whatever the rest of the graph does, and a runtime that opens the
file at all rejects the whole graph: "Type 'tensor(float16)' of input
parameter (/Constant_output_0) of operator (Resize) is invalid", on the GPU
as much as on the processor. The script widens it back and asserts it did.
The tensor the app builds stays float32 and the model's two Cast nodes are its
own edge, so nothing in superRes.ts or App.tsx has to know which copy it got:
205 nodes, 101 fp16 weights, io still float.
`openSession` asks for the model only where the adapter advertises
`shader-f16` — a provider without it emulates the type on the same file at the
same speed, so the smaller download would be the only thing gained. The order
is fp16 on the GPU, fp32 on the GPU, fp32 on the processor, each attempt
falling through on its own failure.
Measured on the rebuilt container (BASE=http://localhost:8090):
- fp16 vs fp32 on a 128x128 tile, same graph: max abs diff 0.0025 (0.65/255),
mean 0.00028, psnr 71.0dB.
- sr-f16-chooser.cjs 4 PASS / 0 FAIL: on a forged adapter advertising
`shader-f16`, the fp16 file is the FIRST model asked for; on one whose device
refuses, the fp32 file is fetched for the processor and the 4K export still
lands (7,555,377 bytes, 19.6s), no console errors.
- superres-test.cjs 32 PASS / 0 FAIL, sr-crop-export.cjs 0 FAIL,
web-smoke.cjs 0 FAIL, sr-model-probe.cjs 0 FAIL.
- npx tsc --noEmit clean.
ponytail: the speed of the fp16 path is NOT measured — this container has no
WebGPU adapter (not even lavapipe/swiftshader, headed through xvfb), so every
export here runs the wasm fallback. sr-model-probe.cjs on a machine with a GPU
is what would show it.
Also worth noting for the next person: in a browser with no working adapter,
the runtime builds the device BEFORE it fetches the model, so no probe in a
GPU-less container can observe which model was chosen — a stub whose device
throws leaves the network silent. The chooser probe forges a device good enough
to be accepted for exactly that reason.
The WebGPU execution provider has no PReLU kernel. The model is 34 convolutions
with a PReLU after every one of them, so an export that took the GPU path was
split 33 times: each activation came off the chip to be activated on the
processor and went straight back, a 64-channel map in both directions, per tile.
A machine with a good graphics chip was not exporting any faster for having it.
PReLU(x) is exactly Relu(x) - slope * Relu(-x), and Relu, Neg, Mul and Sub the
provider does implement, so scripts/realesr-gpu.py writes the 33 activations out
as those four and drops the slopes nobody reads any more. The model file is the
output of that script, not the file as published.
One 256x256 tile through the model before and after, on a WebGPU session: the
runtime no longer reports nodes left off the preferred provider (it did, once,
before) and the processor path answers bit for bit what it answered before. The
warning itself cannot be switched off from here - env.logLevel is read when the
runtime module initialises, before any of this runs - so the graph was fixed
rather than the lines hidden.
The server still never sees a photo, so the model has to run in the page.
Real-ESRGAN x4v3 ships as a 4.9MB ONNX in public/models and is loaded
lazily on the first export that actually needs it; the wasm runtime is
copied next to CanvasKit at build time and stays lazily fetched, cached
for 30 days. Vite is told onnxruntime-web is external-wasm so no 28MB
asset lands in the bundle.
UNCHANGED keeps the old path and the tier cap; 2K/4K/custom upscale only
when the request is larger than the photo being edited, otherwise they
resize down. Guests keep UNCHANGED and 2K. Tiling is 256px with an 8px
overlap, so memory follows the target size rather than four times it.
`docker/` now holds the whole web build — frontend (Vite + React + CanvasKit),
backend (Fastify + SQLite) and the compose file — so the folder can be moved to
another machine and run without the React Native project:
cd docker && cp .env.example .env && docker compose up -d --build
Only `${WEB_PORT:-8090}` is published; nginx serves the SPA and proxies /api to
the `api` container over Docker's DNS. Photos never reach the server.
The shared render code is vendored into `docker/frontend/shared/` and aliased to
a CanvasKit shim, so the app's own frameUtils/toneShader/jpegDpi run unchanged.
Fix the all-black render on GPU surfaces: `MakeWebGLCanvasSurface` creates a
separate WebGL context per call, and a texture from one context cannot be
sampled by a surface on another — so any pass that drew a snapshot onto a second
surface (output sharpen, screen sharpen, polaroid/wallframe cards) came out
solid black, while the raster fallback was correct. Use one shared
GrDirectContext + MakeRenderTarget instead.
Verified in headless Chromium against the running stack: 12MP JPEG in, preview
mean=120.5 sd=60.5, export 2048x1536 mean=107.2 sd=62.1, JFIF density 300/300,
EXIF present, no console errors; health/signup/login/me/recipes all 2xx through
the nginx proxy.