Commit Graph

1 Commits

Author SHA1 Message Date
3dtours a9030fc0c0 web: give the GPU path a half-precision upscaler
The GPU export now runs the same upscaler in half precision. The chip is handed
2.34MB of weights instead of 4.88MB, and where its shaders can multiply in fp16
it does twice the work per pass.

`realesr-fp16.py` is the conversion, run on what `realesr-gpu.py` already
wrote (the PReLU-rewritten model), never instead of it. onnxconverter-common's
`keep_io_types` needed two of its own mistakes put right:

- It rewrites the consumers of the graph input but misses the one that never
  goes through the network. This model adds a Resize of the ORIGINAL photo to
  the upsampler's output, that Resize reads the graph input directly, and the
  runtime refuses a graph whose final Add mixes fp32 and fp16. The consumer is
  rewired onto the cast that `keep_io_types` should have sent it through.
- It also half-precisions Resize's `scales` — ONNX defines that input as
  float32 whatever the rest of the graph does, and a runtime that opens the
  file at all rejects the whole graph: "Type 'tensor(float16)' of input
  parameter (/Constant_output_0) of operator (Resize) is invalid", on the GPU
  as much as on the processor. The script widens it back and asserts it did.

The tensor the app builds stays float32 and the model's two Cast nodes are its
own edge, so nothing in superRes.ts or App.tsx has to know which copy it got:
205 nodes, 101 fp16 weights, io still float.

`openSession` asks for the model only where the adapter advertises
`shader-f16` — a provider without it emulates the type on the same file at the
same speed, so the smaller download would be the only thing gained. The order
is fp16 on the GPU, fp32 on the GPU, fp32 on the processor, each attempt
falling through on its own failure.

Measured on the rebuilt container (BASE=http://localhost:8090):
- fp16 vs fp32 on a 128x128 tile, same graph: max abs diff 0.0025 (0.65/255),
  mean 0.00028, psnr 71.0dB.
- sr-f16-chooser.cjs 4 PASS / 0 FAIL: on a forged adapter advertising
  `shader-f16`, the fp16 file is the FIRST model asked for; on one whose device
  refuses, the fp32 file is fetched for the processor and the 4K export still
  lands (7,555,377 bytes, 19.6s), no console errors.
- superres-test.cjs 32 PASS / 0 FAIL, sr-crop-export.cjs 0 FAIL,
  web-smoke.cjs 0 FAIL, sr-model-probe.cjs 0 FAIL.
- npx tsc --noEmit clean.

ponytail: the speed of the fp16 path is NOT measured — this container has no
WebGPU adapter (not even lavapipe/swiftshader, headed through xvfb), so every
export here runs the wasm fallback. sr-model-probe.cjs on a machine with a GPU
is what would show it.

Also worth noting for the next person: in a browser with no working adapter,
the runtime builds the device BEFORE it fetches the model, so no probe in a
GPU-less container can observe which model was chosen — a stub whose device
throws leaves the network silent. The chooser probe forges a device good enough
to be accepted for exactly that reason.
2026-09-25 20:01:51 +07:00