a9030fc0c0
The GPU export now runs the same upscaler in half precision. The chip is handed 2.34MB of weights instead of 4.88MB, and where its shaders can multiply in fp16 it does twice the work per pass. `realesr-fp16.py` is the conversion, run on what `realesr-gpu.py` already wrote (the PReLU-rewritten model), never instead of it. onnxconverter-common's `keep_io_types` needed two of its own mistakes put right: - It rewrites the consumers of the graph input but misses the one that never goes through the network. This model adds a Resize of the ORIGINAL photo to the upsampler's output, that Resize reads the graph input directly, and the runtime refuses a graph whose final Add mixes fp32 and fp16. The consumer is rewired onto the cast that `keep_io_types` should have sent it through. - It also half-precisions Resize's `scales` — ONNX defines that input as float32 whatever the rest of the graph does, and a runtime that opens the file at all rejects the whole graph: "Type 'tensor(float16)' of input parameter (/Constant_output_0) of operator (Resize) is invalid", on the GPU as much as on the processor. The script widens it back and asserts it did. The tensor the app builds stays float32 and the model's two Cast nodes are its own edge, so nothing in superRes.ts or App.tsx has to know which copy it got: 205 nodes, 101 fp16 weights, io still float. `openSession` asks for the model only where the adapter advertises `shader-f16` — a provider without it emulates the type on the same file at the same speed, so the smaller download would be the only thing gained. The order is fp16 on the GPU, fp32 on the GPU, fp32 on the processor, each attempt falling through on its own failure. Measured on the rebuilt container (BASE=http://localhost:8090): - fp16 vs fp32 on a 128x128 tile, same graph: max abs diff 0.0025 (0.65/255), mean 0.00028, psnr 71.0dB. - sr-f16-chooser.cjs 4 PASS / 0 FAIL: on a forged adapter advertising `shader-f16`, the fp16 file is the FIRST model asked for; on one whose device refuses, the fp32 file is fetched for the processor and the 4K export still lands (7,555,377 bytes, 19.6s), no console errors. - superres-test.cjs 32 PASS / 0 FAIL, sr-crop-export.cjs 0 FAIL, web-smoke.cjs 0 FAIL, sr-model-probe.cjs 0 FAIL. - npx tsc --noEmit clean. ponytail: the speed of the fp16 path is NOT measured — this container has no WebGPU adapter (not even lavapipe/swiftshader, headed through xvfb), so every export here runs the wasm fallback. sr-model-probe.cjs on a machine with a GPU is what would show it. Also worth noting for the next person: in a browser with no working adapter, the runtime builds the device BEFORE it fetches the model, so no probe in a GPU-less container can observe which model was chosen — a stub whose device throws leaves the network silent. The chooser probe forges a device good enough to be accepted for exactly that reason.