web: ask the GPU for the fast one before an export runs

onnxruntime hands a webgpu session whatever adapter the browser picks by
default, which on a laptop with both is the one built into the processor: the
export then waits on the slow half of the machine for no reason. The runtime
reads `env.webgpu.powerPreference` when it builds the webgpu session, so set it
to high-performance; a browser with nothing to honour the preference with
still falls back to the threaded wasm path exactly as before.
This commit is contained in:
2026-09-23 08:36:03 +07:00
parent 28a688fd82
commit a2c12a2638
+3
View File
@@ -51,6 +51,9 @@ function load(): Promise<Loaded> {
ort.env.wasm.numThreads = self.crossOriginIsolated
? Math.min(8, navigator.hardwareConcurrency || 1)
: 1;
// A laptop with two GPUs would otherwise hand this to the one built into
// the processor. The model is the export, so ask for the fast one.
ort.env.webgpu.powerPreference = 'high-performance';
const session = await ort.InferenceSession.create(MODEL_URL, { executionProviders: ['webgpu'] }).catch(() =>
ort.InferenceSession.create(MODEL_URL, { executionProviders: ['wasm'] })
);