web: ask the GPU for the fast one before an export runs
onnxruntime hands a webgpu session whatever adapter the browser picks by default, which on a laptop with both is the one built into the processor: the export then waits on the slow half of the machine for no reason. The runtime reads `env.webgpu.powerPreference` when it builds the webgpu session, so set it to high-performance; a browser with nothing to honour the preference with still falls back to the threaded wasm path exactly as before.
This commit is contained in:
@@ -51,6 +51,9 @@ function load(): Promise<Loaded> {
|
||||
ort.env.wasm.numThreads = self.crossOriginIsolated
|
||||
? Math.min(8, navigator.hardwareConcurrency || 1)
|
||||
: 1;
|
||||
// A laptop with two GPUs would otherwise hand this to the one built into
|
||||
// the processor. The model is the export, so ask for the fast one.
|
||||
ort.env.webgpu.powerPreference = 'high-performance';
|
||||
const session = await ort.InferenceSession.create(MODEL_URL, { executionProviders: ['webgpu'] }).catch(() =>
|
||||
ort.InferenceSession.create(MODEL_URL, { executionProviders: ['wasm'] })
|
||||
);
|
||||
|
||||
Reference in New Issue
Block a user