In an automated build system, I could see per device kernels, made more interesting by fusion of ops.
The area that obviously needs more attention are the 947 tests that did not result in reliable output. This could be the test, it could be that the ORT WebGPU is wrong, or the new kernel is wrong, or the test harness had issues, hard to know from the outside. Perhaps the Fleet thing is meant to improve this.
Very excited to see such attention at making WebGPU first class in ML!