And with their actual understanding of the hardware limitations of GPUs (memory bandwidth) and the parallel work on things like cutlass (if there was ever an unportable thing :-), the coming *Dx libraries (the explosion of cuBLAS/Solver/FFT to allow kernel fusion and new in-kernel linear algebra shenanigans) the slow but steady introduction of sparsity everywhere, I can't see how anyone can but play catch-up.