Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I think WebGPU is mostly for running inside the browser. If one has the option to use a cloud container + GPU, running LLM inference directly with CUDA/ROCm/TPU will be possible and runs more efficiently.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: