Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It's certainly possible to write LLM inference on the GPU in a deterministic way, but it's somewhat nontrivial and trades off against performance, so by default most LLM inference engines aren't deterministic even at zero temperature. The classic post about that is https://thinkingmachines.ai/blog/defeating-nondeterminism-in...
 help



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: