Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> I suspect the economics favor centralized servers, if you only look at the aggregated cost to serve X number of users' tokens.

The economics of real-time, low-latency inference of very large near-SOTA models will heavily favor a centralized setup. But if you can afford to wait for your answer - be it a day, a week, or even more at the extreme low end (or if you just stick to leaner models for your relatively quick replies) the economics start to shift in a very clear way. A slow-going local inference setup relying on cheap SSD offload does not need the high power input of a datacenter rack, and the cooling load is outright trivial - even when working on many requests in parallel, which (in a SSD offload context) is what maximizes throughput even for local inference. These are serious problems for centralized inference that will probably limit the scale at which it can be applied.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: