Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
CuriouslyC
69 days ago
|
parent
|
context
|
favorite
| on:
Jamesob's guide to running SOTA LLMs locally
You can improve that with speculative preload. I'm sure models could be designed and tuned around efficient SSD offloading to keep throughput pretty high.
searealist
69 days ago
|
next
[–]
It would apply equally to GPU or RAM inference as those are also bandwidth constrained on decode, so people already try to optimize for it.
rsalus
69 days ago
|
prev
[–]
surely the supply of unified memory will rise to meet demand before this is needed
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: