Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
syntaxing
16 days ago
|
parent
|
context
|
favorite
| on:
Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
With MTP? I get 25-30 TPS on a strix halo. 50+ on a M5 max should very doable. Dflash (2) will push your TG even further
Casteil
16 days ago
[–]
It's a bit deceptive to state inference speeds without mentioning the additional things you're doing to achieve them
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: