Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The whole point of thinking is to throw more compute/tokens at a problem, so it will always add latency over non thinking modes/models. Many models do support variable thinking levels or thinking token budgets though, so you can set them to low/minimal thinking if you want only a minimal increase in latency versus no thinking.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: