Hacker Newsnew | past | comments | ask | show | jobs | submit | Schlagbohrer's commentslogin

I hold out hope that a multipolar world will be more peaceful and stable, with no one national ruling class able to dominate other nations so easily, and developing countries having more options for whom to ally with to their own benefit.

Unfortunately the UN, the organization that was created exactly for that, is just a puppet of the current hegemon because the money printer can buy anything and any soul. So a hard reset is needed in world politics in order to find a peaceful way again. It baffles me how China and Russia don't get together and form a new UN, like AN (Allied Nations) then both organizations UN and AN would fight for their members rights and get into agreements that would benefit all

Wow... I want an episode of "Well There's Your Problem" to cover this now.

Reminiscent of the Therac-25 scandal. https://en.wikipedia.org/wiki/Therac-25


I feel weird that I like your typos, because clearly AI did not write your post. Typos have become downright charming and nostalgic for me.

Ahahah, thank you.

Yes, I'm not a native English speaker, and I'm definitely not an AI. ;)

I can't say the same thing because I just don't notice typos (mine or others).j I just assume my English is bad.

Isn't the definition of "being human" is "not to be perfect"? In French, it kinda is. We say "He/she humain after all" to mean that someone made a mistake.


Ahahah :'D

Yes... I'm not english native. And definitively not an AI.


Absolute treasure of a website with these graphs, thank you for sharing this. Huge help for me to find a faster model (smaller quantization) for my VRAM.

Many people who go enjoy the opportunity to get away from family

They don't see security cameras because there are very, very few in Germany and the few that exist are in train stations, or private areas like INSIDE a business, or INSIDE a private estate, pointing inwards since they're seldom allowed to film outwards

And there is a big warning if there is a security camera around. And they get destroyed pretty often.

Source: living in Berlin.


Me too, LLMs are bad at generating 3D models still but that and EE/ME CAD are the next frontiers.

Maybe they could try conquering GUI layouts and CSS first

That table assumes cache hit rate of 95% or better. Am I understanding this correctly that people really are doing such repetitive prompts (compared to each other, across the concurrent user base at that time) that only 5% or less need actually be computed by the intended LLM?

That is shocking. Is it per-token I wonder?


Every tool call is essentially entire prompt so far sent again with the response and that's why cache rates are so high for agentic workloads. This really bites when using expensive models since most models are 1/10 for cached input.

We have been running a lot of agentic benchmarks with the various loops and tool calls on longer threads - we routinely see 90%+

Just checking now: recent runs tau3[1] was at 96% and toolathlon[2] was at 90%

[1] https://www.induction.ai/docs/benchmarks/tau3 [2] https://www.induction.ai/docs/benchmarks/toolathlon


If you are using their coding plan for coding, then yes you can easily hit such cache rates, with a good harness.

I’m getting 97%.


I am surprised at the lack of open-weights models in the >35B, but <200B range. I keep thinking about devices like the NVIDIA Spark and AMD Ryzen Halo, which have their 128GB of combined memory, but there are so few models made for that range. Nearly all the open weights distillations are for larger customer bases with <24GB VRAM.

Qwen3.8 Flash Next just released which hits that range.

Also, Deepseek V4 Flash can be run relatively well in hybrid 2-bit quantization on 128gb devices, with way better results than you'd expect for a typical 2-bit quant.

Those are currently the 'smartest' options for that memory level.


Qwen Flash Next 3.8 … even at 3 bit quant it is very solid.

The market is too small.

Yes exactly, I am hoping that when the AMD Ryzen Halo gets wider release and more consumers have these 128GB devices, the market will justify a wider range of quantized model sizes. And then when I win the lottery and can buy one I'll have lots of nice options!

The only (still in prototype stage!) "competitor" for those GB10/Ryzen Al Max+ 395 (in my region, borderline unobtainable) systems seems to be the Xiaomi AI Cube.

The frontier AI companies have occaisionally been using their latest and greatest (and most power hungry) models to research large efficiency gains for their older, lesser models.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: