I hold out hope that a multipolar world will be more peaceful and stable, with no one national ruling class able to dominate other nations so easily, and developing countries having more options for whom to ally with to their own benefit.
Unfortunately the UN, the organization that was created exactly for that, is just a puppet of the current hegemon because the money printer can buy anything and any soul. So a hard reset is needed in world politics in order to find a peaceful way again. It baffles me how China and Russia don't get together and form a new UN, like AN (Allied Nations) then both organizations UN and AN would fight for their members rights and get into agreements that would benefit all
Yes, I'm not a native English speaker, and I'm definitely not an AI. ;)
I can't say the same thing because I just don't notice typos (mine or others).j I just assume my English is bad.
Isn't the definition of "being human" is "not to be perfect"?
In French, it kinda is. We say "He/she humain after all" to mean that someone made a mistake.
Absolute treasure of a website with these graphs, thank you for sharing this. Huge help for me to find a faster model (smaller quantization) for my VRAM.
They don't see security cameras because there are very, very few in Germany and the few that exist are in train stations, or private areas like INSIDE a business, or INSIDE a private estate, pointing inwards since they're seldom allowed to film outwards
That table assumes cache hit rate of 95% or better. Am I understanding this correctly that people really are doing such repetitive prompts (compared to each other, across the concurrent user base at that time) that only 5% or less need actually be computed by the intended LLM?
Every tool call is essentially entire prompt so far sent again with the response and that's why cache rates are so high for agentic workloads.
This really bites when using expensive models since most models are 1/10 for cached input.
I am surprised at the lack of open-weights models in the >35B, but <200B range. I keep thinking about devices like the NVIDIA Spark and AMD Ryzen Halo, which have their 128GB of combined memory, but there are so few models made for that range. Nearly all the open weights distillations are for larger customer bases with <24GB VRAM.
Qwen3.8 Flash Next just released which hits that range.
Also, Deepseek V4 Flash can be run relatively well in hybrid 2-bit quantization on 128gb devices, with way better results than you'd expect for a typical 2-bit quant.
Those are currently the 'smartest' options for that memory level.
Yes exactly, I am hoping that when the AMD Ryzen Halo gets wider release and more consumers have these 128GB devices, the market will justify a wider range of quantized model sizes. And then when I win the lottery and can buy one I'll have lots of nice options!
The only (still in prototype stage!) "competitor" for those GB10/Ryzen Al Max+ 395 (in my region, borderline unobtainable) systems seems to be the Xiaomi AI Cube.
The frontier AI companies have occaisionally been using their latest and greatest (and most power hungry) models to research large efficiency gains for their older, lesser models.
reply