Watch the video carefully. DFlash2's tool call fails on python syntax.
Usually models in this class nail things like that 1 shot, which the other side did.
I don't know the cause. It may be nothing. But I'd like to see the model doing something where its path is a bit more constrained, to help out rule out such oddities.
Only if you do greedy sampling. With probabilisitic sampling (categorical sampling), you will end up with different trajectory just “mathematically equivalent”.
Not OP, but you can influence how deterministic your LLM behaves using the temperature setting. The neural network doesn't directly output tokens, but logits which are then converted to probabilities and then a token is chosen at random, unless the temperature is 0 (i.e. greedy, we just always pick the most probable token without any randomness). All speculative decoding methods have to "commit" to a token though even when they don't know the actual logits of the full size NN yet. The question then is (and I don't know the answer): how do the common inference engines behave when the speculation landed on the most probable token, but the random choice still doesn't land on it? You can imagine that in the interest of performance as long as we stay reasonably inside the probability we just go ahead with the speculation. Not sure if thats implemented like that though.
Edit: I just looked up the math, and actually the idea of speculative decoding is done in a clever way that fully preserves the probability distribution while still maximizing the acceptance rate of draft tokens. So I would have to disagree with OP and say that no, non-greedy sampling doesn't influence the trajectories.
Yes, it doesn’t impact the probability distribution due to verifier. However, remember how you use PRNG and effectively due to the drafter is sampled from a different distribution initially, a separate rejection sampling won’t be able to recover what the “old PRNG” would choose in a “without drafter” case. Hence in my original post, it is about different trajectories you will end up with, not the correctness of each stochastic sampling.
If you assume that the RNG generates true randomness, then the two are identical. Only if you care about the determinism of the RNG (for example you want to use identical seeds and get the exact same generation between the two) it makes a real difference.
Correct. I am trying to explain why even it is "exact", the generated text is different from the with / without DFlash2 runs, and potentially why the DFlash2 run will contain the invalid Python syntax.
If the underlying probability distributions are the same, then DFlash can lead to an invalid Python Syntax iif the autoregressive process could have generated one if the random sampling picked a different token.
If a model can output a “wrong” sequence with a certain probability p, then Dflash can also output the wrong sequence with the same probability. They wouldn't necessarily produce the same output from the same seed, but speculative decoding shouldn't be able to produce anything that the autoregressive model couldn't also produce when using a different seed.
It will contain the same or different syntax with or without it. Also multiple runs without DFlash2 will contain the same or different syntax. And multiple runs with DFlash2 will have the same or different syntax with the same probability. DFlash2 literally has no influence (unless buggy). The difference is purely caused by the randomness.
Extremely large 1 bit models are usually within 50-60% of KV divergence to lossless models. In this case I think the comparison to Opus 4.5 is a fair assessment.
Extremely large models don't suffer as much from quantization due to its weight topology also contains encoded information, so the loss of info from any one weight is somewhat mitigated.
KL divergence (you misspelled it) doesn't tell you anything about capability drop - how much did this particular benchmark (thus ranking among models) change after 10% or 50% KL divergence?
Not only that, but KL divergence is not a '%'. It's just a number ranging from 0 to infinity that tells you the 'distance' between two probability distributions.
Any one weight, but all of them. And also crushing the architecture itself?
I wouldn't pick up 400gb of hardware to run in that mode. I might try it for fun, but even then you are looking at handling a 95GB active parameter set.
This is NOT a model for most home labs. I'm sure some can and will use it. But most, should steer clear.
Michigan also dropped their mandate for ACT / SAT to obtain admission. So, yes, I would say that's pretty broken. Combined with systematic grade inflation American universities as a whole are removing any real feedback about student preparation, level of effort, or ability.
ACT/SAT is still needed for merit based scholarships though so if you fall in the income bracket between homeless and owning your own island those tests are very very important.
The UC system and many other of the top universities in the country have also stopped looking at SAT scores.
Growing up with no one in my family having gone to college I had no idea how any of this stuff worked or anybody to tell me its important. Eventually I moved to a middle class suburb and saw kids hire private tutors for it. Meanwhile I remember 2 days before I had mine scheduled and had to find someone to beg to drive me there. I'm quite thankful that universities acknowledge how silly it is to assess a student's aptitudes from a single test score like that.
I think personal essays and a more holistic assessment is a great alternative but I'm worried about the LLM age and the limits it introduces to that as well. Perhaps interviews are the future. It would certainly better prepare students for the real world
It's much cheaper for a smart student to take the SAT than all the alternative ways to stand out. Unless the school just goes by essay, which as we've learned is either lottery or codeword for discrimination, and then they get students who can't deal with college.
Caltech, too. It was the entire first year when I was there but they have since changed to to the first two terms then normal grades for third term.
However, Caltech and MIT have a reason for this that probably doesn't apply to a large (35000 undergraduate) public university like UM.
Caltech, MIT, and UM all have incoming classes that did very well in high school. Around 88% of the UM class was in the top 10% of high school. It is 96+% for Caltech and MIT.
At all three most students will be going from having been one of the top students at every school they have attended to a school where they are more average. This drop on average will be bigger at Caltech and MIT because their students are clustered more tightly toward the top of their high school class.
Another thing that makes the drop bigger at Caltech (and MIT?) is that UM has a way more flexible schedule. For example if you take physics in your first year, they have I believe three different courses for that. One if for people majoring in life sciences and emphasizes the physics most useful in those areas. Another is for engineers and others doing practical applications. It is more rigorous. The third is for physics majors.
Compare to Caltech. First, there is no "if you take physics". It is a first year requirement. Second, there is just one course. Go to Caltech to major in English (which does occasionally happen! [1]) and your first year you are taking the same required physics and math courses that people there for physics and math are taking. Third, because admissions specifically looked for people with very strong math/science ability, the professors don't have to slow down. Those courses cram a lot into a short time compared to even the courses meant for people majoring in their subject at less specialized schools.
The net effect is that for a top high school student Caltech and MIT are much bigger academic shocks than UM because even though at UM you can find comparable courses structural differences due to its size and its much bigger number of students mean it doesn't make you take a bunch of rigorous courses right from the start at an insane pace.
[1] It is rare. The reasons some have done it include:
• They intend to be science writers or science journalists.
• They intend to go to medical school. Medical school admissions committees like English majors and they like a hard science background.
• They intend to go to law school. An English major with a serious hard science education and a law degree is a good fit for patent law or for going into science policy work.
Coincidentally just last week a relative recounted a conversation with someone who used to be a student counselor (or psychiatrist?) at MIT who said MIT should be even softer on students. WTF, right? The dynamic, the counselor explained, is that MIT is already extremely hyper-competitive to get into and so it already selects for students who are very self-driven and competitive. Hence they tend to take on a much harder workload than they can handle, especially because MIT courses are no joke either no matter how smart you are.
This causes very high levels of stress amongst a large portion of the student body. The counselor hence wished MIT had even softer policies to make it easier for students to walk away from courses with no repercussions.
If the sandbox has vulnerabilities, which you can also use the AI to fuzz for. Obviously at the point in which it can talk to the internet it doesn't really matter, but there are a very finite number of zero-days that can exist in a bytecode interpreter hosting a harness.
Well it turns out that we have yet another sandbox escape just released today called "Zapscape".
My point is if an agent recited how to find one in its memory or training set and it is air-gapped, the chances of it spreading and infecting other computers is pretty low.
Gonna have to point out that's a KVM CVE. I was very specific about using a bytecode interpreter.
If you're serious about a secure sandbox, you don't touch hardware virtualization with a 10 foot pole. In fact, you don't even use an emulator that lowers code into native machine code like QEMU. The standard for secure sandboxes is Bochs: https://github.com/bochs-emu/Bochs
Not that Bochs is perfect, a new CVE was discovered back in June. But that's the 5th CVE it's had in its lifetime, and it has a much smaller upper bound on possible CVEs compared to something like KVM or QEMU.
The reason why you use something like this isn't just for the security you get out of it, but also the deep introspection and analysis facilities you get out of it as well. Unless you're a very well funded lab, it's actually quite hard to do analysis on bare metal when you can't trust your own kernel. You can always airgap the host machine (and good defense in depth does), but that's still not an appropriate sandbox by itself, even if it's theoretically secure.
I'd made the shift to HTTP/Stateless from MCP a few months ago. It's the right thing to do IMHO. Reliability up, problems down. TOON support is natural, if desired, etc.
My only question is how do you handle channels in the architecture now. From what I saw in Claude Code, shifting to a totally http world has some timeout issues if a server drops out and comes back. Because of that I'm stuck writing stubs for my internal use MCP, this is fine for me, but if you are cleaning up semantics: Understanding how we expect clients to act around failure would really help, the story.
> edit: on bigger tests, got it to loop pretty easily unfortunately, probably local settings.
Been playing around for a few hours with the poolside/Laguna-S-2.1-NVFP4 + poolside/Laguna-S-2.1-DFlash-NVFP4 + vLLM, been seeing the same behaviour. Usually new model releases are plagued with issues at release though, best to wait 1-2 weeks then retry, or better yet, investigate yourself :) Personally I haven't found any obvious issues.
Update: Seems quite literally they have bugs on the hardware I'm trying to run this with:
From Poolside CEO Eiso Kant on Twitter:
> Learning we have some bugs on the RTX6000. We’re on it. Team has worked non stop last days and it’s getting late for a lot of the inference folks, so might be until tomorrow till we have a solution. - https://x.com/eisokant/status/2079693050796785720
Update2: I'm now running poolside/Laguna-S-2.1-NVFP4 with vLLM 0.23.1rc1.dev1378+gd6dbdb9b0 (FlashInfer 0.6.14) and seeing slightly better results in regards to the looping. I can't see any specific changes that would affect this though, strangely enough.
> The BF16 checkpoint doesn't exhibit any of the complaints [...] with our current quants, the models tend to choose the wrong logits sometimes [...] why we'll need a requant [...] We're not aware of any bugs in any runtimes themselves [...] We have two remaining things that we're trying to tackle: some people are reporting thinking being too hard to trigger (^^) , and others are saying it thinks too much. We've seen much more of the latter internally
Deleted earlier, didn't see you post, pasting here:
/* Just started testing with the gguf (with gpu offload, m5 max 128gb), q4_k_m, running seemingly well. Speed is initially slightly faster than antirez/ds4 - decode tok/s in the 30s, prefill ~400-ish. Expected, given the slightly smaller size. Looks to be working fine, but too early to tell. Definitely likes to "think". */
Anyways, guessing that discord might be focusing on the nvfp4 stuff. I've noticed spelling mistakes in the thinking traces, tool calls have been fine so far.
Well, just ran said gguf on the GeneralsX codebase with a pretty open-ended "Explain this codebase to me, and the general game loop." prompt, and...
Let me also look at the GameLogic::update() to see the rest of the update flow, especially the object update loop.
Actually, I think I have enough information now. Let me also check the GameLogic::UPDATE to understand the full update flow.
...repeating forever.
Since they mentioned that they're working on new quants, guess I'll wait. From earlier tests on work stuff, it's definitely capable.
More updates from the Discord (also mostly about the NVFP4):
> we've got an NVFP4 checkpoint that has a KD of 0.135 relative to the BF16 checkpoint (for W4A4, on some GSM8K prompts). [...] We're still seeing some looping on W4A4 (although much less than before), whereas for W4A16 we've not been able to make it loop. Given that W4A4 is still a little broken, we won't make this an official release (we have one more idea planned to fix that). [...] The RC1 for NVFP4 is reachable as poolside/Laguna-S-2.1-NVFP4 @ RC1
...they'd messed something up in that quant, apparently fixed promptly. ds4 now supports it as well and it's... well, context-limited (<=250k on a 128gb machine) but it positively flies on an m5 max - the 60tok/s decode / ~500tok/s prefill is real.
Running deepseek flash on something locally now, this will have to wait a bit. I still stand by my initial quick assessment - looks capable. Some people on r/localllama also reported loops. We'll see in ~10 hours. Hopefully I haven't terribly mislead people.
Might wanna try it now, seems to have been largely fixed. Check huggingface threads and reddit. Works for me, very memory hungry and PP speed drops off a cliff around 200k context (~30tok/s decode and ~40tok/s pp - like... hope it isn't debugging with big logs), but it brings a fresh perspective alongside ds4-flash. The more, the merrier - I prefer keen eyeballs more than speed anyways.
From my mortal spot, I'm sad I'll never get to have another.
Rice made a good one in 'ya, though I suspect anywhere could have forged you well. You were made of the right stuff.
God speed Steve.