1. It uses non-idiomatic terminology in several places.
2. It repeats the same finding over and over (141 flops per byte, for example), without going deeper.
3. I stopped reading about a quarter of the way through because it felt like it was never going to stop teasing me about what it was going to tell me and actually tell me it.
4. It seems to assume the reader has a lot of context that isn't explicitly laid out (and which the reader wouldn't get just from reading the prior work, which is cited).
For example, I understand some of what it is saying because I used some similar techniques to benchmark things in the past (running at multiple scales to estimate overhead + marginal gains with a linear regression), but I wouldn't expect anyone who hasn't personally done that to follow the prose.
> 4. It seems to assume the reader has a lot of context that isn't explicitly laid out (and which the reader wouldn't get just from reading the prior work, which is cited
I've had this complaint well before LLMs were used. People writing about topics they have a lot of knowledge in the subject tend to make the assumption only other subject knowledgeable readers will read it. Or that it never edited by a real editor that would enforce rules like spelling out acronyms on first use. Or forcing additional information when too many details have been left out on the assumption it would already be known.
There's plenty of this type of writing to have trained the bots that way
You obviously haven't read it, because it is clunky garbage.
> 19.4 Pacing compiles after a failure
> A failed compile is not free of side effects on the shared compile service. A compile that fails restarts the
service, which takes a few seconds to come back, and failures that keep arriving faster than the service can
restart between them keep it from making progress, so unrelated compiles slow down until the failures stop.
The effect is a function of how fast failures arrive, not how many occur: failures spaced out past the restart
interval cause no degradation at all. On detecting a failed compile, wait at least one restart interval, roughly
15 seconds, before the next compile, so a burst of failures cannot accumulate. No hard failure-count cap is
needed.
The whole document is less nutritious than a wonderbread miracle whip sandwich.
Personally I'm not in the habit of printing and eating articles I read, but in the unlikely event that I did I find it even less likely that I would be concerned with its' nutritional content. (/s)
The idea that an LLM can discern intent on any given prompt is farcical. I might be researching nukes to commit an atrocity, or to prevent one. I might be asking about laundering money to commit a crime, or to prevent one. I might be researching the Nazis because I want to commit a genocide, or I want to read up so I know how to prevent one. Same with cybersecurity. Same with anything.
In my opinion, these companies should put their effort elsewhere. Obviously if all someone is doing on their platform is looking up how to build a nuke, where to buy uranium, the best city to explode it in, etc. please report them to the authorities. If someone is clearly just using LLMs to write hate speech they go post on the internet, ban them. And so on.
This cat & mouse game trying to have LLMs police inquiries is ridiculous to me.
> The idea that an LLM can discern intent on any given prompt is farcical.
Yes, and: the LLM is a "brain in a jar". It doesn't have any ability to verify ground truths outside itself, other than maybe calling out over the internet. Therefore it is easy for humans to lie to. You could call this an "Ender's game" attack, after the book in which a hyperintelligent kid is playing "war games" that end up being the real war.
I don't really agree with it but the government is moving towards making you ID yourself to use frontier AI - i.e. only US citizens are going to be able to use Claude Fable supposedly. In that regime the AI companies would in fact know if you are a money laundering expert or a normal software engineer.
> The idea that an LLM can discern intent on any given prompt is farcical.
Not really though. For most people in most situations it's just not going to give you that info. Software security is a niche where its a bit strange in that there is 100X the amount of white hat users than bad actors and there's open source etc.
The idea that checking for a US ID could possibly stop actual foreign bad actors from using it is also farcical. Millions of stolen identity documents can be bought on the dark web for relatively cheap. North Koreans have been hiring real American citizens for years to infiltrate tons of US tech companies as employees.
And ya, it's pretty easy to hide your intent once you have access.
I think your really anchored on anyone successfully breaking restrictions means any restriction is impossible. So your starting from the position that if it is possible for any actor in the world to get past a restriction, then the whole restriction is a farce.
KYC for example does stop most money laundering and financial crime. The most resourced actors like governments/ cartels often find ways around and it is a game of cat and mouse. Normal citizens don't really stand a chance to get around most of them.
Like it feels like your logic is that we shouldn't do background checks for employment because North Korean spy agencies get past them sometimes?
Hiring an employee, and to a lesser extent opening a bank account, are much higher-touch processes than taking on new users for your massive-scale internet app. With bank accounts and KYC, transactions can be reversed, traced, frozen, etc. after the fact. You can't "take back" API responses the same way.
Clearly, there's no such thing as a perfect exclusion rule at any of these scales, but the false-negative to false-positive ratio seems like it will be way higher if Anthropic starts trying to verify IDs.
> I might be asking about laundering money to commit a crime, or to prevent one.
Or, much more likely, the same pattern of tokens happen to exist in a completely different discussion, either as a direct metaphor, or as a reality of linguistics. Hell, "laundering" itself is a metaphorical word.
The absurd notion is that any speech should be policed in the first place. If there really is such a thing as dangerous information, then it must be removed from the training data. Any other strategy simply launders the risk.
at their scale they could also just run a large on-premise or rented (basically still cloud, but cheaper) GPU cluster and run through that. fixed costs, even license a SOTA model’s weights if you’d like
The problem isn't really Uber, Microsoft or Nvidia, it's all the smaller none IT companies that also have developers on staff. They are screwed. $1500 per seat per month is just way to expensive, but they also can't afford to build and maintain their own on-premise solution. If Microsoft can't afford to run CoPilot for their own developer, what chance does any of their customers stand?
If the large, well founded IT companies in the world believes the current AI cost is to high, then Anthropic, OpenAI and CoPilot have no actual customer base. AI is then relegated to very profitable niche business, but that can't fund the R&D for the models.
So on the lower end that's (1500 USD ~ 1300 EUR) close to half the total expenses of such a developer, on the high end here around 15-20%. That's quite significant, depends on whether their productivity also improves (if that's what the orgs care about).
And we’re not even the country with the worst pay out there, but pay the same for tokens, cause regional pricing isn’t a thing!
I wonder how this plays out. Perhaps programmers in these countries will use cheaper models like Deepseek and they will be able to compete better, so offshoring continues?
> Perhaps programmers in these countries will use cheaper models like Deepseek and they will be able to compete better, so offshoring continues?
Even here, companies don't really trust Eastern providers that much, so they'd be looking for someone in the EU running DeepSeek instances, which might come with a bit of markup. Those orgs would also sometimes be weary of OpenRouter which to me seems like shooting yourself in the foot by being so picky.
That said, DeepSeek V4 Pro (with Max reasoning) is pretty okay and I'm using it instead of Opus 4.8 (my Max 100 USD subscription weekly limits ran out today) and it can do stuff passably (even better than Mistral's offering and has nice context window), but compared to the amount of work I can get done with Anthropic's models, it keeps occasionally fucking up and I have to go back and correct it, so lots of token waste. Maybe it's close to SOTA from 6-12 months ago, though, which is pretty cool on its own, though - just less confidence in its output.
It's like trying to limit the costs and therefore not gaining the maximum added value from the technology. Similarly for those trying to run stuff on-prem, we don't really have the electrical grid here for large scale inference in-country, nor is anyone exactly salivating at the idea of dropping multiple tens of thousands of EUR to build out something passable. I do host some stuff on a bunch of Nvidia L4 cards (Qwen3.6 35B A3B) and while the model has its uses, it's also a far cry from SOTA.
So I guess it depends - compared to an Anthropic subscription it kinda sucks, but then again if you have to pay for Anthropic's tokens those are robbery and then DeepSeek looks like a no-brainer alternative.
In hcol locations yes, but in south of spain you can get full time talent for that figure. It's also an entry-level salary in eastern europe, with ukraine and turkey even being somewhat cheaper.
Why are smaller non-IT companies "screwed" because they can't pay out the nose for their developers' AI usage? They're non-IT companies, developers are presumably not on their critical path, or not their bottleneck. Developers can keep on writing code the old way, or doing it with a more reasonable AI spend. I don't see how this "screws" any company.
> even license a SOTA model’s weights if you’d like
Yeah, I bet all labs releasing SOTA models are more than happy to remove the main way they make money and let you run it locally, especially if you're a big spender like Uber who seems very willing to throw money into the sea as an experiment.
Anthropic and OpenAI license to the public clouds. Google reportedly licenses to Apple. licensing to Fortune 100 companies running on their own infra is an obvious next step
it is a race to the bottom and I’m not sure the labs win that race. we’ll see!
I'm not sure the labs will win either. I wouldn't be surprised to see OpenAI & Anthropic just get acquired, either by Microsoft or Amazon and their models just become another product offering in their public cloud and and some hybrid on-prem offering like Azure Stack HCI or Azure Stack Hub (already basically a "cloud in a black box" that could become "AI in a box")
I believe Databricks series L round raised $4B in late 2025, but earlier this year they raised another $5B so technically they've maybe completed series M round and are "on" series N round now? The press releases are a bit confusing to me.
It's semantics, but the latest raise might have been a follow-on to Series M, not a new round (to be clear, I know nothing about their finances, just speaking from experience at another company).
all are by the DuckDB team except three third-party owners. I’m unfamiliar with Vortex, but presume it’s like LanceDB and MotherDuck with a serious company behind it. and presumably the DuckDB team trusts them not to ship malware in their extension
Thanks for the link. Good to know that they are at least signed by a key. But I really like my software not changing on me at all. I'd rather have all of the modules I need locally and static.
Also creates fun situations like getting on a plane then realizing that your extension isn't available!
It seems that nixpkgs at least fails to run the extension but more by luck than design. I hope they find a way to vendor the extensions locally.
Should be compulsory reading. Actually, now that I think about it, this would make a great interview question in the AI era: “what did you think of Soul of a new machine?”
Those two books are probably the two best about tech projects I've ever read. I worked at Data General as a product manager for about 13 years and know many of the individuals although I joined a few years after the book was written.
What strikes me is the stories that never get told. I met a retiree at a Java meetup once who had worked at Zilog during the z8000 era. He was surprised to meet someone who knew about that.
Especially, pre-web and pre-blogs there's a great deal of tech industry history that largely doesn't exist any longer unless it was especially notable and/or some author decided to spend a year or two writing about it.
Honestly, when I read this book in 1982 or so it changed the trajectory of my life and career. It is an story of a bet-the-business project that occurred in real life. After reading I though I was late to the tech party, missed out on so much, and yet, here we are, 45 years later with incredible advancements and a career I couldn't have imagined. Tracy Kidder was a gem of an individual, just a wonderful person who was truly interested in others. I hope you'll read the book soon.