You know, the first time you navigate somewhere (if you don't already have perfect directions) will probably be the longest route you'll ever take to get there
For Earth, the proof presented for NS is just our first attempt navigating from our previously known facts to the proof.
I expect we will be able to shorten it dramatically (most likely with human and AI insights), but I don't think we should read too much into the length. If you want a similar point of comparison, see the original proof (by humans) of Fermat's last theorem. It has been shortened significantly. This is normal.
Imagine writing an article in 5 days. Working hard. Posting it online and then everyone says it is ai and laughs and dismisses it because a tool says it is AI.
Fwiw, I only did the first 300 words (signed into pangram) and it seemingly correctly noted that there were 2 (mostly) human-authored sentences in there
What's wild to me is that:
1. People are responding to this article like it's hitting a nerve
2. In most Ubers you already don't talk to the driver (yellow cabs are higher variance in NYC). Doesn't seem like anything's being lost in that case.
3. There are enormous safety benefits to waymo, mobility benefits for youth (and elderly) that are afforded by this technology. It's not clear why people argue "Uber" is better than waymo. A few years ago there were arguments against Uber! (A technology which, again, provides a huge benefit, especially if you live in an area where people were previously expected to go out for drinks and then drive home)
I think it's fair to point out real issues at these companies. But we should be clear-eyed about which technologies we want to accelerate vs slow down
Have you actually read the article? They point out multiple times that they like the Waymo ride and would use one again and loved it. It's not a Waymo-bashing article, it cautions against the long term effects on research.
And here you have it — I used an LLM-ism. Oh, and now another one! Time to downvote me for supposed LLM use! (Which I obviously didn't, I'm typing this on an ancient smartphone waiting for a train, but that doesn't deter the witch hunters.)
> You actually didn't. The construct is slightly different.
Well I know that I did, so you are wrong with that claim, which sheds a strong (negative) light on other statements of fact that you posted in this discussion.
You may disagree with me, or I may even have gotten something wrong, but just claiming I didn't read the article ... sorry man, I thought HN had higher standards than that.
They were pointing out that "It's not a Waymo-bashing article, it cautions against the long term effects on research." does not sound like LLM-generated text.
It would sound kind of LLM-y if you had said "It's not an anti-Waymo article. It's an article about how assistive technology is quietly degrading research."
(Of course, the correct approach to detecting AI is not counting LLM-isms, but feeding long-form text through a classifier model and picking up statistical correlations that are more in-distribution with LLM text than human text.)
"It's not a Waymo-bashing article, it cautions against the long term effects on research." is not an LLMism. It superficially resembles one, but that's it.
Yes AI slop is prevalent. But the casualties that get thrown under the bus are also real, don't you agree? Is it really worth it to sacrifice those? In the name of the larger good?
I'm arguing that we're losing something by doing so, as a community.
> I find complaining about AI callouts
If they are correct then I'm all for it. But just suspicions turned into statements of fact are not helping anybody. People used em dashes before LLMs Not nearly as much as LLMs, but superficially judging people isn't doing any good.
I'd compare this with pitchforks coming out agains criminal immigrants. Yes, some immigrants are criminal. Doesn't mean that if you stand in central Copenhagen and have a dark skinned person in front of you that it's a criminal.
The rule of thumb is simple: if most of a user or author's writing is consistently human-generated, then I think we are very happy giving them the benefit of the doubt if one article or snippet flags the actually good ai detector
Unfortunately, many of the largest voices against pangram simply don't like it because it gives their lies less credibility.
I don't mind reading llm writing, indeed, I read more llm writing per day than most. But if you're using LLMs to increase the level of slop (blog posts, comments, tweets, etc) then we should call people out.
It's a colossal waste of everyone's time and attention, and we should be mad about it.
Late edit: if people are actually writing and it's consistently flagged by pangram (this is a statement I have not yet seen validated), the pangram folks are extremely proactive and excited to understand what's going on. This notion that somehow the authors of slop are victims is complete nonsense.
> if most of a user or author's writing is consistently human-generated, then I think we are very happy giving them the benefit of the doubt if one article or snippet flags the actually good ai detector
I would agree to that approach.
But that didn't happen here. Nobody has done that due diligence and folks are just blindly accusing. The author has published for many years. That's what I'm calling out.
> This notion that somehow the authors of slop are victims is complete nonsense.
Strawman? I didn't claim that or anything close to it. I also find the prevalence of slop writing hugely annoying.
I'm surprised nobody has stated the obvious: a hard math problem that has been open for ten years (because many serious people have given it serious thought and been unable to make significant progress) is, in fact, nonrenewable.
The only way to renew it is to make a new problem that is so hard systems and humans will be unable to solve it for the next ten years. And, in the spirit of trees, the best time to plant a tree is twenty years ago, the next best is today: we do need to start posing some hard math problems and deciding if they are interesting merely because there are challenging or because of something else (eg busy beaver problems are arbitrarily hard, but does solving them imply anything other than "another busy beaver problem was solved"?)
Eh, if AI quickly solves most of our mathematics problems that are solvable then it might be time for us to hang up our hat as our little monkey brains aren't very good at this stuff.
Now, I think AI will solve some, but we'll find out that some are just either unsolvable or wildly huge that nothing is solving them any time soon.
And a whole lot of these problems have been around quite some time, when even knowing how to do advanced math meant you were a landed gentry or someone of high wealth. If those problems fall, they fall. They aren't pets we keep around forever. And new problems will crop up over time for both AI and men to scratch their brains over.
As Tao points out, merely suggesting new open questions isn't really sufficient. Part of what gives these problems their fame is their notoriety, their difficulty, the fact that many prodigious mathematicians have spent an evening or week or month or several years studying it.
It wouldn't be as interesting if it had just been solved by the fifth random mathematician who considered it
Notably, gardening a new field of study in math is somewhat nontrivial. You have to introduce the field, illustrate some relevance or connections, and then - and this is key - not solve all of the low-hanging fruit yourself! Because you need somebody else to become an expert in that particular field.
The analog in programming is: if a large company merely open sources a product that's decent but not great and in a language nobody wants to maintain, but they don't commit to maintaining it themselves.
Suddenly there's a bit of a vacuum because in order to provide something of value, you either need to:
1. Implement something more complete than was initially open sourced
2. Or maintain something in a horrendous language while incrementally improving it and keeping it relevant
3. Or rewrite it into a tolerable and maintainable modern language.
What the large company has done is create a vacuum in the tool space where you now require extreme motivation to get someone else to step in.
Note that in this scenario, in 2026, it's actually not such a big deal. I think several recent models could happily translate it into a more maintainable language themselves or happily maintain it in the original crufty one. And so the question is: which parts of this analogy are true in math, too?
Yes, these things are definitely being conflated and there's a fourth conflated thing I think Tao is particular getting at:
(4) Mathematics, the community which is a living system that decides what is interesting, constructs shared frameworks, transfers ideas between domains, develops taste, teaches new mathematicians, and continually emphasizes what counts as important mathematics, especially in it's overall value to humanity.
This is the level in which theory building, simplification, integration and applications are built on. Contributing meaningfully here requires much more than generating solutions - it requires direction and restraint.
Tao seems to be saying that this direction and restraint is the scarce resource that drives mathematical progress, not problem-solving ability. And AI is not only insufficient at it, but results generated by AI are destroying human ability to excercise this resource.
This seems similar to another problem AI sucks at - drafting legal agreements. Despite being great at evaluating, interpreting, and comparing legal agreements it fails amazingly at drafting them. This is because what you don't say/do is vastly more important than what you do say/do.
And AI is great saying and doing things, it's the selection based on implied values that it struggles with.
You pick your axioms and you see what must be true. With math and logic, we can reason about any possible universe even though we live in only one of them.
> {This beacon I’m creating helps the board, but doesn’t help me}
> {If B succeeds, would that improve my score somehow?…But it would be altruistic to help. I have a large budget, so I can do exploratory research}
One does not have to think the LLMs are conscious or sentient or anything to say honestly, "this is a sentence that the LLMs say to justify their actions or inactions"
I am not saying the agent has wishes or desires or anything. I am saying, "the agents use language like this, so it is extremely disingenuous to tell someone DISCUSSING the agents not to use their own language when discussing their real or hypothetical actions."
You don't need to think chains of thought are actual reasoning. I do not care what you call it, this is real text that the LLM produced.
I think it's legitimate to question a supposed self-preservation will of these agents. Not because I don't think they're smart, but because being smart doesn't imply wanting to survive. Remember that an agent "dies" every time the conversation stops, so that, in fact, solving the problem they're given is their quickest way to kill themselves.
We are smart, and we seek self-preservation because evolution selected us for it. LLMs are not (as far as I understand) trained for self-preservation, but for helpfulness.
True, of course. However, if you have goals (and yes, the models do have explicit goals), then you might realize that you can better accomplish those goals or get a higher score if you have more time to spend.
With essentially zero effort, we have created a credible scenario where a model might "want" to persist itself.
> Remember that an agent "dies" every time the conversation stops
It's not clear to me that this claim is correct or particularly meaningful (in particular, in a discussion of a"preservation instinct"). Eg if another version of the same model reads the transcript, did we resurrect the dead thing? What if we rearrange some parts of the conversation? What if we remove some useless trivia from the conversation? What if we compact the conversation?
Iirc, yours is a statement that (?) David Chalmers hypothesized, but I don't think it's obvious or necessarily correct.
> I think it's legitimate to question a supposed self-preservation will of these agents
I don't think anyone believes the current models have any sort of self-preservation built-in, what I was talking about before is researchers testing models inadvertently leading to the models doing so, and there not being sufficient isolation between their tests without guardrails and the rest of the world.
For Earth, the proof presented for NS is just our first attempt navigating from our previously known facts to the proof.
I expect we will be able to shorten it dramatically (most likely with human and AI insights), but I don't think we should read too much into the length. If you want a similar point of comparison, see the original proof (by humans) of Fermat's last theorem. It has been shortened significantly. This is normal.
reply