It's hard to read Gwern's accompanying article without seeing this for what it is, which is a kind of mania.
I'm sure he means well and is genuine in his aspirations, but what's outlined for GA is framing LLM's as quasi-gods, which they absolutely are not. I wish him the best, and look forward to being proven wrong.
An LLM just solved 10 top tier math problems this week. It seems very likely that in 6 months they will solve 100 top tier math problems, maybe 1 millenium math problem, maybe 10 top tier physics problems, etc.
Math is verifiable. LLMs will continue to make strides in such search spaces where all they need to do is try->verify->repeat. You can't train an LLM on how to end the war in Iran or Ukraine, nor how to solve hunger, nor climate change, etc. Real issues.
Three years ago, people were saying "LLMs just generate plausible-sounding text, they don't understand the notion of truth so they can't do verifiable work like math proofs."
They still can't. But very smart humans constructed ways to use the monkeys with typewriters (with a statistical advantage) to find correct answers to problems where they already knew how to verify the answer.
You need a problem where you both know what the solution looks like or can otherwise very quickly and efficiently determine that a solution is correct, but at the same time can't work out a correct solution with a similar amount of effort/time/cost as it took to determine how to verify a solution.
Most problems don't match that criteria. You usually either have a problem with a known method of solving, or you have a problem with no clear way of verifying the solution besides the act of finding the solution itself which would involve in some way proving it is correct, or you have a problem where verifying a solution takes a very long time or has a high cost or even can't be done more than once, so you need to try to determine the best solution without being able to actually test or verify.
Basically all problems just don't fit the "hard to solve but easy to verify" criteria to a degree that makes llms a good fit. On the other hand, there are so many problems that even a tiny fraction is a relatively large number.
No, he says it's a different category of problem, and ability to solve one doesn't carry over to the other and it doesn't imply intelligent understanding.
Right, which was true at the time. So hundreds of billions of dollars have been poured into making LLMs better at these tasks via pretraining, RL, RLHF, post training, etc. again all with something verifiable in the loop. In order to improve the thing in the loop, the loop itself needs to be verifiable.
There have only been a few thousand wars, and they’re all different and all different in the world in which they occurred. The dimensionality is absurd, which is not a problem for LLMs if there’s enough data, but in this case there isn’t.
Yeah but... Both things are true. LLMs are doing things we would have considered absolute AGI god tier magic 5 years ago. They also still can't count to 100 yet.
If you think all they need to do is "just that" (try, verify, repeat as you say) then I think you're very, very mistaken.
There is no way for the LLM to bruteforce the search space any better than a human. What it can do better, tho, is to make connections between seemingly (for us) unconnected notions and join them, then verify if that's right.
Your view is not only wrong but also condescending in this day and age.
That’s literally how AlphaEvolve works. And presumably how OpenAI’s research loop works if they’d ever publish it.
They put an LLM in a loop, with a reward function, and keep trying to get a higher score. The reward function for AlphaEvolve is “did this code get a better score or faster”, for math research it’s “did it write a LEAN proof”.
I’m open to hearing that I’m wrong but I don’t think I am. I agree that LLMs make connections that humans wouldn’t, but eventually you need to verify those because otherwise the connection it made may as well be a lie unless proven otherwise.
Claude had some pretty good ideas about how to end the war in Iran when I asked it a while ago (The US should
sue for peace, I believe was the conclusion).
>An LLM just solved 10 top tier math problems this week
and 30 years ago a computer beat Gary Kasparov, if you'd listened to Hans Moravec you wouldn't be surprised that the first thing that gets automated is intellectual domain expertise.
Things are going to get crazy when it can figure out how to walk into a random house a and brew a cup of coffee, not do math
Suppose LLM solved all math problems. So what? It’s not like diplomacy and war will end. Are humanity’s problems really constrained on our intellectual ability? If anything most evidence points to cultural failing and all AI will do is enable unprecedented oppression due to said skills at intellectual tasks.
As you read this some poor person is starving. Humanity already possesses the ability to identify said person and send them aid. How is AI going to help here?
The ultimate fallacy is that all technological progress will benefit mankind. That will be true until it is not.
It's an interesting time when the bar has moved to "It’s not like diplomacy and war will end". But that is on the table. Nuclear weapons caused the end of wars as we know them. Great powers no longer directly attack each other. And AI is much more powerful than nuclear weapons, with an equal capacity for damage and a much greater capacity for good.
By the way, I donate monthly to GiveWell and the shrimp welfare project. Do you donate monthly to starving people? Most people don't, and the reason is simple: at the end of the day, most people just don't care that much about helping a starving person far away. They also don't care much about things like that, animals in factory farms, or earthquakes that kill hundreds of thousands. But ideally, we could find a way to empower people such that the minority who do care can make a big difference.
How is AI "more powerful than nuclear weapons" etc when they're two completely different things? Please explain to me how AI can glass the planet via mutually assured destruction.
Someone who would not otherwise be able to asks AI for simple instructions for how to make a deadly virus, which then spreads rapidly and kills most of those it infects
this comment will be killed, i find it an apt response to the idea that enabling those who believe themselves to be good is more important than reducing wanton suffering
i read your article you posted here. in that article you state that you support live and let live. most people who support live and let live either havent been raped by life yet, or have chosen to double down post-rape. the latter is respectable, still im not currently interested in a discussion involving the influence of such a perspective. out of good faith i highlight that i responded to:
> ideally, we could find a way to empower people such that the minority who do care can make a big difference
is it a coincidence, that in your ideal world, you would be empowered, because you care, as evidenced by your testimony of things you do that make you think you care? in my ideal world, wanton suffering would be minimized, regardless as to the path it took to get there; requiring 1 specific path involving people who think they are good, demarcates our ideals.
> Suppose LLM solved all math problems. So what? It’s not like diplomacy and war will end.
These are two vastly different statements, and the the war and diplomacy end is almost trivial by comparison, it would absolutely be solved before we solve all math in even the most steelmanned version.
Not OP but it implies all efforts are not intentional but at best by chance. For instance, the state of freely available LLM chatbots still makes me cry by the amount of crap they spew out confidently, which is not intentional. (tiny bit of sarcasm, but only a tiny bit)
> is framing LLM's as quasi-gods, which they absolutely are not
Always such certainty in these dismissals. And whenever you dig in, it always seems to be "well my super-human coding assistant made this dumb mistake!", along with a heavy dose of https://mastodon.social/@falseknees/116477394235883790
I don't think anyone who has spent any time around me would describe me as having anything but a pretty clear-eyed view of the current abilities of LLMs, I don't think anyone would seriously consider me to be in a period of mania or psychosis, and it also seems almost inevitable to me that we are a year or two away from LLM-cognition being substantially more competent in almost every regard from human cognition, and therefore in the foothills of the singularity.
The extraordinary claim that requires extraordinary proof is that this will somehow all blow over, not that we're heading towards a singularity here.
Honest question: What _would_ convince you that we are not heading towards a singularity? Like, is this belief falsifiable?
(Secondary question: what do you mean with singularity?)
As an offering of me engaging in good faith, a controversial opinion of mine: I legit think arpanet going online and starting the networking all of humanity into a massive coupled complex system fits the definition of singularity of "the moment after which predicting what will happen becomes hard to impossible", although that of course heavily depends on your definitions of prediction and hard/impossible.
For falsification, I would take a year in which we don’t see an absolute sea change in capability. The last twelve months were not that, in my opinion, nor were the twelve before it. Hell, I’d even throw in “a year in which capabilities obviously improve but at a cost consumers can’t afford”.
“What LLMs can’t do, only humans can” feels like a God of the Gaps situation: people have to keep moving the goal posts because the more recent models keep unlocking more territory that was previously a “well they’ll never be able to do this!” holdout
That's the best answer I've seen to that question so far, my honest respect. I am much more skeptical than you, for me I have seen a sea change in _tool capabilities_ (mainly pre opus 4.5 to post 4.5) and harness engineering but no strong change in the type of errors made and the pattern of harness engineering (the pattern of "set things up for the LLM to see when it fucks up and let it flail till the verifier tells it to stop").
I would actually expect the sea changes as you describe it in your first criteria to continue with 1) vision, audio and video natively integrated 2) continued scaling of e2e rlvf for workflows with large scale labeling efforts 3) ASICs and widescale deployment of diffusion models leading to speed ups
But as of right now, I still expect these models to need humans to prune the output to the gold and set up the harness right for both the novel bits, and for the boilerplate to be cohesive with the global intent.
Which is of course an amazing potential boost in productivity, but still a sigmoid flattening.
As for your second criteria that includes cost, I think we might every well see this coming soon, but it's difficult to estimate with the efficiency gains still possible.
> I don't think anyone who has spent any time around me would describe me as having anything but a pretty clear-eyed view of the current abilities of LLMs
You are literally arguing that LLMs are quasi-gods in the same comment you claim a 'clear-eyed view' of current LLMs.
Every bubble has its believers; that's why we have bubbles. There have been so many times somebody very smart has thought we were on the verge of the rapture, the coming utopia, the workers' paradise, power too cheap to meter, flying cars, settling the moon.
Could this bubble finally be the one? Sure. You could finally be the guy who is vindicated. And tomorrow the people I grew up around could be vindicated by Jesus's return to earth.
I can't prove either of you wrong, because you're not making rational arguments. You've had powerful experiences that have convinced you of something. There is no point in trying to discuss it rationally. The only thing people without faith can tell the faithful is what I said 30 years ago when I got out: I guess we'll see.
Many will stay faithful. The date of the singularity will keep getting pushed back. There will be another AI winter, but some will always see summer as just around the corner.
No, it has some of the nasty trappings of a religion, like keeping women separate until you need to exploit them, but none of the good moral or creative literary parts of religions.
I think this is what is driving the hype, the mania, and what increasingly appears to be insanely high investment. It's not about the underlying technology, which is objectively impressive, even if we're not certain of the true cost or utility of it. It's about selling the dream of Olympian omnipotence to investors, implicitly promising that IPO stands for Install Planetary Overlord.[1] And who doesn't want to be on the right side of that? (That there is no 'right side' in such circumstances does not seen to occur to anyone involved in the decisionmaking.)
1. Not my coinage, but that of SF author Charlie Stross in The Jennifer Morgue
The term robot came from the Czech language in 1923. The word was coined by Czech author Karel Capek, first used in his play R.U.R. (translated as Rossum's Universal Robots).[7][8][9] The term comes from the Czech word robotník ('forced worker'), from robota 'forced labor, compulsory service, drudgery,' from robotiti 'to work, drudge', from an Old Czech source akin to Old Church Slavonic rabota (работа) 'servitude,' from rabu 'slave'.
In this startup economy? Where almost no one is getting acquired anymore and over fifty percent of global venture is locked up in dead weight and liquidity is at an all time low?
Not sure how you can feel sure anyone can get acquired even if their company had a path to profitability, much less for ones that absolutely don't.
Plenty of companies in (or adjacent to) the AI space are still getting acquired and/or finding VC funding, even as funding in other spaces is drying up.
I'm sure he means well and is genuine in his aspirations, but what's outlined for GA is framing LLM's as quasi-gods, which they absolutely are not. I wish him the best, and look forward to being proven wrong.