Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> If this is truly AGI (subject to one's definition of AGI still)

Scoring well in a benchmark that's called AGI does not make an LLM AGI.

 help



The goalposts of AGI will shift forever. If you showed our current capabilities to someone from 2016 it would be declared AGI.

Is anyone from 2016 still alive today?

If so I'm hoping we can track them down and have them tell us if they think this is AGI.


As someone who spent countless nights tweaking Edge Detectors (looking at you, Canny), morphology operators, etc., building models to recognize 10 handwritten digits, let me tell you: the current set of LLMs (even the smaller ones) seem like magic. I had never imagined a computer would do such things in my lifetime.

Exactly people can say whatever they want, but current level of LLM is AGI level to me. It is already on par with senior programmer if the instruction/prompt is right.

Once we have 1000 tps, i am sure robots etc.. will also start working like magic.


I don’t know, I am writing a modest 30 page paper with Fable and even after rounds and rounds of feedback and improvements there are so many things that are just plain wrong or weirdly out of place or just stupidly written that Fable 5.1 doesn’t seem to have any awareness of by itself that I don’t think it’s AGI, I think a human researcher can easily outclass it in writing and problem understanding. It definitely has super human capabilities but it lacks awareness or self reflection in my opinion.

For example it should be easy to tell it to not write a paper in the style of a clickbait SEO article or use all of its stupid hallmark AI writing patterns “it’s A, not B!” And a smart human that would be told that would be easily able to comply with that but the model needs to be told in a very detailed way and it seems to lack even basic capabilities to reflect on this, when explicitly given a sentence it will be able to rewrite it but otherwise it’s mostly blind to it. That’s to me a hallmark of it being overtrained on the specific tasks or problems so it appears very smart but once you go off script it still shows that it’s not a “real” mind.

Of course it’s amazing and has super human capabilities in many areas but if you honestly think it’s better than Einstein like some people suggest why can’t it write a simple “good” academic paper even after giving it specific examples and instructions.

Maybe that’s what makes these things dangerous, they have super human capabilities in some areas but apparently lack self awareness, taste and meta reflection abilities. The only reason people aren’t afraid more is that they don’t act in the physical world yet, imagine giving it a body, superhuman strength and letting it care for your child when it has a strong “urge” to comply with your exact request and little to no self awareness and human basic instincts.


100%

"You're absolutely right to call me out on that. I shouldn't have stopped the baby crying by killing it, that's on me."


>It is already on par with senior programmer if the instruction/prompt is right.

It's magical to me as well, but I don't feel like it's AGI.

Because in my experience a Senior Programmer does not need the right prompts to deliver the right outcome! :-)


Doesn’t the G stand for “general”?

An AI model that’s human-level at programming is an incredible achievement. But it isn’t general intelligence. It’s highly specified intelligence.


Though what has a program that is really really good at edge detection have to do with AGI? The community just spent decades to perfect edge detection. That's great! But let your edge detector try to fry an egg and then tell me again that it's AGI.

I'm still here, nearly 50 years and counting. If you had asked me what I imagined AGI would look like back in the 90's, I would have told you "A system that can do everything we can: see, hear, think, do.". If you had shown me GPT-6 back then, I would have said "It looks like a really powerful program, but that's not really what I had in mind.". That's AI, but it's not quite general.

Likewise, it's very impressive and useful, but it is obviously not AGI to those of us from that era.

If anything, the fact that it is so powerful is almost a concern, because I think we are still way underestimating what these systems will be able to do when we give them more cognitive capabilities.

At the moment we are something like, having had great success with propellers and have promised we will fly to the stars.

People love to say 'this is the worse they will ever be', then extrapolate to conclusion that they will continue to accelerate at the same rate of progress of last few years .. it may, maybe, or we will hit a ceiling, might be a temporary one, could be 5 years or 50 years ..


And then you'd ask it about an area you're knowledgeable in and realise it routinely makes stupid mistakes.

Or you'd ask it to add a new page to your website and shout at it to use your existing brand colours instead of inventing some and realise it's not AGI at all...


> And then you'd ask it about an area you're knowledgeable in and realise it routinely makes stupid mistakes.

This honestly doesn’t happen to me much anymore. In what areas do you find LLMs routinely make stupid mistakes?


Origami design will be my personal test bed for the coming years.

It's objectively very difficult and technical, it's spatiovisual, it's artistic, learning resources for it are sparse and most just learn by the FAFO method, current AI sucks terribly at it, and it's not likely to ever be specifically targeted by benchmaxxers.


Scientifically useful physics simulations. Every model absolutely sucks at them.

Or, as someone else points out in another thread here, academic writing. It's one of the things newer models seem to have actually gotten worse at. Even when you give them detailed instructions on how to write and what to avoid, the "load-bearing", "A but not B" and journal-like writing make it in anyway, with the supposed AGI having no ability to reflect on how blatantly unacademic (and often unreadable) its writing is.


In the last day of coding it has:

- Created useless pydantic schemas with all fields Optional[Any]

- Created a REST endpoint that silently mutated on GET (unsubscribed users from a mailing list)

- Failed to log costs in my app so users could have bankrupted me, etc, etc.

Good job I actually review its code.


It's really bad at game design

Anyone that is not impressed by what ChatGPT or the likes are doing now is being either dishonest or is incapable of being impressed.

Only the translation and language understanding capabilities are enough to be impressed, and they are 2 year old already. Now, the AI do see, draw, speak, listen, think, work, etc.

Someone from the 90's would simply not believe that the AI would be a machine but would think for sure that a human is behind. The only odd thing would be that this human would both exhibit high intelligence and stupidity at the same time.


Compared to what we had in 2016 with RNNs, this is effectively “AGI”

OK, so its way better. that doesn't make it AGI.

If I can't give it an arbitrary task and have it solve that task eventually, it's not a general intelligence.


Are you guaranteed to solve an arbitrary task eventually?

I believe so. AIs are shockingly good at a lot of domains, but there's still a lot of pretty basic stuff they don't really "understand" at a conceptual level and (currently) they can't learn to get better at them.

(obviously it might take years for me to get good enough at something, or if you set the "arbitrary" task as something ridiculous, but lets work in good faith here and think of something the average human could do after learning about it)

If we progress to the point where an LLM instance can meaningfully learn to get better at something overtime without retraining, then I will accept that is basically AGI. Right now, they still seem to be pretty boxed into their training, even if you can prompt them to act differently.


Compared to 2016, it can do a lot of things, but it still fails for example with recommending a setup for my Raspberry Pi to have a 4G connection with some parameters (I want to use as a gateway between VoLTE calls and SMS, and my SIP server somewhere else). It failed miserably. I bought stuff according to its recommendation which was more or less a waste of money, twice. With miniscule knowledge compared to theirs, or even hobbyists', I could figure out all the details at the end, and order something which really works, but only after I sit down for 4 hours, and dig through exactly what I needed, because LLM lied flat out what Sixfab 3G/4G HAT can do. (Of course, not just LLMs lie, SixFab lied about something else too)

Of course, it's a moving goal post, because we have no clue what general intelligence is exactly. But it's definitely not general yet. Now the goalpost is to achieve that kind of level of thinking which I did in that 4 hours. When it reaches it, we will find something else it clearly lacks. Until we can't. Then, and only then we reached AGI. Until you see comments, reviews, etc about things which it cannot do, until then it's not general.


True. "AGI" has also become a marketing term. Achieving AGI has become valuable, so companies will move the AGI goalposts, over and over again, so they can achieve AGI, over and over again.

If you came at it from the perspective of imitating what the human brain does, we now have a very very powerful speech center and short term memory, and vision catching up. The other parts are missing. I‘m sure that’s being heavily researched.

If the only difference between a human and LLM is a human needing to tell LLM to try harder then I think we are already there.

If i suddenly travel to 1500s i would also be considered genius(in some way)

I bet you’d think they are not even conscious.

It seems like AGI until it does something that leaves you scratching your head. Getting a simple thing wrong.

I hear the T-rexes were still roaming the earth trying to eat us cavemen in 2016.

talking about self proclaimed, it's about as much AGI as openAI is open.

But they declared it...

“I declare bankruptcy!” - Michael Scott

I DECLARE AGI!

"Homer, you can't just declare Artifical General Intelligence; you need to like, make something or something...mmmmrrrhh"


did they?

"""

In a closed a press briefing earlier today, OpenAI co-founder and president Greg Brockman offered an unusually direct formulation of that message, ending the session with: “Welcome to the AGI era.”

"""


What test do you propose as the actual go/no-go gauge to verify if some model is or is not AGI?

If you’re talking some nonsense, silly singularity… than whatever, don’t care.

But if you’re asking when a model has a sustainable general intelligence, for me, it’s pretty easy…

When it makes financial sense to run it 24 hours a day.


I can have Astra run a large-scale infrastructure migration 24/7 (much of the time waiting for results), completing it in weeks or even months faster than I could before agentic AI.

What does it mean to run a model 24 hours a day?

Aren't we way way past that already? QPS to any of the frontier models for a given point in time is most likely (far) greater than zero.


For whom? That is a fantastically ill-defined test. Everyone here is comfortable throwing around this or that is or isn't AGI which is fun because, at the same time, nobody seems to have a testable definition.

It makes either position pointless to argue.


I mean the laundromat runs the machines pretty much 24 hours a day but a washing machine is not AGI.

Directly - something can be useful without being AGI.


There can be no such test because “AGI” is (or has become) a pseudo-philosophical/socio-political concept rather than a scientific one.

It always has been. The idea around here that we can actually define intelligence and point to it is sophomoric and incredibly frustrating.

If you’re trying to tell me this is why my mom telling me how handsome I am didn’t translate to the general populous, I could have used this info about forty years ago.

Populace

Christ, thanks. I blame the beers.

Hey now! Keep your reason out of their marketin^H^H lies!



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: