There is some sense of rose-tinted glasses of pre-LLM coding. A lot of human written code, particularly at the enterprise level, was of low quality well before AI automated it.
So much "bad" enterprise code evolved into that state over years or even decades of small changes. Meanwhile, last year I got to watch an LLM-authored codebase speedrun itself into a similar state in only a couple months. And I would say that the enterprise code was actually better. It at least did its job fairly reliably. The LLM codebase was riddled with defects, so much so that it ate up all our time and our feature delivery rate ground to a halt.
There are two observations that really eat at me:
1. Studies seem to indicate that agentic coding uses 2-10x as many lines of code to accomplish the same task.
2. One of the only really well-established empirical results in software engineering is the strong association between LOC and defect rate.
This is happening all over the place right now. There is a ton of greenfield happening, which further adds to the illusion of speed. Eventually you produce a big old pile of shit that even with the help of the LLM is weird to reason about, and it slows way down. Many such cases.
Writing code at enterprise level is insanely difficult. You are constrained by budget, staff, legacy databases/environments, business rules hiding all over the place, and people.
You can't just rewrite everything. So over many years people are touching small parts of the pie.
Yeah, enterprise code has that trope of being enterprise-y, verbose and bad. In my experience, that has always been the opposite. At the big corps/FAANGs I worked at, a single line of change can adversely impact millions of paying customers, so a lot of the verbosity and harnesses exists to dampen the failure modes.
Most of the terribly written stuff has always been at startups, where devs fling nearly anything across the finish line, if it barely works the happy path.
I don't doubt that. But humans still need to be responsible for understanding what they're shipping. And IMO you get your best understanding by actually writing some code. Even if you don't actually ship what you wrote.
Let’s not romanticize it too much... A lot of enterprise systems are built by developers copying an old AbstractBeanFactoryFactory from a 2011 stack overflow thread without really understanding it :)
Nobody is. It's the AI slop which is supposed to replace this shit which barely worked with equally shit shit which doesnt work which people are romanticizing.
Most of the human written code was slop, but the really fundamental and successful stuff we relied upon and which we didnt want to throw away? yeah, not so much. most of that was actually really good.
those EJB monstrosities were routinely swapped out by some saas written in python by somebody who did it properly and werent responsible for a lot of late and over budget projects which barely worked or didnt work.
> But humans still need to be responsible for understanding what they're shipping
I don't necessarily disagree. That said...
Why?
I've been grappling with this myself. There is an easy/obvious answer, but I wonder how stable/permanent it is. If you feel strongly about this, are you willing to unpack your judgement?
Yeah... like we all get to start green field projects and write all the code we should understand. Many of us cut our teeth on bad legacy stuff with no proper documentation made by "engineers" long gone. At least a LLM can easily make sense of this mess.
Indeed. And not fair comparisons ”look at the quality of this small one-shot Claude hobby project. The quality is less than this major open source project written by some of the best developers in the world”
To be fair the pitch has frequently been that Devin/Claude/Astra/whatever is some sort of superhuman bottled John Carmack that will single-handedly replace entire teams of developers.
Yea, that extreme side exists too. Truth is inbetween. AI with the instructions from a dev that knows what it is doing writes better code than most regular 9-17 devs.
1. People didn’t wear that as a badge of honour though.
2. A lot of it wasn’t. Low quality code/speed serves a purpose for point solutions and scripts etc. That’s not the same thing as writing a core system and if the user doesn’t put any credentials in for an S3 bucket then it falls back to giving information about your own S3 bucket (as I’ve seen just this week).
3. Plenty of companies you can discern the difference between mission critical systems versus “business” systems where if it falls over it’s annoying but not the end of the world.
This was always due to pressures by management and the company environment, not the workers themselves. It's hard to blame the people writing code when they have to deal with nontechnical leadership that wants to have a feature factory or never given appropriate resources to solve problems.
Blaming workers is always an excuse by poor management.
The pressures from management and the company environment are not always a bad thing. It really depends on whether the pressures are coming from a logical business perspective or whether they are just coming from stupidity or ignorance. In a business environment, taking a long time to ship great code can mean that the company goes out of business, and then the software developers have a lot of great code and no income.
Wait. I didn't say anything about making people miserable. And I'm not talking about some weird kind of pressure like yelling at people.
I just mean the normal, almost inevitable kind of pressure that comes from the business situation. Management has to somehow figure out strategies to handle the pressure and it has to communicate these strategies and the reasoning for them to the engineers in a constructive way.
Generally, this is the pressure to be competitive and make money. It does the developers no good if they spend so much time writing great code that the company goes out of business.
Nah, you're still blaming workers and not leadership. If leadership is okay with not training workers (something American corporations would do in the distant past) then it's not fair to continue to blame workers when leadership is clearly aware of the problem and would rather pocket the money than help workers.
These companies pay management more than workers for a reason, if you can't even admit that they are to blame then what are you trying to do here? Just attack workers for what reason exactly? Being anti-worker is a great tell to never trust a person.
We're talking about professionals here. People who (at least in the US) often make several multiples of the median worker. Competence is assumed, and every company I've been at has had programs to pay for additional school if the employee wants it. IIRC at least one had an explicit book allowance, and I don't doubt that I could expense books right now if I asked. Do surgeons and lawyers complain so regularly that management doesn't train them? Or are software engineers just this desperate to be seen as "not a real professional"?
Who even is supposed to be training us? We're supposed to be the experts. Unless you mean mentorship, which is also generally already a thing at any company that has more than a handful of engineers.
There is no professional developer in the US. There are no licensing requirements to write code. There are no repercussions against developers that write code that immiserate or kill Americans. The person slinging wordpress plugins at an agency is equivalent to the person writing malware at Meta in the eyes of the government.
No one is training us because there are no regulations in our industry. Sorry but I still reject the premise of the other poster.
Also yes, in other industries professionals do complain when they aren't given their mandated time to learn while working on the job.
We are given time though. Maybe not at some small sweatshop? But at least at big companies all the jobs are salaried, and either no one cares how exactly you spend your time, or they explicitly say it's fine to spend some time learning, and will pay for materials or courses.
They're generally not going to teach you how to program or what the latest technologies are or what the best ways to do things are because that's literally what they're paying you for. But IME they definitely give you plenty of time and autonomy to be learning things at work.
I do blame management for letting these people through the interview process and then not firing them. But that’s independent of the fact they exist.
Training doesn’t solve every problem, the worst programmer I ever worked with that a PHD in computer science. Everything he made was horribly slow, wildlife overly complicated, and buggy. Worse he wouldn’t listen to anyone correcting his issues. He’d store numbers in the database as strings to be database agnostic etc.
Let me throw a curve ball at you: do you accept the premise that most modern corporations are centrally planned economies under the rulership of monarchies, oligarchies, or general authoritarians? If so, do you think introducing democracy into the workplace would help alleviate issues you care about?
You do not like bad workers, management doesn't care. They pay bad workers the same as you, bad workers can get promotions the same as you, you will also get laid off with the bad workers as well; or maybe even worse, the bad workers get promoted into management themselves. How do you want things to change in such an environment?
You have no authority to do anything meaningful as a single worker, what if you were given a voice to actually make these claims and have other workers decide what to do based on your voice?
Workplace democracy seems like an interesting concept to explore if you truly want to create better environments with beneficial outcomes to all, not just the few:
> do you accept the premise that most modern corporations are centrally planned economies under the rulership of monarchies, oligarchies, or general authoritarians?
No stock owners are ultimately in control of public companies. That doesn’t fit any of the models you just described.
Similarly companies are not independent of government control which inherently separates them from management systems associated with governments. A CEO is limited by the law in ways that a dictator isn’t.
I've only worked at companies with great leadership. This is in the Nordics with very strong worker protection. And yet most of my colleagues including myself have been pretty terrible and write dirty code. It's not a management issue and it's not anti-worker to acknowledge this fact.
Your great leadership doesn't seem to care, so either your management knows better than you or maybe you should push back on the notion that you worked with "great leadership."
Only poor leaders ignore their workers, which is what you seemed to have actually experience.
I would argue the average code quality of LLM's today is much higher than pre-AI code quality. It's better documented, more readable, and has fewer bugs. There was a brief period where frontier models were still worse than the average developer, but that period among frontier models is well past us.
Are you saying that pure vibe coding by a non-technical person produces better code than pre-LLM developers, or that experienced dev + AI produces better code?
Both of those things are very different, and AI shouldn't be the one taking the credit if it's the second case.
I've been doing development, in one way or another, since the 90s. I've worked with dozens of teams from enterprises to startups. Hundreds of developers. The quality of work has been all over the place, but the majority was not great.
I'm arguing that what people today call "AI slop" is already higher quality than what most developers created historically, and the fact that tests and documentation pretty much come for free now means that the floor has been raised.
The quality of AI generated code is not great. Yes, it will get better. It's already better than 65%+ of what regular devs can do AND it is faster to produce, iterate, and release.
This is off-topic, but I strongly dislike AI written documentation.
When I see AI house style my eyes glaze over. Just this morning I reviewed an RFC from a colleague that he said was a spec for a web service. The document had no introduction, no context, it described endpoints for 2 distinctly different services instead of 1, and made no effort to reconcile why there are 2. It was scattershot with details, some of them important, some completely irrelevant. It was replete with typical LLM-ism.
Basically, it was a dump of a conversation he had with an LLM. As a document to build shared knowledge, it was nearly useless. The only feedback I could provide was a polite "I do not understand what you are trying to build".
But, supposedly, another engineer is already working on implementing this spec. I assume the other engineer just cycled this "spec" into his LLM, and off the two of them went. \o/
They are trying to pull me into their project right now, I stood up some containerization infra for them. But, oh boy, do I not want to join. I looked over their codebase, by LOC the codebase is 35% comments, and a lot of the comments are contradictory, there are dependencies that are not used, there is no tooling of any kind (no type checking, no linting, no PR process), there is no auth (this code is already running in production lol -- they have public endpoints exposed that can be used to scrape/mutate internal company data). Another 30-40% of the codebase is unit tests that test trivial stuff like whether their framework's serializers and ORM work, ex: x=DB.create_x(arg=1), assert(x.arg == 1).
At the intuitive level, I do not understand people who say coding is solved... To me it seems like LLMs are a multiplier (LLMs are amazing, sci-fi level shit), but if you multiply a negative number or 0, you get something that is <=0. Making agentic coding work requires a lot of discipline & expertise.
Oh man, the "unit tests" that test whether the framework/browser/language is doing what it's supposed to drive me insane. Those have their own tests already! Test the unit under test, that's why it's called that!
AI documentation is practically worthless ime. I forbid it in my projects. It's almost always more useful to not have any documentation and read the code than to rely on AI docs.
I wanted you to be wrong, and to be able to make this an example of us over-reacting to certain trigger words created by AI, but unfortunately I just scanned the first couple paragraphs with pangram and it reported 100% AI, so you're probably correct.
I get a funny feeling in my stomach over the idea that common and effective means of communication (i.e. it's not X it's Y) have become faux pas to use because of AI. I think it's something about these phrases being taken away from us more-so than the AI inventing them.
Pangram should be paying HNers for how often we pitch needing to use their product by name to believe things as obvious as "a long form news article cramming every AI trope it can fit from start to finish" being AI written.
At this point I'm surprised when a news article isn't largely AI written, let alone one using the default tone! I don't even mind it as much as others seem to, it's just turning into more and more of a rarity for a news article to not be these last few years and so is now what sticks out.
> Pangram should be paying HNers for how often we pitch needing to use their product by name
I kind of assumed it is, given how it suddenly seemed to start getting namedropped in multiple comment threads. And it's often in response to someone saying content is obviously LLM (with examples) and the shill wedges pangram into the conversation "omg you're right, I didn't believe you but I checked this new product and wow it agreed with you" as though that adds anything to the discussion at all.
Pangram gets brought up a lot because if I read an article and think it's blatant AI slop and want to communicate that fact, a natural impulse is to provide some sort of objective corroboration rather than just asserting that I have superior taste and thus am able to tell.
Also Pangram is basically the only AI detector that's actually put in the work to build an accurate classifier, so if you try to discuss AI writing detection without being specific that you mean Pangram you'll get half a dozen commenters screaming about how awful some of the snake-oil salesmen like GPTZero are.
I mention pangram because it’s good at what it does. Otherwise, people like to claim “That’s just how I’ve always written. I always say That’s the seam and the seam undercuts the other pillar. It’s not just correct, it’s also verified lol; it’s just how I speak” and it’s obvious bullshit. Depending on the politics of the poster, people will also support this kind of garbage, but it’s all drivel.
The authors are so often dishonest that it’s hard to believe that the things they’re saying are anything but a random hallucination.
Em-dashes would be a loss, they can be a better flowing version of a parenthetical. “It’s not X it’s Y” is not a loss. This phrase is a symptom of a situation where the author wants to subvert expectations but doesn’t have the space, ability, or faith in their audience to organically set up X as the thing to be contrasted against.
I suspect it became an LLM tell because it is over-represented in text that’s easily available to the models but that most people don’t actually want to consume: marketing text, LinkedIn posts, that sort of thing.
AI detection in long-form content is a problem well-suited to training a classifier model. We have tons of verifiably-not-AI text from before 2022, and you can create tons of verifiably-AI text. As language drifts over the next few decades, it may get more difficult. But right now it is quite a manageable problem for languages with large pre-2022 text corpora available online.
The reason why AI detection tools other than Pangram are awful is because they are not really trying to solve the problem -- they just want to appear good enough to convince people to use them.
Imagine writing an article in 5 days. Working hard. Posting it online and then everyone says it is ai and laughs and dismisses it because a tool says it is AI.
Fwiw, I only did the first 300 words (signed into pangram) and it seemingly correctly noted that there were 2 (mostly) human-authored sentences in there
What's wild to me is that:
1. People are responding to this article like it's hitting a nerve
2. In most Ubers you already don't talk to the driver (yellow cabs are higher variance in NYC). Doesn't seem like anything's being lost in that case.
3. There are enormous safety benefits to waymo, mobility benefits for youth (and elderly) that are afforded by this technology. It's not clear why people argue "Uber" is better than waymo. A few years ago there were arguments against Uber! (A technology which, again, provides a huge benefit, especially if you live in an area where people were previously expected to go out for drinks and then drive home)
I think it's fair to point out real issues at these companies. But we should be clear-eyed about which technologies we want to accelerate vs slow down
Have you actually read the article? They point out multiple times that they like the Waymo ride and would use one again and loved it. It's not a Waymo-bashing article, it cautions against the long term effects on research.
And here you have it — I used an LLM-ism. Oh, and now another one! Time to downvote me for supposed LLM use! (Which I obviously didn't, I'm typing this on an ancient smartphone waiting for a train, but that doesn't deter the witch hunters.)
> You actually didn't. The construct is slightly different.
Well I know that I did, so you are wrong with that claim, which sheds a strong (negative) light on other statements of fact that you posted in this discussion.
You may disagree with me, or I may even have gotten something wrong, but just claiming I didn't read the article ... sorry man, I thought HN had higher standards than that.
They were pointing out that "It's not a Waymo-bashing article, it cautions against the long term effects on research." does not sound like LLM-generated text.
It would sound kind of LLM-y if you had said "It's not an anti-Waymo article. It's an article about how assistive technology is quietly degrading research."
(Of course, the correct approach to detecting AI is not counting LLM-isms, but feeding long-form text through a classifier model and picking up statistical correlations that are more in-distribution with LLM text than human text.)
"It's not a Waymo-bashing article, it cautions against the long term effects on research." is not an LLMism. It superficially resembles one, but that's it.
Yes AI slop is prevalent. But the casualties that get thrown under the bus are also real, don't you agree? Is it really worth it to sacrifice those? In the name of the larger good?
I'm arguing that we're losing something by doing so, as a community.
> I find complaining about AI callouts
If they are correct then I'm all for it. But just suspicions turned into statements of fact are not helping anybody. People used em dashes before LLMs Not nearly as much as LLMs, but superficially judging people isn't doing any good.
I'd compare this with pitchforks coming out agains criminal immigrants. Yes, some immigrants are criminal. Doesn't mean that if you stand in central Copenhagen and have a dark skinned person in front of you that it's a criminal.
The rule of thumb is simple: if most of a user or author's writing is consistently human-generated, then I think we are very happy giving them the benefit of the doubt if one article or snippet flags the actually good ai detector
Unfortunately, many of the largest voices against pangram simply don't like it because it gives their lies less credibility.
I don't mind reading llm writing, indeed, I read more llm writing per day than most. But if you're using LLMs to increase the level of slop (blog posts, comments, tweets, etc) then we should call people out.
It's a colossal waste of everyone's time and attention, and we should be mad about it.
Late edit: if people are actually writing and it's consistently flagged by pangram (this is a statement I have not yet seen validated), the pangram folks are extremely proactive and excited to understand what's going on. This notion that somehow the authors of slop are victims is complete nonsense.
> if most of a user or author's writing is consistently human-generated, then I think we are very happy giving them the benefit of the doubt if one article or snippet flags the actually good ai detector
I would agree to that approach.
But that didn't happen here. Nobody has done that due diligence and folks are just blindly accusing. The author has published for many years. That's what I'm calling out.
> This notion that somehow the authors of slop are victims is complete nonsense.
Strawman? I didn't claim that or anything close to it. I also find the prevalence of slop writing hugely annoying.
I once heard someone suggest that this should be on the OS level and I'm slowly coming around to the idea. There should be some kind of OS level flag that can easily broadcast to products that a child is using the device. No ID verification required, the parent is responsible for setting it, and all downstream applications are regulated to respect it, If applicable to the product. This removes the incredibly invasive ID verification, and places responsibility on the people responsible for the minor. Companies like meta then can be sued not for failing to detect children (this is, evidently, not working on any platform trying to enforce this currently), but instead can be judged in a black and white manner (i.e. is the child version of meta too predatory to children?)
This is almost right, but it shouldn’t be “is a child using the device?” It should be “is this a child-locked device?”
There should be a children’s Internet, just like there are children’s libraries, and child-locked devices should give access to it. Adults should be able to use the children’s Internet to see what’s there and children should be able to use the adult Internet when supervised by parents and teachers.
Technically, the only thing cooperating websites need is an http header indicating that the client is a child-locked device. Websites can disallow creating accounts or logging in from child-locked devices when they’re only appropriate for adults. There can be laws prohibiting advertising on the children’s Internet, etc, and legit websites will have to follow them. At no point does a website need to know a child’s age or anything else about them. Vendors selling devices are responsible for not selling unrestricted devices to children, but that’s easier than making every website do it.
Since the Internet is still a dangerous place, for non-cooperating websites, child-locked devices do still need the usual whitelists and/or blacklists.
There should not be an HTTP header indicating that the client is a child-locked device. That puts the onus on the server to respect the header, and HTTP doesn't require any action on unrecognized headers. Also, it reveals to the server that the client's user is likely vulnerable to manipulation — exactly the opposite of what you want!
Instead, there should be an HTTP header indicating that the server is an adult-only website. Then, child-locked devices can refuse to show the content to their users. Moreover, this can be more granular than just a single adult-only bit.
If the current age verification controversy was intended to protect children rather than destroy anonymous speech, it would be focused on requiring the implementation of PICS or something similar.
That standard is mostly used to indicate “adult” websites (porn, basically) which is useful, but rather a different thing than what I’m advocating. There are many websites intended for adults that should not be blocked when Google’s safe search is turned on.
It should be possible to write a web crawler for the children’s Internet. Whether it’s a client-side or server-side header doesn’t matter for that use case.
Roughly distinguishing between children and adults based on behavior can’t be stopped and shouldn’t be considered a privacy violation. For example, see how Lego does it: https://www.lego.com/en-us
A client-side header would let them do a redirect instead of a dialog box. There could be alternatively be a server-side header that causes a redirect, but either way, they could distinguish child-locked devices from unlocked based on the destination of the redirect.
100% agree, and no OS should be forced to implement this "child-locked" signal. the existence of OSs that do implement it should be enough (if you want to lock your child's device, use a lockable OS).
There are many kinds of proxies. Proxies would have to be blocked if they don’t cooperate. Obviously a child-locked device shouldn’t let a client-side proxy be installed.
There could be lock levels to this. By default everything is unlocked (level = 0), but if a content provider / host receives a signal/header with lock-level > 0 they should be required to honor it. Something like this would require government which means it will probably never happen. Much more power asymmetry to just track you.
No, that leaves vulnerable adults unprotected. It should be "Is my thing in one of the categories of things that this device says this user is not permitted to do? If so, I shall not permit this user to do the thing.".
Nothing stops software authors from providing pre-built bundles of categories that they believe fit certain types of vulnerable people [0], but the fine-grained control must be there so that guardians can choose to set things up for those they guard so to adequately protect them while minimizing the amount of stuff that they're blocked from.
[0] Like: "Overly-trusting human who needs protection from scams", "Dementia-damaged adult who cannot be trusted to manage their finances", "Median sixteen year old USian", etc, etc.
Sure, the categories can be expanded. For example, movie ratings aren't just "children" and "adults." But explaining the simple case seemed like enough for one comment.
If you want to make a child-safe website that is OK, it can go be in it's walled garden g-rated brand-safe reality. But that is retarded. The rest of the internet still exists, and will not stop existing. Porn/defense distributed/much more insidious things will still exist.
All the age verification is is creeping totalitarianism by governments.
Building this functionality into an operating system is an extremely slippery slope to even further centralized corporate and government authoritarian control of people doing what they want with the personal computing devices that they own and possess.
I think you may be misunderstanding a bit. There's no forced verification. It'd be more of an RFC that gives parents the ability to communicate their underage child is using the device without revealing or verifying any further information. Or do you suspect that giving any ground will cause the ID verification and centralized behavior?
I think the concern is a forced bit of functionality in all operating systems. It would quickly lead to a "and you need a license to sell your device so we can check that it can't bypass mandatory chuld-safety guards". Just like how some politicians are proposing mandatory licensing for releasing any AI model.
You need licenses for a lot of things though, particularly for things where there is substantial harm to be done when done incorrectly. This is just how it goes as things get more popular. Think back to 1903, there was no FAA then, but also no need for one. But it would be super disingenuous to argue now that airspace should be totally unregulated.
I think the problem with this debate in general is that people aren’t recognizing the harm, and are clinging on to old ideas about how the world should work, ideas that just don’t acknowledge the reality of how things change as technology changes, or even just spreads. Rejecting the idea that there are harms, and thus nothing should be done, just ensures you don’t have a seat at the table at all when it comes to the inevitable decision to do the regulation.
I think you aren't recognizing the harm of collecting this information. I might be okay with this if companies were banned from using age for targeted advertising.
However, the other harm is the recent IDScan hack which leaked 153M people's IDs. And regulation is not a solution here, we don't know how to implement a regulatory regime that will prevent these sorts of privacy disasters. Even if IDScan gets fines which kill the company (and I suspect they will not) it's not enough of a deterrent because no one will pay enough to actually provide proper security here.
Licensing for selling anything with a computer in it would cripple competition from small incumbents. They're mostly just regulatory capture for the established players that created the problem in the first place. Age restrictions are also an implicit statement that they are allowed to continue doing the same harmful things to adults.
Meta, Google, Anthropic, OpenAI etc can afford to pay for and deal with licensing. They also created these issues and would very much prefer not to have to mitigate the harm they do to adults (i.e. people with money to spend). Furthermore, it'd be great if they didn't have to worry about small time competitors emerging and growing too fast. Licensing under the guise of "think of the children" is perfect for them.
I suspect that communicating a flag value of boolean true/false that "this computer is in use by a minor" from the operating system (via browser or app) to a remote SaaS service like something run by Anthropic will be seen as insufficient by the SaaS, so they'll necessitate ID-scan/live-selfie verification anyways.
Meanwhile, the age flag in the operating system will be used for other forms of authoritarian control. It will have actually accomplished nothing other than limiting peoples' fundamental civil liberties.
That's a valid take. The issue I'm wrestling with is the inevitable attempts to point a finger at who is responsible when bad things happen. If you claim it is on the tech companies to know if a child is online, then they will take the safest path for them by forcing ID verification in independent adhoc manners. This is what we are seeing now, and to me this is the worst situation. If you shift the responsibility more to the parents (this child was using a device that didn't send the `PARENT_CONTROL` flag, therefore we assumed they were an adult) it's a completely different conversation. Furthermore, if something happens to a child AND it's obvious from traffic logs that the platform willingly knew a child was being talked to, then that is also a completely different situation than the first.
Leaving it all completely deregulated and/or letting platforms implement it themselves to varying levels of success and personal invasion feels like the worst option to me.
I think you need to explain the solution you're thinking of in detail.
"an RFC that gives parents the ability to communicate their underage child is using the device without revealing or verifying any further information." is no different to "turn on age lock?" that you see on website now. It requires the parent to be present, engaged and 1 step ahead of their kids, and if we could depend on that, then we wouldn't even be having this discussion.
He was saying, vendors put that flag in the OS and parents set that flag in the OS and internet services respect that flag from the OS. No website anything, all in the OS. No staying ahead of anything, one and done. This, as noted above, requires that services respect the flag, which they do not universally do today. That requires legislation, something that will be a part of any workable solution.
When I was a kid, I spent a lot of time rebuilding custom kernels and installing Puppy, DSL and other niche small operating systems on whatever random bits of hardware I could come across. I assume a lot of other HN users did the same. My parents would have had no opportunity to set the flag, meaning everywhere would have to assume default locked/blocked.
> This, as noted above, requires that services respect the flag, which they do not universally do today. That requires legislation, something that will be a part of any workable solution.
This is advocating for a variety of closed internet that is legally locked by default and would require a level of legal and international cooperation that would never happen. I'd argue that the resulting network isn't the Web any more.
An alternative one could consider is that the ISP could DNS-block sites for "kid SIMs"). On the OS part, this would require protection against picking alternative DNS, but that's already standard. Other "anti-circumvention" features might be required, such as factory reset.
For all the protections that the ISP cannot reasonably provide, one can consider creating a nation-wide (or EU-wide, or US-wide) official label, that device makers can obtain if they meet certain requirements. Parents can them make their choice. We have that for food etc., so why not for internet access mobile devices. Because from my perspective, the main problem that needs to be solved is that parents cannot monitor 24/7 what their children do on their phones.
Yet another solution is to create a top level domain for kids, restrict kid devices to that domain, and allow sites enter this domain only if they meet defined criteria.
Oh, and prevent ISPs and phone makers to sell 'for kids" products at a higher price. People can whine all they want about free market, but it's to protect our children.
Whenever I see people say "oh we should just get the ISPs to block it!" I wonder if they work at any ISP, or are just an end-user consumer of their services... As someone who does work in the ISP/telecom infrastructure sector my first reaction is hell no. I do not want US/Canadian ISPs to go down the path of what has happened with, for instance, Spanish ISPs DNS-blocking football related things as a result of court orders.
I am not, indeed. However as programmer I was sometimes told to do what the customer wants, even though it is complex or inapt. If this DNS blocking is inapt or misused, it's a political problem so let the citizens deal with it (I know this seems like wishful thinking given the current poor state of politics at least in the EU, but hopefully we will fix that while it's still in our hands).
Besides, I read about what happened with Spain and football, and from my understanding the issue was that they blocked not DNS but IP addresses. And problem domain is also vastly different, a priori: in that case, they faced sites that were actively trying to avoid blocking, while in our case, I doubt pron sites and such will try to do the same, because children are not their target audience.
An alternative could be to create a top level domain (TLD) "for kids" and restrict kid phones to it. Then in your TLD you can admit sites that meet minimum criteria (no pron, no terrorism, content and chat moderation for SNS).
Apple verified me by the age of my account. Claude asked to use the age verification when I opened it on my iPad. Haven’t been asked for any other kind of age verification by Anthropic.
I firmly believe that tech companies being responsible for parenting our children is a deeply misguided response to a failing society and overworked parenting generations.
I benefitted greatly from early access to computers and the full scale of their capabilities. It was a constant issue in my family limiting screen time, and who knows what the exact rules should be for everyone. For me it was a dynamic ruleset I was actively involved in the development of.
I could make compelling arguments to my Mom, that will not be the case with Meta, or Apple, or Anthropic.
Accepting "OS level" age validation is surrendering to a system meant to protect the industry, not our kids future. And while there are serious conversations we should be having about how to control technology use, it should be about parental controls and sane defaults, not age validations and legally restricted access by providers.
I hate how even the most progressive-minded people on most fronts are falling into this trap because big tech has failed them and they are angry at it all.
Having had a similar upbringing, I strongly agree. I have a hard time not lashing out at people that fall into this trap given how important that early access to computers ended up being to achieving my dreams.
Personally, I'm not in love with the idea that something harmful to children is totally okay for adults. If social media is that corrosive to young people, maybe we should take a look at that and make it less corrosive for everyone.
Some kind of "child lock" on the device makes more sense to me; a parent or whatever can child lock the device before handing it off, no need to collect images of the child in question and store them in some big pile at a private company.
This kind of thinking just leads to an obvious analogy: alcohol. Alcohol is harmful to children and adults. Yet when the United States tried Prohibition it was rolled back. It turns out that people want to do things harmful to themselves.
> There should be some kind of OS level flag that can easily broadcast to products that a child is using the device.
Anything like that should be configured by the parents. Your child is your responsibility. I shouldn't have to take a picture of my passport just because someone is a neglectful parent.
Doesn't this break if the child uses a VPN, or some other software to spoof? For this to really work Windows would have to lock down all HTTP, and really all networking, only allow approved software/browsers, etc. Is there a reason this wouldn't be the case?
The problem is that it would be a client side identity verification mechanism and won't work for websites.
As an example this would means that if this exists in iOS, everytime a safari loads a webpage it has to send this additional verification identifier as a header or a query param to the web server that serves the request and based on that the web server deny the request if the identifier is present (that means a minor is trying to access the web page).
But this client side identifier can easily be bypassed via a proxy that strips the identifier from the request because not matter it's configured in the OS of the device, ultimately request would just goes out as HTTP and that can be manipulated. That's why it has to be a server side verification of the age/identification.
I think the appeal here for parents is it’s a single point to configure. It’s not perfect, but perhaps better to configure one device than every account individually. I’m game for it.
Please correct me if I'm wrong, but this appears to require Devin to use? I'm disappointed to see I need to use a bespoke platform to interact with this agent, to the point that I probably won't be trying it.
But I don't want to use your CLI. I already have my own harnesses and workflows. The friction is too high to "just try out" a new model like this. It would be preferable if I can evaluate it over, say, open router like all the other models and then decide from there if it's worth downloading a bespoke tool chain for only 1 lab's models
It’s preferable to keep inference capacity available for users using main Devin products than openrouter atm. Might change in the future. Even OAI is cutting off new plan signups to keep up with demand.
> SWE-2 is free to use for users like yourself for the next month, and almost all usage should be supported via our CLI (https://docs.devin.ai/cli)
Thanks for pointing that out; I wouldn't have found out about it otherwise. I bought the 20$ subscription and have been using SWE-2 for a few days now, and it's actually good. I don't think it's Fable-class, but SWE-2 Max feels substantially better than K3 and on par with Opus 5 Max.
I just gave it a try and it doesn't appear to be free, it used up some of my on demand usage. It does say 75% off though. Seems like for Pro subscribers SWE-1.7 is free, maybe SWE-2 is free for them?
If its weights are open, that covers a multitude of other sins. Sufficiently-strong performance on the part of the new model would justify adapting existing tools to work with it.
LLM's are a funny technology because on the one hand this is all undeniably impressive at the rate of what's changed from them, and yet despite that I find myself disappointed by the lack of breakthroughs for things I don't find interesting. I like math and programming, and LLM's are pretty good at it, when are they going to get good at folding laundry for me? I think a lot of robotics work promises to solve this category of "boring" breakthroughs, and I'm optimistic we'll be able to achieve it, i just wonder when
The bottleneck is not really the intelligence here.
We can build robots that do the things you want. Arrange a visit to Amazon's robot warehouse tour.
We can't ship them because they break all the time with current technology. It would be a tough sell to have to being in a 100kg robot for servicing every few weeks.
This was cars in the first several decades of automobiles. The tide shifted as soon as you could just drive the car to a neighborhood dealership for servicing. It's fun to imagine the logistics of that for robots but the material science and engineering has to advance a bit.
Just have two robots, and teach them to service each other. Problem solved!
Only sort of kidding, tbh having bots service themselves (and being intentionally made in a way that they can service each other) just makes a lot of sense.
An automated service station could be quite compact; it wouldn't need plumbing, lighting, human-comfortable climate control. Parts can be modular, and when a station gets low on spare parts, a self-driving truck could come by to pick up damaged parts and drop off replacements.
We're not there yet, but I think we're a lot closer than most people realize.
Really? All you need to do to stop this is to stop governments from hiring robot "labor". Surely a democracy can make that decision, even when it means paying a little bit more, no?
I think the parent's point - and what I more or less agree with - is that the problem is the hardware.
Human arms and hands are incredibly intricate. Reproducing their facility with hardware requires a large number of actuators and finicky fine parts. This isn't a software problem. Industry solves it with maintenance schedules.
There's probably nothing in your house that has as many moving parts as a robot needs. Your car maybe, and pretty much all it does is rotate wheels.
Robot vacuums are massively simpler and they fucking break all the time. Wheel motors or their position sensors, belts, plastic gears, contacts that corrode, a circuit board someone decided to not comformally coat and a cat puked on it...
I've learned that I spend significantly less time folding laundry than many of the people commenting about modern ai powered robotics. It's not meant literally is it? For instance keeping floors and counter-tops clean seems a much bigger time sink for me.
There's definitely a wide variance in laundry. Laundry is something that I can do to kill time while boiling pasta on the stove, for my partner it is an entire afternoon task.
It's just a general "thing I don't want to do" not the thing that takes the most time. Could be taking out the garbage. It's just an example menial task.
Pretty soon. Sunday Robotics had a 3 hour stream of folding clothes with 99% accuracy. You can watch it for yourself. There’s a lot of “hand” companies with very compelling videos just over the last two months. Then there was Figure’s multi day livestream of package manipulation that was very impressive. Physical LLMs are definitely coming. Given enough training data we know LLMs can output coherent data in any space, it’s just a matter of time.
I feel the iPhone or ChatGPT moment for robotics is coming soon. Lots of different companies doing interesting things. What’s missing is somebody putting it together into a compelling package.
I'm starting to think that a ChatGPT moment for robotics would require abandoning the current brute force approach to artificial intelligence. ChatGPT's trick was simply more data and compute. A stochastic parrot will eventually become very impressive. But we don't have an equivalent of billions of lines of code for robotics.
The economic value in automating high skill and expensive labour that the most profitable corporations rely on is what attracts the most investment and effort.
Once you lose your job you'll find plenty of time to do all those menial tasks you hate wasting time on.
Jokes aside, programming and mathematics are just well ordered environments. Already digital and verifiable. Perfect for a GPU cluster to power through.
The real world is big slow and complex in comparison.
f folding clothes, i wanna just have a robot be my personal chef. the amount of different tasks and capabilities a robot will need to make any meal that I can in my kitchen is huge and i feel like it’s still a while from being solved.
It may sound strange, but cooking is one of the most intense and attention-consuming tasks I encounter.
I have to do it every day too.
So I fully agree with this line of thought... Many a time I have considered that I would happily spend more on a personal 24/7 chef than I ever would on a car. Cars to me are utilities and should simply be efficient and optimized to purpose - food is luxury and taste, it is sublime experience and art.
Maybe that's why I can't make it, treating every recipe like a strict command chain isn't how art is done. Can my taste buds be scanned?
I mean honestly anything is taste and art. People care about very different things. Some people see their clothes as the biggest expression of themselves, whereas others don't care about their clothing but see their cars as art and a special experience thats separate from the purpose.
But yeah, I think a lot of food is just practicing. And if possible, learn from someone who's food you do really like, and experiment. Just try making some dish you really liked, you'll be surprised at how far you can get.
Unless we spend a bunch of money generating data, I can't imagine the machine steps to fold laundry are very big in the general training sets. Someone is going to have to find a hardware system, cheap enough to make it economically feasible then train it. As far as tasks people will pay a lot of money for a robot, this seems low on the list.
future is on the way, three to four generation ( one each year ?) will unfold this, primarily quantum computer improving material and battery, humanoids becomes standardised and modular enough to be easily replaceable ( think ibm pc ) ( most components are simple injection moulded advance plastics , mass produced in some corner of china, self detection of wear and tear and self replace that part ), other is optical computers ( 100x lower power x 100x speed = local inference ), problem is, when this will become reality, who will benefits more ? who will hold moat ?
It's because existing models are a brute-force approach to intelligence. With enough data and compute, a stochastic parrot will become very impressive.
But with robotics, there's no pre-made dataset that can be parroted. Notice that these datasets, e.g. how to fold clothes, need to be created by humans. That's as if humans needed to write algorithms like quicksort to teach LLMs how to code.
Has anyone been able to get anything substantial done with Fable in the first place? I more or less had totally given up on using it since the alignment checks were so sensitive that it pretty much always threw me back to Opus.
I hear this a lot and I believe it because I've heard it from so many people, but I have never run into this in my work, and neither has anyone I know in real life.
I don't use Fable for a ton of implementation work, but I use it a lot for planning, so maybe that's related to it. For planning though, I've had a very good experience with Fable and implementing with Opus.
I don't mean to sound like I'm dismissing your experience, but are you sure? I've (semi regularly, most of the time I'm even trying to use Fable) started with Fable, proceeded through my planning, and then at some point in the future realized it had kicked me back to Opus without me knowing. It obviously _said_ it had happened, but I didn't realize and just continued. This might primarily be a result of the project I'm working on (anything network related seems to gets kicked back).
I'd guesstimate that ~80% of the time I thought I was using Fable, I wasn't actually. It's also led me to just... not even try, and just start with Opus regardless.
I've found Fable unusable; not because it's bad, but because it... can't be used.
No that's totally fair - I want to say that I haven't, but I guess I really can't be sure. It's very possible. I'll keep an eye out for the next time I use Fable.
FWIW, most of my code only encounters security concepts as standard implementation of best practices. I'm not in a security centric position.
Do your apps do anything with security? I can't hardly use Fable on our authentication service because it constantly trips up and refuses to write tests. Even just doing a security review usually triggers opus.
I do very security cyber dangerous work like building a signup/login form or setting up a certificate. For obvious and good reasons Fable refuses to work on such sensitive stuff.
I asked Fable to build me a set of telemetry scripts for a specific use case so I could get a working posture of an environment. It assumed I was building some reconnaissance tooling for a nefarious reason by default and noped right out. Apparently the powers that be couldn't ever imagine their tooling being used for understanding device state. I can't wait until I need to get an exception from the USG so that I can use sudo to check some processes on a box that's spiking CPU.
I agree and wonder whether its either people who basically never use the model complaining or people who used it once a long time ago and haven't touched it since.
We have access to Fable at our company on our enterprise plans and most of us rarely run into an issue.
Obviously this is gonna vary a lot with what technical domain you work in which is why its important when talking about the classifiers that people specify exactly what types of workloads they were seeing failures with.
You can easily trip it up if you're doing reverse engineering work. From memory, the moment Fable 5 saw anything loosely related to "linux seccomp" it threw a fit.
No, almost anything related to my job is flagged for "cyber" and my company currently has no plan to try and enroll into mythos. I'm not sure if anyone has been able to enroll solo.
It did help with some worldbuilding for my book (it wasn't incredible which gives me some hope for writers). So far opus 4.8 is the most reasonable model.
I heard the only way you are going to get into the CVP program is if you have public CVEs. Doesn't seem to matter if you are in a company account or not according to people that are supposedly in the program.
You don't seem to be alone: FT.com: Anthropic’s best AI model struggles to attract users as cheaper tools thrive. AI lab’s Fable 5 has met with sluggish demand from corporate clients [1]
Combo of that, laziness and load-bearing language + the penchant for making up weird dense conceptual names pushed me to sol 5.6. They seem to indicate it is a less annoying writer in the announcement so I’m curious to try it out, though.
Ironically one of their demos is speeding up inference - do us normies get to do that with Anthropic tech??
Useless for reverse-engineering the software that talks to a ten-year-old video cam + DVR system I was given, really nice for things I actually do in my day job (web dev at an agency).
For a long time, no - it was completely unusually for my work that references biological information about migratory birds/other (innocuous) seasonal phenomena.
About a month or two ago, they must have tightened the black list on bio topics as it became more willing to process requests without visibly downgrading to Opus.
I've used Fable for so much stuff. My experience has been that it can pretty much one-shot most of my complex problems, if I describe them clearly and provide a solid way for it to verify its work.
I get punted down to Opus 5 occasionally (for security-adjacent things) but that's pretty rare.
It's probably a good model for folks doing basic software stuff, or humanities related tasks, but I work in cybersecurity on the defense/detections side and I haven't been able to use it for anything even with being in the CVP. It downgrades to Opus every time.
I have been running Fable with Binary Ninja MCP. It will reverse engineer a binary in a lot of detail if you give it mild direction and I haven't had it flag. I think it assumes since I have a valid binja license I must be responsible lol.
I do think probably ralph looping a binary locally first is going to be best to get 100% recovery of types and function behaviors then letting a smarter model churn the final steps.
The only problems Opus struggles with, Fable won't take on. I was porting some software from Win32 to linux. Opus was running in circles. Fable was going great until it saw some authentication code and bailed.
Yes, I used it to do a big push of my self-maintaining project
All I need next is to push it through enough cycles to trust it (I have high confidence based on results so far) and then I can just setup a cron to do routine maintenance on my project
In my experience, the guards are less strict than they were at first. When Fable came out, it dropped back to Opus 4.8 for about 50% of my prompts. Now it's maybe 20%.
When fable first released it was almost useless. Since then, it's improved a lot. It has been working on my binary ninja MCP server just fine. It flagged for cyber 1 time (no idea why), but it generally works fine.
I have noticed sometimes it likes to gaslight itself into thinking that everything its doing is allowed or allowable, I saw that it thought the game I was reverse engineering was running on a private server (it was not) so it assumed it had permission to do anything lol.
Nope. Always failed within 2-3 prompts. The most basic REST service you can imagine. Cookies are signed, that's crypto, banned. Completely useless model.
There's a couple different ways to look at this. From a charitable point of view to ubisoft, they do not advertise support for Linux. Instead they openly state it's a window's only game. So when a third party (i.e. Valve) makes shim that grants linux functionality, it's not really on Ubisoft to maintain that.
However, from a consumer perspective, this is incredibly frustrating. The average person does not think of "linux vs. windows". They just bought the steam deck and when they bought For Honor it was sold on that "platform". Again, it's hard to blame Ubisoft for this conundrum since they were never necessarily trying to be steam deck compliant to begin with, but it definitely feels wrong to take the game away from a player in this perspective.
All that being said, historically across many games such as Rust, many are quick to blame linux for cheating, but pragmatically banning linux has yet to measurably reduce the amount of cheating that players report experiencing, so I don't agree with the decision at all. It's a footgun disguised as a band aid
Oh man, if you read the threads about that it's such a time capsule of a different era:
e.g. from this thread: https://news.ycombinator.com/item?id=31628342
> I am honestly surprised how little SPAM there is on GitHub in general. Please don’t take that as a challenge!
reply