Hacker Newsnew | past | comments | ask | show | jobs | submit | anileated's commentslogin

Why was this flagged



The issues with LLMs go beyond just IP theft. I would not say PRC making LLMs cheaper is the best outcome (though it is better than nothing). The best outcome would be to make the practice of training on our data without consent illegal, which would simultaneously slow down economic change and make it more organic as well as give PRC companies less capabilities to extract.


> The issues with LLMs go beyond just IP theft.

There is no IP theft because LLM outputs aren't protected, just egregious ToS violations.


> There is no IP theft because LLM outputs aren't protected, just egregious ToS violations

I meant original IP theft that occurs to train LLMs in the first place. But sure that implies that further LLMs based on that LLM are also tainted by that original IP theft.


- Deriving a “no derivatives” licensed item is illegal, no?

- Selling a “no commercial” licensed item is illegal, no?

- Deriving and/or reproducing MIT licensed code without credit is illegal, no?

- Reproducing and/or deriving GPL code and not notifying and/or not making GPL is illegal, no?


I can't make heads or tails of your opinion-free comment, made up of only questions.

My best guess is you're suggesting that Anthropic's model outputs are transitively under copyright (as a reproductions of human work under copyright?), but somehow ownership now belongs to Anthropic and not the original owners, and therefore Anthropic has standing against Alibaba? Not only does this go against what Anthropic argued in court against authors and publishers, such jurisprudence would lead to the immediate shutdown all leading LLMs in the US which were all trained on stolen work.


> immediate shutdown all leading LLMs in the US

They can license training data. They have trillions, look what they are dumping into it, you seriously think they can't afford to license data.

Obviously it would be easier if they do it from the start, but that was their trick, to do it while people don't notice and get big ASAP. Should they get away with it?

Also, it would solve their Chinese problem, because it would make them violate copyright too. Right now it's more like rules for thee not for me so it's hard to take seriously.


My theory is that YouTube blocks some accounts for publishing LLM-generated music, and people who wanted to earn ad money from it get burned and publish LLM-generated posts about it.

I would be on YouTube's side here, except it's possible that their motivation is simply to avoid poisoning their dataset while they train their models off creators videos. Also, the question is how they tell apart what's LLM-generated without false positives.

Maybe there were also artificial listens fraud (it's a problem with their competitor Spotify), but we'll never know because no one who was blocked would publish that honestly.


No one is required to use EUDI: https://ec.europa.eu/digital-building-blocks/sites/spaces/EU...

Companies and providers (like banks) have to support it, but use is voluntary.

Check out the spec and legal framework, it actually makes sense and is open to different implementations, though you might need to certify it.


You are not required to accept anything other than digital ids. So from experience, whatever demands euid has will be what is required to identify you.


If they have to support something that most everybody has they will soon stop supporting alternatives that are not required by law. What then?


There are attempts make it almost mandatory through mandatory age verification. Which would mean that you'd have to submit to privacy violations or be cut off from a sizeable portion of the internet.


My prediction is that eventually services for people NOT using the digital ID will be so degraded to be almost useless or seriously disadvantageous.

Kinda like the discrimination DB does for people using paper tickets vs those using the DB Navigator app.


CEO of Roblox was once asked whether he would ever put prediction markets inside Roblox, he gave a straight face answer: https://youtu.be/XpIXRgMlPo4?t=2122


In case you don’t want to watch the video: his answer is yes BUT he needs to figure out how to do it legally in the different jurisdictions that control kids gambling.


And well, and he wants it done for a "educational" purpose without Robux (which I assume is the in-game currency), slightly missing context from your TLDR there.

But as mentioned in https://news.ycombinator.com/item?id=47334696, for sure he's still out after dat monez, he's the CEO after all.


Where does he mention “without Robux”?


Literally seconds after the linked segment. Give it a listen for a minute.

Edit: To be precise, he says "no free Robux":

> no free Robux, no free prizes, just a game called the dress to impress predictor where it's not like trying to get kids money or anything like that

You could have also read the comment I linked earlier.


Brilliant conversation.

"Would you let kids gamble?" - "It sounds very fun and obvious." "To be clear, we think it's a horrible idea!"


The CEO does not say it’s a horrible idea. The interviewers say that. CEO says it’s “brilliant idea”.


And then of course he continues (although won't make for a good social media rage point so understandable you didn't provide full context):

> Well, I actually think it's a brilliant idea if it can be done in an educational way that's legal.

> no free Robux, no free prizes, just a game called the dress to impress predictor where it's not like trying to get kids money or anything like that.

Still, probably what he sees in his mind is "more money yay" as always, he is a CEO of a for-profit company, that's what they do. But still felt disingenuous not to include the full context, doesn't even make it "that much better", he still seems like a scumbag with it.


We can mince his words. In the end anyone who organize gambling for profit is scum. If you do it for charity I know form experience it is massively profitable, but.. I am not sure it is worth it for society.


Hey now, silent auctions and raffles are great for small communities and aren't prone to degeneracy. I know a lot of fire departments that get a majority of their funding from a mix of these attractors and things like cookouts and public events.


I believe it just enables and validate the bigger actors. I do not know where the line should be drawn, if gambling is ilegal you build an illicit trade, if it is legal that trade just become more evil.

We should be careful with gambling especialy when CEOs are talking about it and only caring about the legal frameworks.


> I am not sure it is worth it for society.

As always, seems to depend on the scale.

Letting anyone in the world place bets on when the next nuke will hit a city? Probably pretty bad overall for society.

Doing raffles for some local tiny organization run by Ada and William so they can continue hosting a ten person event? Probably pretty good overall.


The problem once again comes when you decide to hyper optimize for profit. Ada and William will rely on word of mouth, maybe a few posters to drum up attention to their raffle.

Meanwhile large gambling orgs will run ad spots non stop with celebrities enticing you to join their app with free bonus bets and once you're in they will send you daily notifications to nudge you to place "just one more bet".

Easy to see how one would be relatively harmless while the other could cause widespread addiction.


Yep.

Can't even go to a local baseball game without the shit being shoved down your throat let alone try to watch one on TV.


He's also previously said he wants Roblox to be a dating service lol


Jesus Christ. This man is just a sociopath who doesn't bother to mask anymore.


Eventually the money just inures you.


Let's use correct attribution: AI agents don't hack; people hack.


We anthropomorphize everything. It indeed would be nice to not attribute intent to AI. Could save us from some confusion. Perhaps one day.


I don't think global time would be a problem like many people suggest. If you're in US and talk to somebody in Australia, you will quickly develop an intuition that time @X is night (or whatever it happens to be) over there, just like our other intuitions about how many things (weather, season, how long are sunsets, etc.) are different in different places.

Timezones are failing at all of their jobs. Getting time to correspond to sun position? It can be 7pm here and 7pm there but here it will be fully dark and there it will be still mid-evening. Knowing working hours of shops and government? Everything is all over the place. Everything is fluid and changes with seasons.

Plus, there is this unfair specialness that some countries are at UTC and others have offsets. With global time, everybody gets @0, just for different places it will be at a different sun position. (As long as we find a political way to pick something neutral, instead of saying "that's when the sun is highest in London".)

Finally, we don't have per-latitude calendar and things are working fine for us. It's February here and February in Argentina, and yet life doesn't stop even though it corresponds to winter here but to summer there.


It's worth noting that technically London uses GMT for 5 months and BST for 7 months.

The GMT offset is zero, but it's important to note the difference especially when configuring servers to avoid nasty daylight savings surprises kicking in at at end of March.

There has been talk of moving to a +1 offset all year round for lighter evenings in winter, albeit at the cost of some very dark morning, but given we couldn't even manage Metrication without people still complaining 20 years later, I can't see it ever happening.


I think you mean complaining about metrication 50 years later :-)

The counterpoint is that without the metric system how could we make snarky comments on US-based woodworking videos?


I was specifically thinking of the "Metric Martyrs" who were jailed over refusing to display weights and measures in metric.

The law requiring metric didn't actually come into force until 2000, these cases were early 2000s. Note that the law to this day still allows for imperial measurements to also be displayed, but they wanted to display in solely imperial.

https://en.wikipedia.org/wiki/Metric_Martyrs

The situation another 20 years later is rosier, even the boomers have spent most their adulthood with metric, and they're dying off now.


> There has been talk of moving to a +1 offset all year round for lighter evenings in winter, albeit at the cost of some very dark morning

Why not just offset the office and opening etc hours by +1?


Because society and culture doesn't work like that.

You can't will a culture of closing up at 4pm during GMT and 5pm during BST. That's just even more confusing.


The talk was +1 offset in clocks all year around, in effect dropping DST and changing the timezone.

Also a lot of places and services have different hours during different seasons.


If you're going through the hassle of dropping DST, why not settle on BST as the permanent timezone if that's what the preference is for hours of daylight?

Asking an entire culture to change from 09:00-17:30 to 08:00-16:30 seems awkward and doomed to failure in comparison to simply landing on BST instead.


the obvious solution is to move it by .5 the whole year round.


Australian here!

I constantly forget which way the half hour difference is between Adelaide and Melbourne / Sydney!

Then I have regular contact with offices in London and LA. For some of the year it’s not too bad, and then our clocks switch the opposite way and it gets less convenient! Which way is which I can never remember.

Queensland doesn’t bother changing their clocks at all.

Writing software that deals with Timezones isn’t too bad these days, but supporting it is as it constantly confuses users I find!


You have some fun ones. On the other side of the spectrum is PRC, where at the same hour of day it can be complete darkness on one side and almost technically noon on the other. It's super arbitrary with little rhyme or reason.


I used to think this, but mirroring the sun position makes a lot of sense. If I wanted to meet w/ someone in Australia, I would still need to know extra information (what their equivalent 9-5 working hours are).


You would need to know that person's working hours, so I don't see how you are avoiding something.

Sure, if you talk to someone there for the first time, you would need to learn what time is generally day/night. However, you will know that 2-3 times in. Just like you would automatically know that now it's summer in Oz, or 3 hour short days near Arctic circle, if you talk to anyone from there even very occasionally.

Case in point, we have global calendar with no problems.


My point is either way you need to memorize some info in the first couple of interactions and it really doesn't make sense to go through all of this change to just memorize a different thing.

If you really need to coordinate something across many timezones, you currently have the option to use UTC to specify the time.

Following the sun also gives a lot of context. (i.e. if my flight to China lands at 9p local time, I immediately know that it's going to be night, but if my flight lands at 1PM UTC, I really have no context as to what time of day I'll be landing)


> My point is either way you need to memorize some info in the first couple of interactions and it really doesn't make sense to go through all of this change to just memorize a different thing.

It's at least to make time management in systems much less error-prone and complex, among other things.

> if my flight to China lands at 9p local time, I immediately know that it's going to be night

What does that imply? If you mean "it's going to be dark", not really (you need to have more context to assume it's going to be dark at 9pm, there are places where in summer it's still very much light at 10pm). If you mean something like "buses are going to be running and McDonalds will be open", not really (you'll need to check the schedules anyway).


> It's at least to make time management in systems much less error-prone and complex, among other things.

I'm sure people deal w/ more complex issues, but 90% of it is covered by storing everything as UTC and doing the conversion on the frontend.

> What does that imply? If you mean "it's going to be dark", not really (you need to have more context to assume it's going to be dark at 9pm, there are places where in summer it's still very much light at 10pm). If you mean "buses are going to be running and McDonalds will be open", not really (you'll need to check the schedules anyway).

It implies a lot, including that less things are likely to be open, it's likely to be dark, and that I'll probably want to get to bed within a couple of hours to wake up for whatever I'm doing the next day at a reasonable time in the morning (i.e. how to adjust to the daylight cycle there).


Things you will have in context when traveling: "it's going to be cold", "it's likely to rain", "it's going to be government conference so there will be extra delays with transport", "it's going to be %holiday% so everything's going to be closed all week", etc.

You're so used to it you don't even question that, and if you add to that "@x is when sunset usually happens"... somehow I think the world will not come crashing down.


I don't think the world would come crashing down, but what's the point if it's just another thing to memorize?

The total cost/effort of timestamp translation in systems is not as high as people make it out to be.


> just another thing to memorize

Not knowing what time it is for my Australian colleagues at all without checking my phone every single time is worse. Remembering N timezone offsets (remember DST and half-hour offsets, too) is worse. Doing UTC translation or adding "my time/your time" every time is worse.

If you talk mental overhead, current system is like 10x of that than global time.

We agreed to meet at "8pm their time" but unless I literally put it into my TZ-enabled calendar app every time the chance I mentally translate it to my TZ wrong is unacceptably high. With global time, meeting would be @123 and that's it. I can keep it in my head or write it down on paper, no confusion and full precision every time. I don't even need to know if it's day or night if it's a remote call with the other side of the Earth, maybe me or my colleague works late, who cares, but I know what time it is there at any moment.

> is not as high as people make it out to be.

It's not just timestamp translation and all the errors that come from that, it's all the rest of it, waste standardizing timezones and moving them around, having to convert time all the time, missed meetings, etc.


The implication of "you have to have spent $1000 in tokens per engineer, or you have failed" is that you must fire any engineer who works fine by themselves or with other people and who doesn't require LLM crutch (at least if you don't want to be "failed" according to some random guy's opinion).

Getting rid of such naysayers is important for the industry.


"Just show me the prompt."

If you don't have time, just write the damn issue as you normally would. I don't quite understand why one would waste so much resources and compute to expand some lazily conceived half-sentence into 10 paragraphs, as if it scores them some points.

If you don't have time to write an issue yourself or carefully proofread whatever LLM makes up for you, whom are you trying to fool by making it look pretty? At least if it is visibly lazy anyone knows to treat it with appropriate grain of salt.

Even if you are one of those who likes to code by having to correct LLMs all the time, surely you understand if your LLM can make candy out of poo when you post an issue then it can do the exact same thing when it processes the issue and makes a PR. Likely next month it will do a better job at parsing your quick writing, and having it immediately "upscaled" would only hinder future performance.


What would make sense for me is to use an AI to turn implicit context that is only there in the moment into explicit context that is stored in the ticket.

E.g. maybe you have your application open in a browser and are currently viewing a page with a very prominent red button. You hit that /issue command with "button should be yellow not red".

That half-sentence makes sense if you also have that open browser window as context, but would be completely cryptic without.

An AI could use both the input and the browser window to generate a description like "The background color of the #submit_unsafe button widget in frontend/settings/advanced.tsx should be changed from red to yellow." or something.

Sort of like a semantic equivalent to realpath if you want.

I do see utility in that.


I think a URL and screenshot would be way more useful than a bunch of text for that use case.


But maybe the button only appears on that URL if you've first pressed something else, or if you're logged in/out, or maybe that URL has a different token each day that makes it seem like a completely different URL, or...


> You hit that /issue command with "button should be yellow not red".

Wouldn’t it be easier to just open the inspector, find the css class, grep the source code, and then edit the properties? It could be even easier in an SPA where you just have to find the component file.


If you are a web programmer sure. I write embedded code, I know all the things you are talking about in the abstract, but I'm not good at them because none of them have ever been relevant for anything I do. Give me a few hours and I can figure it out (maybe minutes, maybe days? - if you know this area your guess might be better than mine), but it isn't something worth my time.

there are a lot of people who are not programmers at all. I can teach my plumber everything you said (learning it myself is the easy part), but it will take years. In the end they just know "that button I'm pointing my finger at should be yellow not red". How to we transfer that pointed finger to a ticket is the question here.


> How to we transfer that pointed finger to a ticket is the question here.

There’s a reason the Support and IT Technician role exists. They’re there for talking to the end user. And they in turn will write a proper report to Engineering.

If you want to wear both hats at once, that is fine. If you want an agent to be your support middleman, that is also fine. But most LLM proponents are acting like it’s a miracle solution to some engineering bottleneck.


> where you just have to find the component file.

This can be a substantial effort, especially if you're not familiar with the project.


> I don't quite understand why one would waste so much resources and compute to expand some lazily conceived half-sentence into 10 paragraphs, as if it scores them some points.

Because it does. The goal here isn't to create good code, it's to create an impression of a person who writes good code. Even now, when software career is in freefall, for many people in poor countries it's still their only way out of poverty so they'll try everything possible to build a portfolio and get a job and the suffering of your little pet project isn't a part of the equation. Those people aren't trying to get Nobel prizes, they're trying to get any job that isn't farming with literal medieval-era technology.

My very radical personal opinion is that either we have small elitist circles of trust, or the internet will remain a global ghetto.


On the Web, the github.com/*/*/issues namespace is home to the worst bugtracker behavior in the world. A bugtracker should should be restricted to bug reports and (well-informed) proposals and discussion about the bug/bugfix. The bug report should contain, at minimum and at maximum:

1. Clear steps to reproduce (ideally, using the prepared testcase as input, if applicable)

2. A description of the behavior observed from the program

3. A description of the expected behavior

4. Optionally, your justification for why the program should be changed to behave the way described in #3 and not #4

Everything else belongs on a message board, mailing list, or social media.

But this is all totally foreign to, like, 80% of GitHub's userbase (including the majority of the project managers aka maintainers who are in charge of allowing/disallowing the sorts of things that people post as a way of shaping the tone and tenor of the space).


> Everything else belongs on a message board, mailing list, or social media.

There's a reason that collaborative code platform (not just GH but also GL) "issues" end up being used for much more than bugs:

- message boards suffer from the SSO friction issue. No thanks I will not sign up at some phpBB board of questionable admin quality that will get 0wned sooner than later, or have the board owner bombard me with advertising themselves.

- mailing lists are even worse usability-wise because these by design leak your email address, on top of that their management UI often enough is Mailman which means it probably still stores passwords in cleartext, and spam filters, attachment size limits and overeager virus scanners make it a living hell

- IRC suffers from context loss. Netsplit, go for a smoke and the laptop goes to sleep, whoops, you disconnected and don't see what happened in the meantime. Yes, there's bouncers, but honestly, the UX sucks hard. Also, no file transfers to a channel, no native screenshot/paste functionality.

- Discord, Slack etc. solve the pains of IRC but are walled gardens

- Social media... yikes. No, no, no. Eventually, people that follow both you and the author of some FOSS software get pissed off by your conversation spamming their feed. (Too) many are still only active on Twitter which excludes people who don't want to be on that hellsite. Bluesky, good luck finding non-commies there. Mastodon, good luck and pray that your instance operator and the instance operator of the project team didn't end up in some bxtchfight escalating in defederation. Facebook groups, not everyone wants to leak their real name.

- messenger groups (especially Telegram)... blergh. You will drown in spam.

GH/GL are the sweet spot between UX/SO friction (because pretty much everyone who would want to file an issue has an account) and features, and on top of that both platforms have deals with email providers preventing them from getting blocked. That's why these two platforms are so far superior above everything else mentioned.


> - mailing lists are even worse usability-wise because these by design leak your email address...

So does git and GitHub. Last I checked, authoring a git commit with an email address associated with your GitHub account is what makes GitHub attribute that commit to your account. I assume Gitlab works in a very similar way.

"But 'git clone' is soooo much harder than reading through mailing list archives!" Nah.


If you don't want to expose your email address but you still want commits to be associated to your account, Github lets you use a noreply email address [1].

[1] https://docs.github.com/en/account-and-profile/reference/ema...


> Github lets you use a noreply email address

Oh, I was unaware of that. I've not seen anyone use it, [0] but I've only paid any attention to the Big Corporate and Traditional Hacker populations.

Thanks much for the information.

[0] I'm certain that folks do use it, so folks shouldn't bother pointing out people that do.


It's the default on new accounts for stuff when people do things through the GitHub interface.

If you set user.email using git-config on your machine to a real email address and decide to author and publish commits with it, then GitHub will, of course, not be able to stop you (aside from maybe rejecting the commits when you tried to push them). It can't just arbitrarily rewrite the email address in the commit. That would break Git's data model.


I actually was looking into this recently (exploring how much of a PITA changing my github email would be) and found it interesting that, while in principle your GH email is public to anyone interfacing with your commits via `git`, they have gone to some length to avoid displaying it anywhere in the web interface. The docs actually mention it being shown on your 'profile' page but I don't see it anywhere there.


GH has a "Discussions" feature for message-boards attached to your project. Same sign-on as GH. Nobody turns it on.


Ghostty turned on Discussions and made it the first place for users to go report bugs, make feature requests, etc., then promotes those to issues once it's detailed enough for the maintainers to accept: https://github.com/ghostty-org/ghostty/issues/3558

Later, Hyprland followed suit: https://github.com/hyprwm/Hyprland/issues/9854

I'm sure there are other projects doing the same, but those are the two I know about off the top of my head.


> message boards suffer from the SSO friction issue

GitHub is an SSO provider and has been for a long time. This criticism is ignorant.

Aside from that, there's nothing stopping anyone from using GitHub's dedicated message boards for message board stuff, or, before those existed, shunting it all off into the "issues" of a separate "$PROJECT/community-bullshit" "repo" instead of cluttering up the actual bugtracker.

> Social media... yikes. No, no, no.

I'm talking about the appropriate-for-social-media stuff people are already posting on GitHub issues. It's like you started writing your comment and lost the context. People are today already misusing GitHub issues for this. I'm saying keep the stuff best kept to social media and email... on social media and email. Don't clutter the bugtracker with it, and for project managers: don't let other users do it either. (You will lose contributors who know how to use a bugtracker efficaciously and are accustomed to it but have a fixed time budget and don't want to have to sift through junk for the privilege of doing free and thankless QA on your software.)

> You will drown in spam.

The irony. Help. It burns.

For emphasis: Everything that isn't a bug belongs on a message board, mailing list, or social media, and not on the bugtracker. Anyone who can't abide by this simple, totally reasonable request should be booted.


> GitHub is an SSO provider and has been for a long time. This criticism is ignorant.

The problem is, a "sign up to contribute" is a friction source. It will almost always leak my email. In contrast, I'm already logged in to Github.


What I'm hearing is, "Your Honor, I know that I shouldn't have blown through that stop sign, but hear me out: I wanted to. If I didn't, then I would have had to stop and wait for traffic to pass and respect other people's time. Does that strike you as reasonable? I think we can agree that it does not."


What do you think about Discourse in particular.


I don't quite understand why one would waste resources on restating the points of the article in a four-paragraph hn comment, as if it scores them some points.


It isn't a restatement of the article, it's a criticism of the behaviour of the article's author.


The context windows before a prompt is often large and contains all sorts of information though, it wouldn't be just a prompt in isolation.


I was going by this example:

> /issue you know that paint bucket in google docs i want that for tldraw so that I can copy styles from one shape and paste it to another, if those styles exist in the other shape. i want to like slurp up the styles

What kind of context may be there?

Also, the entire repository and issue tracker is context. Over time it gets only more complete.


I don't get it either. The LLM-generated issue from the above prompt is just the same information written more verbosely.


The entire chat log up until the user asks it to generate a summary. And maybe "memories" and custom system prompts for good measure. A lot of potentially private information, in other words. "Just the prompt" only works in a very particular case where you ask it for something out of the blue.


Because the point is for the AI to take the prompt, figure out what that prompt means in the context of the code base, and then make an issue that provides additional information.

It's asking the AI to takes its best guess at what you actually have to do to solve/implement the issue using the code that already exists.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: