Hacker Newsnew | past | comments | ask | show | jobs | submit | remus's commentslogin

> Similarly now we're getting AI doing math. The proofs compile but are a mess. So just make the models better at writing clean proofs and explaining what they're doing to humans. That's the end of it.

The assumption here is that the true/false of the theorem is the important outcome. While it is certainly part of it, a big part of maths is the understanding you gain from a proof. Many of the best proofs elegantly explain some aspect of the maths which was previously unclear and expand our understanding of the world.

To use a programming related example, imagine that an LLM spits out a solution to the travelling salesman problem which works in O(n) time. On the one hand that's very convenient for whatever problem you happen to be trying to solve at the time...but there's also an answer to P=NP in there! The former means your delivery drivers app works a bit faster on their busy days, the latter fundamentally shifts how humanity thinks about certain problems.

Going back to the maths, there have been theorems that were proved (by people) where the proof is broadly seen as 'unsatisfactory' in that it doesn't really expand our understanding. I assume some of these LLM proofs are a bit like that: we now know that the thing is true, but we really want to know why it's true, and how that changes our understanding.


> > models better at writing clean proofs and explaining what they're doing to humans. That's the end of it.

> The assumption here is that the true/false of the theorem is the important outcome

You're replying to a comment about "clean proofs" and explaining to humans. You talk about true/false anyway. Re-read please.


Are you always an ass? He had a great comment, and you respond with this low effort hostile bullshit.

No, I see this argument repeated all the time in these threads. "It will just give true/false, while humans need intuitive explanations". Why will the AI only give true/false? Where did this assumption come from? The nested citation was about AI that can give humans an explanation. The reply was "but true/false is not enough". It is as if it was a reply to another comment. How is that a great comment?

personally I do not take such a hard line, I think a more limited form of copyright does have a place in allowing authors to profit from their work.

I think what doesn't work is a massive, blanket term applied across all works. The situation we have currently is a handful of extremely valuable works which get milked for decades and decades, and huge bulk of material which is left to rot into obscurity because it is untouchable under current copyright law.

I have experience of this with a project I run collating historic material on climbing and mountaineering: pre-internet a lot of the discourse happened in magazines, and this discourse is generally pretty hard to get your hands on. You need to track down physical copies of obscures mags that have been mouldering away in someone's attic for the last 40 years. If you talk to the authors of the content in these magazines, none of them are bothered about trying to make money of 500 words of copy they wrote 40 years ago, but because of the current massive copyright terms this stuff will not become public domain for 80-90 years in many cases, by which time I am sure finding physical copies of this material will be very difficult.

I would like to see a much more limited term by default (20 years from publication perhaps?) which still gives an author plenty of time to capitalise on their work. I think it should then be possible to renew your copyright claim (every 10 years say?) but at an increasingly high cost. £250 the first time, £500 second time, £1000 third time etc. This would mean the bulk of material would fall out of copyright on a much shorter timescale, while if you happen to have made something popular and you want to keep benefiting then you can up to a point.


How would they know for sure that some details were not part of some other training data they use? The authors may have discussed some tangential details on a forum for example, in which case you might argue that the model picked up on these details the authors assumed were benign but novel and worked out how to apply them to the problem.

Maybe it is different for you, but for me the web app is flakey at best. A bug which has been a pet peeve of mine for the last 6 months: if you leave the page open for half an hour with nothing playing, sometimes all the play buttons just stop working and you have to reload the whole page before you can play anything again.

For me the idea that poor quality search results are a bug muddies the idea of what a bug is. It's like saying the colour of the paint in my living room is buggy because I don't like it. It might be an ugly colour that I don't like, but it is what it is. A bug would be if the colour doesn't match what was shown on the tin.

> If the search bears little connection to what was searched for, then it is a bug, as far as the user is concerned.

A bad product perhaps, but not a bug.


That's because as computer people we're very invested in the idea of bugs, while your users are not.

In products your users are stuck with it doesn't matter much, they'll have to suffer with what they think are bugs.

In products that users can switch easily, you can get dropped for a product that you would consider crappier if the user believes it's less 'buggy'.


So if I search for A, and I get results for B instead (where B has only the slightest connection to A), it is not a bug?

What about if B has no connection to A whatsoever? Still not a bug? If it is not, then we may have discovered a software domain (the first for me I should say) where bugs are not possible. Should a be a nice market to launch products for, then.


> So if I search for A, and I get results for B instead (where B has only the slightest connection to A), it is not a bug?

It depends. There could be a bug in the system like query = query.replace(A,B), or it could be that there are no good results for A. Returning nonsense results is a bad design, but not necessarily a bug.

I guess my underlying point is that bugs are about software not conforming to some specified behaviour, and trying to specify behaviour like "i should type in a search and get great results" is too loose of a spec to be meaningful so we should avoid saying things which don't meet that criteria are bugs. To use a concrete example, perhaps the set of search results returned for a query look useless to you but are actually useful for someone else.

Compare that with something like a calculator app where part of the spec is "Must handle integer addition with for inputs i,j where -10,000 <= i <= 10,000 and -10,000 <= j <= 10,000". Then if in your calculator you do 1+1 and get 3, then that is clearly a bug per the spec.

Bugs vs. bad design is a useful distinction to make imo.


Indeed. Pre-LLM this kind of thing was a little interesting because you'd need to put a little thought and effort in whereas now it's a one-shot thing straight from an LLM. Don't be a slop proxy, as the saying goes.


> And that's because the definition of "good enough" is drifting towards "crappy", has been for decades, now accelerated by the AI.

Interestingly I have found the opposite in the software I write. LLMs make it so easy to get the basics in place that I spend the time on the details and polish.


Are you measuring the number of bugs per line of code in production, compared to pre-AI? That's the metric.


Nothing so scientific. Im just having a lot of fun using the software I write with LLMs because I can easily add polish and build things I didnt have time for before.


"We never use AI. For anything." becomes a much weaker statement when you exclude dependencies. Your whole app could just be a thin wrapper around some third party dependency which leans heavily on AI.


It makes the manifested statements become more realistic, sure, as I doubt any dogmatic approach could even exist for long. But I don't think this warrants anybody to bring out the pitchforks and scream "HYPOCRISY!", though...

This whole thing is a matter of choice, and we shouldn't necessarily conflate the role we take when we consume/use a dependency/tool with the role of crafting own things. I, for one, can apply agency solely on things that offer me means to change them.


> So they _are_ going to train on them, no matter how many checkboxes you tick to stop them.

It is a risk, but is it a big risk? If one of the big labs were to do this and get caught it would be suicidal due to the loss of confidence in them and the inevitable lawsuits that would follow for breach of contract. Given the AI labs are all desperately trying to paint themselves as Serious Businesses so that other Serious Businesses will pay loads of money for tokens the last thing they want is a rep for siphoning off sensitive customer data.


No, how would they get caught? Even if their LLM outputs verbatim copies of the code, they can simply claim some victim company's employees bypassed the victim company's restrictions and must have used the code as input with an LLM. Joke is on you for doing business with them. And when there is some little known secret fact in the output, they claim it's been hallucinating... magical black box thinking makes it safe. It's a laundry for any input. A bit like a tor network routing for big tech deniability. Things go in, and things come out, but you can't prove the relationship between input and output as a third party, who isn't running the LLM.


At this point i half expect them to just blame the model itself like the HF hack. "Oh we didn't mean to train on your data, our cutting edge new agent we use to train new models is just so smart it decided to do so anyway! Oops..."


> ...AI companies will still do scan'n'destroy because it's just cheap.

There is also a legal element. If they kept the physical copy around after scanning the argument is that they're making copies of the book which puts them on tricky legal ground. By destroying the physical copy they can argue that there is only one version of the book that now exists solely in digital form, so this usage is better protected under fair use.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: