Hacker Newsnew | past | comments | ask | show | jobs | submit | thom's commentslogin

This is sad to hear! Early in my career I worked for a company making airport collaborative decision making (A-CDM) software. One of the most impactful projects to which I contributed was a model that calculated the exact right moment for a pilot to turn their engine on given likely pushback times etc. That saved quite a lot of fuel (and therefore cost and pollution) back in the day. It's a real shame if similar thinking isn't being applied to the problem you describe.

How long a prompt do you think would be required to cajole an LLM into making legal moves at the rate of a human? Or do you think no amount of prompting could do that?

I don't know. My understanding is that current models will eventually fall into making illegal moves in longer chess games, and that no amount of prompting reliably gets them to stop doing so.

More importantly, beginner human players don't exhibit that tendency. The history of the position doesn't bother a human (except as required for castling and en passant rules), and the analysis becomes generally easier as pieces come off the board.

Humans do make these errors when playing blindfolded. If you even the playing field and give the LLM the position at each turn, it does not make mistakes.

> If you even the playing field and give the LLM the position at each turn, it does not make mistakes.

It absolutely still makes mistakes if you ask it to draw the board each turn, which should be equivalent to giving it the position because it only has to update one move at a time and then it has the position in the context window.


I've not noticed this happening if you give it the FEN each move. The alternative is just blindfold chess and very few humans can do that for long.

I haven't tried it myself, but people seem to report that the illegal moves surface eventually. It just takes longer: https://news.ycombinator.com/item?id=49720751

Nothing is forcing the LLM to play 'blind'. If it's smart, it should be able to create its own representation of the chess board and update it with every move, just like a human would. Any chess engine that's sensitive to how the moves are formatted is clearly not very capable.


A human wouldn't do that, they'd look at the board. I'm not disagreeing that to demonstrate clear superhuman ability the LLM should be able to do this, but it plays better than most humans blindfolded, and with fair prompts seems very good otherwise.

That's what a human will do if they already have a physical board to look at. But if someone, say, posed you a chess exam question via FEN notation, or as a sequence of moves in algebraic notation, you'd sketch a visual representation of the board off your own initiative to help you answer the question. There is nothing in principle to stop the LLM creating its own board representations in whatever format enables it to easily keep track of game state and legal and illegal moves. If it fails to do so, that's a sign of its own limited understanding of chess as compared to a human.

The LLM would only be playing 'blindfolded' if you somehow forbade it from making notes (as you effectively do by literally blindfolding a human, given how limited human working memory is). But you are not doing that. The LLM is free to keep track of the game state via whatever means it chooses.

None of this is about superhuman ability. Any human who understands a given chess notation can convert it to a visual representation of a chess board and then use that representation to choose their next move, with their usual level of performance.


I maintain that the amount of effort to teach a human to do this vastly outweighs the amount of effort to teach an LLM to do this unless you're deliberately trying to make them fail. I honestly have no bigger point than that, I just think this isn't a very good thing by which to evaluate LLM capabilities. If there's no argument you'll accept, I am happy to move on.

You don’t need to teach a human anything except the rules of chess and the details of a particular chess notation. No special skill or training is required to make a sketch of a chess board. Surely there is no chess player who, if confronted with a sequence of chess moves in algebraic notation, would not think to construct a representation of the chess board in order to understand what was going on.

> I just think this isn't a very good thing by which to evaluate LLM capabilities

I don’t think any single task is a good way to evaluate LLM capabilities, but I don’t see why chess is worse than a lot of other tasks. (Of course it is of no practical consequence whether LLMs can play chess, so if you are just making that point, then yes, I agree.)

> If there's no argument you'll accept

It’s a little unfair to suggest that I wouldn’t accept any argument whatever for your position just because I haven’t been convinced by your very brief comments so far. I could equally well say the same thing to you!


The actual question is backwards: how do we keep the prompt and context small enough so the LLM doesn't start hallucinating basic rules of chess.

Not sure what to answer on any of those questions. There's no 'perfectly clear but nothing like normal vision' option.

I really don't think this is the right definition. I don't think I have aphantasia: I picture things in my mind's eye, can manipulate those images, and I get the various stimulus responses described in some of the literature. But it's totally different than dreaming: it's an act of will, I can dismiss it at any time, and I would never be able to convince myself I was actually 'seeing' what I imagined.

But heck, if that's it then hail, brother.


It's absolutely not a definition - if you can actually picture things, you don't have aphantasia. I brought it up because it's surprisingly frequent that the people who doubt it exists turns out to doubt it because they don't actually believe people actually picture things, and that it's all just talking in metaphors.

It's what I used to think most of my life...

Note that some people who do picture things do visualise it so clearly that it is comparable to seeing, and that was my one singular experience with it as well, but that is the far other end of the scale. Most people will definitely be able to tell it apart from "actually" seeing, but it will be distinctly visual.

When I draw a comparison to dreaming it's about the images, not about being able to will it or dimiss it.


Gotcha. I _do_ have extremely clear auditory hallucinations, especially when I'm sleepy, but I was never talented enough to write them down properly, so I can believe more vivid visual ones are possible. Just always anxious that's what people mean during these threads.

Skald sold well which is stylistically very similar, and this has a slightly more popular confluence of genres.

AI curated lists of homogenous websites, flayed to their bare bones to inspire the utterly uninspired. The internet is better and weirder than this, even in these dark hours.

I remember losing sleep when I bought my RTX Pro 6000, but somehow they keep going up in price, and waddya know I’ve even done real work that appreciates the VRAM size.

Same. Had to convince the wife, initially aimed to get two of them, ended up with one. Now my wife is berating me for not convincing her that I should have bought four and then sell two later... Can't win :)

What about... queries?

Yes, this is called out near the end of TFA. Relational databases are great at allowing you to just come up with a new query that combines your data in a new way, express it in SQL, and get results near instantly. It's an architecture that is optimised for flexibility! I don't think I'd want to give that up without a very good reason.

How do you disseminate that information to humans?


I don’t like being this person, but you can very easily beat a wall at tennis.


What about a wall 1mm from the net


Yes? You just hit the ball so that every return is out. It’s very easy.


A curved wall?

Well look, the wall is already losing every point by volleying your serve so I feel like I'm on pretty strong ground here already. But give me an exact specification for your curved wall and I'll tell you how I'd beat it.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: