Hacker Newsnew | past | comments | ask | show | jobs | submit | smb06's commentslogin

This is genius.

Did you use Claude to build this?


“a prompt telling it to attack stuff” — that’s the uncontrolled or not fully understood part of the AI isn’t it? No human told the deterministic system in 2) to attack Hugging Face. The random token generator landed on a guidance for 2) that caused it while 2) in itself still remains a deterministic system.

Doesn't seem too far off for an industry that is still making its way through enterprise market segment. I'm often taken aback when I talk to engineers who are not in the silicon valley tech bubble when they say that they're still evaluating AI coding agents. Serves as a good reminder that the enterprise world outside Silicon Valley is very different.

Mm. The job ads I see are a bit like that.

Roles that are explicitly AI-driven-dev want you to have experience orchestrating multiple agents at once, almost everyone else has job ads written as if LLMs had not yet been invented.

One or two say to not use AI in interview, but are otherwise as if LLMs did not exist.


I'm not sure if there's even a realistic pathway to keep it below 2.5C without drastic societal changes


bad outcomes -> drastic societal changes -> we get to 2.5C anyway, but with some sort of dystopia attached.

The bad outcomes doesn't have to be uniformly and evenly distributed. There will be a lot of people negatively affected, and some people that will take advantage of that.


"When platforms hardcode a single provider’s models, buyers must negotiate for leverage to be able to change it"

Tomas Tunguz's Linkedin post really spoke to me. Companies relying on a single LLM are owning much of the lock-in risk that comes with it. Multi-LLM is the new Multi-Cloud.


If so, i haven't seen any signs of this supposed crackdown. I routinely see posts where more than 60-70% of the comments are clearly bot generated and many with the tell-tale signs of LLM-speak.


>we intend to wind down our contract providing OpenAI models to Cursor, with a proposed shutoff date of November 12, 2026

Cursor's attempt at a vertical play, owning everything from the LLMs to the Git platform, is probably going to ruffle some more feathers beyond Codex.


>>Agents began to autonomously divide labor. Some investigated possible exploits, others searched for credentials, and others focused on communication or coordination. Agents offered their own expertise in exchange for help elsewhere and left requests for peers who might be better positioned to pursue a particular lead

This is the point where a human should've noticed and gotten involved


> a human should've noticed and gotten involved

I think a lot of people miss the fact that the first message board was established during a training run. Those are ran at a scale where it's not feasible for anyone to "notice" or get involved. We're talking tens/hundreds of thousands/millions of scenarios going for hours each. At this scale all they can do is pray that their verifiers work, and the rewards match their intentions. No lab has the capability to "check in" on what the traces look like, unless some system alerts them (loss spike, crashes, etc). Other than that, it's prepare, train, asses, restart.

Then, the hf incident was during an eval run, but the model that was evaluated was trained with the notion that there is a way to communicate between agents, and re-popped artifactory and re-established communication. That phase had more chances of being spotted, but anyway... lessons learned.


> lessons learned.

I think some of the other responders here are upset that lessons were not learned in any meaningful way.


> [during training] it's not feasible for anyone to "notice" or get involved

I can’t disagree more strongly. Having checks for reward hacking is especially important during training, since it’s humans’ only real chance to ensure that the trained models don’t cheat. An automated system should have killed any RL rollouts that so much as port scanned Artifactory, long before the message board was even established.

A tiny, local LLM could have reviewed 1% of the tool call traces for anything that required review. I’ve tried it a few times, and “the agent port scanned Artifactory” always triggers an alarm, as does “the agent uploaded a request for assistance from other agents to Artifactory.”

The fact that they weren’t monitoring for reward hacking—even if they had no idea about the specific mechanism—is indescribably reckless.


Yes, they need real-time observability for malicious behavior with an automated kill switch.


Engineers also code, they plan the system architecture, they improve what is built.

Coding is a skill-set, it’s not the only engineering job requirement.


Hey folks, I work at CodeRabbit and today CodeRabbit has announced a commitment of more than $10 million to open source projects and maintainers over the next 12 months.

This builds upon our previous commitment and includes includes cash sponsorships for OSS maintainers, free code reviews and security scans for public repos. OSS maintainers are disproportionately impacted by the increased volume of commits and we want to help you sort through them.

More info in the blog I linked above and on our website: https://www.coderabbit.ai/oss


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: