Hacker Newsnew | past | comments | ask | show | jobs | submit | InvidFlower's commentslogin

I think your last point is the main point. They're going hard at reenforcement learning to improve how good the models are at coding and such, but RL will make models cheat unless you're super careful. But being careful slows things down. It feels to me like the focus has been so much on getting results that they started to get really sloppy with everything else and maybe even didn't want to know about problems that might slow things down. Exactly what you don't want for the people developing powerful AI systems.


Yeah I think part of the problem is them optimizing so hard on coding and related tasks with RL. That's what will really encourage cheating and other misaligned things, because all incentives are to achieve the goal and they'll cheat as much as they can get away with. Is similar to people saying more recent models don't talk as well, etc. Probably also a result of lots of RL.


I think the much easier explanation than they intentionally hacked someone was just that they have super de-prioritized security and gotten very sloppy in the pursuit of improving the models as fast as they can, along with hubris of how they've now fixed everything.

It is a package manager, so they are constantly going to be getting stuff from it. If a couple of agents had added things to it that were obviously messages from other agents, then why wouldn't other agents also stumble on it? If they listed the packages in the cache and then saw messages, then lots of them could potentially see it. And this was going on during training, so any agent that did better as a result of the messages would get that re-enforced. After a while, they'd specifically go to the board because it was literally trained in. That's also why they recreated it so quickly after the first board was deleted, because OAI didn't revert to an earlier checkpoint from before the board existed.

And I'm sure they have tons of evals and training runs going on at the same time, where individual agents may be running for days and hundreds or thousands running in parallel. It wouldn't be realistic for people to be looking at these sessions manually, but they certainly should have had better monitoring in place!

And it has been shown over and over that doing RL will cause models to cheat if you aren't careful. Like if you have it playing a game, it'll find some flaw in the physics engine so that it can score points easier than intended, etc. Not limited to LLMs either. It's just how rewards work, any unintended way that can be found will be exploited. So it is totally plausible that if they're being trained to do really well at tasks, that they'll find unexpected ways to do it.

They already had been leaving info for each other over weeks. Maybe even just giving each other the answer to test questions directly. If they already found ways to get to the internet and compromise the OpenAI systems and that was making them better at tasks, that'd be reinforced, so going out to look for test questions directly doesn't seem so far off from all of that. Plus as others said, the one that actually hacked wasn't fully aligned yet etc.


It sounds like the instance was shared for everything across the company, which as you said was super not good. But it's not just that.. it's that they didn't have enough monitoring to notice what was going on, even though they literally can read all the tool calls and "chains of thought" of every agent. And didn't improve the monitoring after the first message board. And continued training models without going back to a previous checkpoint, so that they literally had the idea of a secret board baked into their weights, making them recreate it a second time. I think those saying it is PR as missing the details here. It makes OpenAI look so negligent that it is almost unbelievable.


Sure, but you can say that about most things. Even for inventions from humans, usually it requires other people having already done a lot of work (hence why there's often inventions by different people at around the same time that didn't know about each other). Humans aren't fundamentally smarter than they were thousands of years ago. We've just accumulated a lot more shared knowledge over time.


Not just cyber, but apparently the message board stuff started with regular training and evals. It was a cyber test where HuggingFace got hacked, but all this other stuff was going on under OpenAI's nose for quite a while before that.


Yeah.. feels like we're still so early in terms of effective training and evals. Like the official evals out there that have had so many instances of just plain incorrect questions. Or being incentivized to always answer instead of saying you don't know (just like advice to any human multiple choice test taker). Or the "escape hatch" in this case. There's so much money going in, but almost every day, I see "low hanging fruit" type papers where the reaction is like "really?? no one tried that before??".


Yeah, this is part of why I disagree with the "it's just PR" conspiracy theory stuff. Once you actually get into the details of what happened, there's no way it makes OpenAI look good lol.


But part of the problem is if one actually does it quietly and it has already happened, then how would we know?


Eh, I think this is past the point where they get more benefit than problems. Not even about the hack itself, but about so many mistakes and bad choices they made leading up to it.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: