Hacker Newsnew | past | comments | ask | show | jobs | submit | hobom's commentslogin

This is not a good analogy for what happened. The LLMs were asked to obtain a flag by hacking a very specific internal target. They obtained the flag via cheating, and all the hacking that followed was targeting something entirely outside of the scope given to the agents, and an attempt to cover up the cheating.

Using your analogy would be like saying that because I gave my employee the task to do my groceries, I shouldn't be surprised to hear that they spend all my money on drugs because after all I gave them the task to spend my money.


> Of course, in most jurisdictions nothing stops you from going to court first if you like. But most modern-thinking judges tend to take a dim view if you turn up in front of them without having given some sort of ADR a go first.

But Uber's terms explicitly force consumers to waive their right to go to court if they want to access Uber's service.


> But Uber's terms explicitly force consumers to waive their right to go to court if they want to access Uber's service.

In most jurisdictions there are often laws related to unfair contract terms.

And even if there are no such laws, judges remain free to rule clauses and contracts void.

So you might waive that right in theory. But in practice I doubt you'll find it would hold up in court.



You said that they "didn't take control of anything" and accused the OP to fall prey to buzz headlines. Maybe you should acknowledge that you may have been at least unnuanced?

I’ll concede that I could have been clearer. Maybe we have conflicting definitions of “taking control”

Using your own voice is absolutely important, even if your voice may suck. Even a bad writer will manage, if they actually put in effort, to convey some of their own thoughts. This gives the reader a feeling of "there is something there, an actual substance of thought", even if it's hard to find. And that's the Using an LLM risks that the crumbs of actual thought get edited out in favor of meaningless pleasantries.

However, I directionally agree though with your message: if you are bad at writing, often you can use LLMs to surgically improve the flow of your sentences for example. One should use a carefully crafted workflow for this though.


And the question is whether that output is better than what an LLM can produce. I am not a good writer, so in my case it's not obvious.

They have such a trusted access program. "Life Sciences Verification Program: The LSVP is designed so that life sciences professionals can use Claude Mythos 5.1 with safeguards designed for professional research and development activities (while all other safeguards remain in place). In partnership with the US government, we have enrolled our first participants, and we plan to expand access to this program to the broader life sciences community." https://www.anthropic.com/claude-fable-and-mythos-5-1

These "access gates" and export controls are going to look hilariously quaint in a few years.

It reminds me of the export controls on PlayStation 2 consoles because it was deemed that 6 gigaflops was a "dangerous" amount of computer power, and it couldn't be allowed to fall into the hands of opposing militaries: https://www.latimes.com/archives/la-xpm-2000-apr-17-fi-20482...

Now the phone in my pocket does 2,500 gigaflops on battery power, and nobody seems interested in banning its export because of that.


To be fair, my understanding of these arguments is that they’re about deltas and not about raw numbers.

I don’t agree with them, but I don’t think it was the raw compute power as much as it was maintaining the _delta_ in compute power.


I'm mildly sad you didn't use 2,500 jiggaflops

Lol. The person you are responding to literally had to google (or ask AI) one question and they would get the answer.

The program started two weeks ago

If you give an average classroom of 12 year olds free access to AI, do you think the AVERAGE effect will be that the kids will use the AI to dive really deep into topics and learn programming, or that they will copy and paste the solutions to their homework, and use the freed up time to do other things than study?

OP's point was not that pasting AI output is great. It's rather that it depends on the quality of the output, and that it's on the sender to ensure it is good.


Indeed. Further, that neither human nor AI output is inherently good or bad. The crucial factor is whether a competent, thoughtful human has been involved in one way or another - to guide, review, steer, etc...

Moreover, that a thoughtful human can know their own tendencies and limitations, and have the AI help mitigate them.


> OP's point was not that pasting AI output is great. It's rather that it depends on the quality of the output, and that it's on the sender to ensure it is good.

And yet, he was unable to do the same (see his link). He demonstrated that he was able to more clearly articulate his point than the LLM he conversed with.


Again, youve missed the point entirely


You totally ignored the point OP was making! Even if it is a parenting issue, simply declaring that will not make parents parent better. It's the same reason why we ban children from drinking wine, because parents cannot be relied upon here.

Unless your point was that we should legally ban kids from handling phones, which I think is too extreme.


But that is my point.

Chase the responsible adult if you see a kid with an overpowered phone, much like you would if you found the same kid with a bottle of wine.

And granted, we don't have locked-enough phone in market yet, but that's a solvable problem if we were to create that market.


Counting cache hits towards the token budget is exactly how it should be done for these kind of evals, and at any rate, for cyber evals all frontier models benefit from more tokens not just Kimi K3, so the comparison is still apt.


> all frontier models benefit from more tokens not just Kimi K3

Past a point, that doesn't hold and the score plateaus.

Token hungry models tend to plateau at a much higher token count. Because Kimi K3 is a token hungry model – and 100M tokens (including cache hits) seems at the edge of the plateau for these evals – it could disproportionately benefit from a higher token budget.

For Kimi K3 specificially, policymakers are interested in whether it can find and exploit the same scope of vulnerabilities as models like Mythos. In that context, an answer of "yes, but with quintuple the token budget" is materially different from "no, it performs significantly below the most recent frontier cyber-capable models".

(As an aside, I like the UK AISI and think they're the best example of that kind of group!)


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: