> Why do you think this uncracked code was so simple to solve?
I never said this. All I said is we don't have the conversation and therefore we can't determine how easy or hard of a problem it was.
> And it's very possible Terra or GLM could crack it, turn off their web access and try yourself.
I'm questioning why this should be labeled "Astra" breaking anything implying it required "the best" model to do it when in fact any other half-decent model might have been able to do this as well.
EDIT: Okay seems like the actual prompt is published, just not on the same article that was linked. Maybe I'll give it a try.
which is crazy because this was grok's competitive advantage, worse than OpenAI models but better than everything else, now it's less efficient than Opus or Fable 5.1
the AA numbers are generationally bad. double token use (the one thing Grok was good at was low reasoning usage!) to gain 5% in the benchmark score. with reportedly a larger model. maybe it shows gains IRL but wow, I've never seen a new generation model look so underwhelming compared to the last.
It still works, bots can solve it but it probably increases the cost of that web call by 10x or 100x for that bot, so it won't bother. Had a recent bad experience with removing recaptcha.
I get reCaptchaed-to-death all the time in iOS and MacOS using both Firefox and Brave. I use a big name VPN which probably makes it worse, since I'm routinely blocked outright by Cloudfare services, assuming cluelessly that I'm a bot.
My assumption is in the long term that efficiency will break interpretability, which will lead to questionable alignment. As efficiency drives capitalism and evolution we'll run headlong at it and try to deal with the risk as a side effect.
reply