Hacker Newsnew | past | comments | ask | show | jobs | submit | gavinray's commentslogin

Did anyone read the blogpost? They did publish a pre-print:

  > Our work to understand the primary function of ARTs is ongoing. However, we think it is important to share such findings early, both to demonstrate Claude’s capabilities and to give the broader community insight into what we’re working on. We have released a pre-print (here) that discusses this in more detail.
https://www-cdn.anthropic.com/22573675ada52a8ca8a97a1a4b4326...

I don't recall that being there when I first read their post a few hours ago... No way to confirm, sadly, as they've posted it to their own domain

Earliest snapshot is from 3 hours ago, more or less, it already shows the pre-print: https://web.archive.org/web/20260923180804/https://www.anthr...

archive.is has a snapshot from 18:47, showing the link was indeed there a few hours ago. https://archive.is/XM0Nw

I just looked at it. I really hope they're not thinking of sending that to an actual bioinformatics, computational biology or molecular biology journal! So embarrassing...

(I love how Anthropic boast about building a lab, but don't seem to realise that you have to test your hypothesis in the lab! Right now, all their "spectacular" assertions are untested and unproven.)

I realise that this will only improve from here, but gods Anthropic has no idea about the biological sciences right now.


Um. Take a look at the authors. Every single one of them is an expert in this field. They did test this in a lab.

If you want to complain about things like this, it really helps to be specific. Given the author list, it's unlikely they made any truly spectacular errors (and also possible the system they studied is not interesting).


This comment chain is depressing. People who hallucinated things they really wanted to be true. Guess it’s not a uniquely AI problem.

Can you elaborate on what you're thinking? I don't see any support for this being a hallucaination; from what I can see, it's a pretty typical "early biological discovery".

I was poking fun at evolarjun and epihelix for "hallucinating" that this was just a marketing whitepaper (they did publish a pre-print) and that the preprint was "so embarrassing" (you said that the authors are actually experts). It kind of seemed like they had some opinion about AI and this paper, and wishfully concluded things that were not true to support their opinion.

What's with all this pre-print business. It became very prevalent during covid, where it felt like every week some new pre-print was published that discusses some new aspect of the virus. These papers would then be used in arguments and put forward as proof of whatever claim the arguer was making.

Every man and his dog can publish a pre-print and in my opinion it's academically worthless.


Preprints are just a way to sacrifice rigor for accessibility and velocity. You can throw out "here, this is what I'm working on, here are the quick and dirty findings" really fast and with little friction.

This does skip the academic "checks and balances" like journal selection and peer review - but it can also help anyone else who's working on the adjacent topics.

If a field is moving fast, and you think there can be some value in your work for others in the near term? Preprint. If your work is too incomplete or too minor to warrant trying to polish and publish it, but you don't want to table it? Preprint. Too deep in corporate structures to care about academic "street cred", and want your work to be accessible? Preprint. Have an exciting early finding that you want to push out there, and are willing to take the rep risks of being wrong about it? Preprint.

There's a reason why preprints came to be the lifeblood of ML.


Academia isn't my thing but I also wonder if there isn't an aspect of putting a stake in the ground? So that if someone beats you to publishing you at least have some record of being on that track.

That is definitely a big motivator for publishing preprints. Journal submissions can take up to a year. Comference submissions take months. If the field is moving fast, claiming a finding early can become an important career move.

In older days, academics would just share notes on their work and word wouldn't usually spread widely before publication.

Preprints may be the better model. But public visibility means that non-experts now get to see the good and the bad research equally, but they won't have the domain knowledge and skill to distinguish one from the other with confidence.


I have a similar to pagerank method I use to evaluate such papers. I look for the references to see how many authors are using their own references (past work), the idea being that people do not jump too far, they make incremental progress.

For the pre-print I could only find only one author who has a single referenced article.


The pre-print authors all have long publication histories.

Pagerank was inspired by academic citation networks; it just turned it in a recursive matrix problem (of which there was some prior literature).


> references to see how many authors are using their own references (past work), the idea being that people do not jump too far, they make incremental progress

The authors are not using their own prior work in the paper, thats the point I was trying to make. I have worked in biotech lab for couple years and its one of the criteria's people use to consider some ones work useful and worth the time.


I think you fundamentally misunderstand both the results published here, and how scientists operating at the highest level of academic research operate. None of these authors has to worry about citing previous work to get the attention of biotech labs.

Journal articles are usually behind a paywall. Preprints are a fully open workaround allowed by most journals.

> Every man and his dog can publish a pre-print and in my opinion it's academically worthless.

Sure but if you look at the authors names and see they have 50 other published papers, you can get a rough idea that it's probably equivalently good to their other work.

Until you've done it yourself, it's hard to grok just how bad the peer review process is. It's like...5% better than nothing.

Honestly you could argue peer review is worse than nothing, as it also filters out actually quality work that violates some dogma of the field.


The review time on top journals is multiple years now (your paper will go through many review loops each of which takes months). It's just totally unworkable for active research, whether you're a student or a corporation.

Just going to post this. And I think it's safe to say that no, they didn't read it.

  > "I think people generally choose Deno because of the stability and security-first architecture."
The appeal of Deno to me was always it's Go-like `std` library, so that building basic apps didn't result in a spaghetti web of dependencies like in Node.

  > "Agents can only maintain or INCREASE entropy in a system. Humans are uniquely capable of decreasing it"
This is one of the most ridiculous statements I've ever read.

Can you expound on why "This is one of the most ridiculous statements I've ever read."? I think a deeper discussion into the capability of agents to decrease entropy would be very valuable.

A program/repository is a big chunk of text

Any text edits (insertions, deletions) that a human is capable of doing, so is an agent/LLM

To think that reducing entropy/complexity in a system is some inherently human capability is both anthropo-centric and full of hubris.


I disagree. A program/repository is not a big chunk of text. It's more like logic to perform a particular function represented in text which is a symbol of language.

"To think that reducing entropy/complexity in a system is some inherently human capability is both anthropo-centric and full of hubris."

Humans invented LLMs and Agents etc, designed, programmed, etc. etc. Seems they are more an extension of humans -- tools that rely on humans.

An LLM/Agent left to it's self on a code base does cause entropy. I've only ever seen them not cause entropy on the first prompt.


The Immortal Jellyfish, and functionally the naked mole rat.

That you can go 300km/h in an F1 car, doesn't mean you can go that fast on a bicycle, even though they're both in the genus "wheel-based vehicles for human transport". The jellyfish is a rather different animal. Trees can grow very old, but there's a reason why you didn't use those as an example: too far from human biology. That jellyfish may be at a similar distance.

The process of science is to determine what exactly these inferred limits actually are and how many can be overcome.

It would be a scientific process to find the limits, but I don't see a scientific reason to find out how to overcome them. Science' job description is oberving and explaining, not altering nature.

In some sciences like astronomy you just really don't have the the ability to.

With medicine it's really the entire point, it's science we learn to improve humanity.


Perhaps that explains why medicine is often such bad science. It's more engineering, really.

We have the ability to convert to becoming immortal. It actually happens to some of us. It sucks though. We call this cancer. Turns out the immortal cells become selfish. But you can scrape some off yourself, throw them in a petri dish, supplement growth media, and they will divide forever.

*sucks for the rest of the body. I'm sure the cancer itself is pretty "happy".

The lobster, though they die from being too large to molt.

  >  Ten minutes earlier, the A.I. agent had called my dental insurer, given them my member I.D., answered a security question and waited on hold before transferring me to the representative
I have been willing to pay a lot of money for a mobile AI tool that can make outbound calls for me.

I dread making medical appointments or other reservations. Didn't know Muse could do this, will have to give it a try.


Yeah I totally want clankers handing down my data left and right. Any reasonable medical provider lets you book via app anyway if you are either too scared to call or don't want to wait.

See also: "meat proxy", "reverse centaur"

The "zero day" is something they call a "ClickFix Attack"

Upon Googling "ClickFix":

  > "A ClickFix attack is a social engineering technique... It typically compromises devices by manipulating victims into copying and pasting malicious commands directly into system-level tools"
I'm sorry, that's not a zero-day, that's idiocy that's as old as time.

The exploit is a local zero day exploit, meaning the machine needs to already be compromised. For it to be a remote zero day exploit you need to do the ClickFix attack. The idea is that it lets you access much more machines and resources if the one machine with Muse is compromised. More info here: https://x.com/dps/status/2102248329111634067

We filed a bug report but the original maintainer seems to have dropped offline. The community's had some success in correcting bugs with low-level hacking, but it's hard to make progress without the source code.

> > It typically compromises devices by manipulating victims into copying and pasting malicious commands

> […] that's idiocy that's as old as time.

And constantly being reinvented: `curl -sL some.unverified.stuff.sh | bash`


Reminds me of playing Runescape back in the late 90s, you could double your money by just pressing alt-F4

Well if nothing else, idiocy that's as old as time started on day zero.

words don't seem to mean anything anymore.

clickbait headline should be changed, not a 0-day.


  > Postgres makes those files searchable. memory.entries stores chunks and line references, memory.embeddings holds 384-dimensional vectors, and memory.claims tracks evidence, confidence, and status. 
Is each Muse instance running it's own Postgres??

That seems wildly wasteful, especially since earlier in the article it states that the Muse instance has a SQLite database and schema already...


Wasteful but FAR more secure. Surprisingly secure. If they wanted to follow through on their promise of sandboxed and encrypted data which even they couldn’t access, this is one way to do it.

> which even they couldn’t access

What would prevent them from getting the data out of postgres?


E2E encryption, but I don't know the architecture. I.e. Meta doesn't have the keys. My comment above is based on an interview I saw with Zuckerberg where he wanted to use the Signal model to ensure the data was secure "even from Meta."

Looks like RAM shortage hasn't hit meta yet. They are probably burning money to grab userbase.

My instance claims the container has 2 vCPU and 8gb of RAM. I got it to set up a Minecraft server with access over Tailscale

It's the best free vps on the market right now. Comes with a coding harness and a few hundred million tokens on a decent model. Can't last long but it's fun.

I’m trying to rein in my hyperbole but if Meta is able to follow through on the promise, Muse is world-changing. I share the scepticism. How on Earth can they offer this for free?

They have paid subscriptions and it gets you using other Meta services more so they make money from ads. If that's not enough they might put ads and affiliate links when you use it for shopping.

The VM is about the same specs as a 9EUR/month VPS from Hetzner but it only runs a fraction of the time, they spin it down when Muse stops using it. Meta is a hyperscaler with optimized infra and they don't need to make a profit because VMs are not their actual service. I think it costs them at least 10x less than inference for most users.


Hyperion.

I guess I'm poor then because my homelab has nodes with these specs.

ECM (Extracellular Matrix) has been studied for similar properties:

https://www.biorxiv.org/content/10.64898/2026.08.05.742898v1...

"Extracellular matrix particle treatment induces digit regeneration in soft-tissue preserved amputation (SPA) model of adult mice"


A few months ago I asked why semantic representation rather than text wasn't used, since natural language seems quite a lossy representation for semantic concepts:

https://news.ycombinator.com/item?id=47195212

I wouldn't have thought to use it for LLM-to-LLM communication, though


My guess is that if you go with something other than readable text as the “thought layer”, observability becomes impossible.

Which is not necessarily something the humans training a model would want re/ alignment.


Indeed this is one of the key problems here. If models start doing CoT in "neuralese", or use it when talking with each other, we lose what little observability we have.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: