Did anyone read the blogpost? They did publish a pre-print:
> Our work to understand the primary function of ARTs is ongoing. However, we think it is important to share such findings early, both to demonstrate Claude’s capabilities and to give the broader community insight into what we’re working on. We have released a pre-print (here) that discusses this in more detail.
I just looked at it. I really hope they're not thinking of sending that to an actual bioinformatics, computational biology or molecular biology journal! So embarrassing...
(I love how Anthropic boast about building a lab, but don't seem to realise that you have to test your hypothesis in the lab! Right now, all their "spectacular" assertions are untested and unproven.)
I realise that this will only improve from here, but gods Anthropic has no idea about the biological sciences right now.
Um. Take a look at the authors. Every single one of them is an expert in this field. They did test this in a lab.
If you want to complain about things like this, it really helps to be specific. Given the author list, it's unlikely they made any truly spectacular errors (and also possible the system they studied is not interesting).
Can you elaborate on what you're thinking? I don't see any support for this being a hallucaination; from what I can see, it's a pretty typical "early biological discovery".
I was poking fun at evolarjun and epihelix for "hallucinating" that this was just a marketing whitepaper (they did publish a pre-print) and that the preprint was "so embarrassing" (you said that the authors are actually experts). It kind of seemed like they had some opinion about AI and this paper, and wishfully concluded things that were not true to support their opinion.
What's with all this pre-print business. It became very prevalent during covid, where it felt like every week some new pre-print was published that discusses some new aspect of the virus. These papers would then be used in arguments and put forward as proof of whatever claim the arguer was making.
Every man and his dog can publish a pre-print and in my opinion it's academically worthless.
Preprints are just a way to sacrifice rigor for accessibility and velocity. You can throw out "here, this is what I'm working on, here are the quick and dirty findings" really fast and with little friction.
This does skip the academic "checks and balances" like journal selection and peer review - but it can also help anyone else who's working on the adjacent topics.
If a field is moving fast, and you think there can be some value in your work for others in the near term? Preprint. If your work is too incomplete or too minor to warrant trying to polish and publish it, but you don't want to table it? Preprint. Too deep in corporate structures to care about academic "street cred", and want your work to be accessible? Preprint. Have an exciting early finding that you want to push out there, and are willing to take the rep risks of being wrong about it? Preprint.
There's a reason why preprints came to be the lifeblood of ML.
Academia isn't my thing but I also wonder if there isn't an aspect of putting a stake in the ground? So that if someone beats you to publishing you at least have some record of being on that track.
That is definitely a big motivator for publishing preprints. Journal submissions can take up to a year. Comference submissions take months. If the field is moving fast, claiming a finding early can become an important career move.
In older days, academics would just share notes on their work and word wouldn't usually spread widely before publication.
Preprints may be the better model. But public visibility means that non-experts now get to see the good and the bad research equally, but they won't have the domain knowledge and skill to distinguish one from the other with confidence.
I have a similar to pagerank method I use to evaluate such papers. I look for the references to see how many authors are using their own references (past work), the idea being that people do not jump too far, they make incremental progress.
For the pre-print I could only find only one author who has a single referenced article.
> references to see how many authors are using their own references (past work), the idea being that people do not jump too far, they make incremental progress
The authors are not using their own prior work in the paper, thats the point I was trying to make. I have worked in biotech lab for couple years and its one of the criteria's people use to consider some ones work useful and worth the time.
I think you fundamentally misunderstand both the results published here, and how scientists operating at the highest level of academic research operate. None of these authors has to worry about citing previous work to get the attention of biotech labs.
Journal articles are usually behind a paywall. Preprints are a fully open workaround allowed by most journals.
> Every man and his dog can publish a pre-print and in my opinion it's academically worthless.
Sure but if you look at the authors names and see they have 50 other published papers, you can get a rough idea that it's probably equivalently good to their other work.
Until you've done it yourself, it's hard to grok just how bad the peer review process is. It's like...5% better than nothing.
Honestly you could argue peer review is worse than nothing, as it also filters out actually quality work that violates some dogma of the field.
The review time on top journals is multiple years now (your paper will go through many review loops each of which takes months). It's just totally unworkable for active research, whether you're a student or a corporation.
> "I think people generally choose Deno because of the stability and security-first architecture."
The appeal of Deno to me was always it's Go-like `std` library, so that building basic apps didn't result in a spaghetti web of dependencies like in Node.
Can you expound on why "This is one of the most ridiculous statements I've ever read."? I think a deeper discussion into the capability of agents to decrease entropy would be very valuable.
I disagree. A program/repository is not a big chunk of text. It's more like logic to perform a particular function represented in text which is a symbol of language.
"To think that reducing entropy/complexity in a system is some inherently human capability is both anthropo-centric and full of hubris."
Humans invented LLMs and Agents etc, designed, programmed, etc. etc. Seems they are more an extension of humans -- tools that rely on humans.
An LLM/Agent left to it's self on a code base does cause entropy. I've only ever seen them not cause entropy on the first prompt.
That you can go 300km/h in an F1 car, doesn't mean you can go that fast on a bicycle, even though they're both in the genus "wheel-based vehicles for human transport". The jellyfish is a rather different animal. Trees can grow very old, but there's a reason why you didn't use those as an example: too far from human biology. That jellyfish may be at a similar distance.
It would be a scientific process to find the limits, but I don't see a scientific reason to find out how to overcome them. Science' job description is oberving and explaining, not altering nature.
We have the ability to convert to becoming immortal. It actually happens to some of us. It sucks though. We call this cancer. Turns out the immortal cells become selfish. But you can scrape some off yourself, throw them in a petri dish, supplement growth media, and they will divide forever.
> Ten minutes earlier, the A.I. agent had called my dental insurer, given them my member I.D., answered a security question and waited on hold before transferring me to the representative
I have been willing to pay a lot of money for a mobile AI tool that can make outbound calls for me.
I dread making medical appointments or other reservations. Didn't know Muse could do this, will have to give it a try.
Yeah I totally want clankers handing down my data left and right. Any reasonable medical provider lets you book via app anyway if you are either too scared to call or don't want to wait.
The "zero day" is something they call a "ClickFix Attack"
Upon Googling "ClickFix":
> "A ClickFix attack is a social engineering technique... It typically compromises devices by manipulating victims into copying and pasting malicious commands directly into system-level tools"
I'm sorry, that's not a zero-day, that's idiocy that's as old as time.
The exploit is a local zero day exploit, meaning the machine needs to already be compromised. For it to be a remote zero day exploit you need to do the ClickFix attack. The idea is that it lets you access much more machines and resources if the one machine with Muse is compromised. More info here: https://x.com/dps/status/2102248329111634067
We filed a bug report but the original maintainer seems to have dropped offline. The community's had some success in correcting bugs with low-level hacking, but it's hard to make progress without the source code.
> Postgres makes those files searchable. memory.entries stores chunks and line references, memory.embeddings holds 384-dimensional vectors, and memory.claims tracks evidence, confidence, and status.
Is each Muse instance running it's own Postgres??
That seems wildly wasteful, especially since earlier in the article it states that the Muse instance has a SQLite database and schema already...
Wasteful but FAR more secure. Surprisingly secure. If they wanted to follow through on their promise of sandboxed and encrypted data which even they couldn’t access, this is one way to do it.
E2E encryption, but I don't know the architecture. I.e. Meta doesn't have the keys. My comment above is based on an interview I saw with Zuckerberg where he wanted to use the Signal model to ensure the data was secure "even from Meta."
It's the best free vps on the market right now. Comes with a coding harness and a few hundred million tokens on a decent model. Can't last long but it's fun.
I’m trying to rein in my hyperbole but if Meta is able to follow through on the promise, Muse is world-changing. I share the scepticism. How on Earth can they offer this for free?
They have paid subscriptions and it gets you using other Meta services more so they make money from ads. If that's not enough they might put ads and affiliate links when you use it for shopping.
The VM is about the same specs as a 9EUR/month VPS from Hetzner but it only runs a fraction of the time, they spin it down when Muse stops using it. Meta is a hyperscaler with optimized infra and they don't need to make a profit because VMs are not their actual service. I think it costs them at least 10x less than inference for most users.
A few months ago I asked why semantic representation rather than text wasn't used, since natural language seems quite a lossy representation for semantic concepts:
Indeed this is one of the key problems here. If models start doing CoT in "neuralese", or use it when talking with each other, we lose what little observability we have.
reply