[I work on Claude Code] I broadly agree with the author’s point: plan mode was useful, and is no longer useful.
In Claude Code, all plan mode does is add a little reminder to every user message along the lines of “you’re in plan mode, please don’t code yet”. It’s something I came up with late on a Sunday night many months ago, when I got tired of asking Claude to plan with me first before coding in each new session. Something people might not realize is plan mode has always been a prompt — it has never changed the toolset because doing so would break the prompt cache, and so would be expensive for users.
This worked well for a while, until a few months ago, using early versions of Fable, I realized that I wasn’t using plan mode anymore because the model just got it, and because for the increasingly complex work I asked the model to do, planning had become interactive and iterative. With Opus 5.5, I feel Opus has gotten to that point too.
For codebase understanding, I sometimes ask Claude to generate an artifact that explains some aspect of its changes. For complex diffs to core parts of the system, I will often ask it to make diagrams or even interactive demos so I can better understand the change and alternatives considered. I don’t do this very often, but it’s a useful way to explain code when you need it. I ask Claude to attach these artifacts to its PRs also, so others can understand and future Claudes have the context.
It looks like a primitive chess engine, trying every move, no matter how stupid, until it works. Relying on its ability to do millions of operations rather than having a plan.
People will try stuff too, but once there is an opening, they will consolidate, generalize, simplify,... before going to the next step. The agents didn't, it is a huge, vaguely directed mess.
Also, it looked so "loud", querying millions of URL with weird requests. The sandbox as weak as it can get, and there is absolutely zero smart extrusion detection or it would have found it. They used their best AI for attacking, but nothing for protection.
A certain nuclear power plant had a Windows NT 4.0 machine running as late as 2007. The reason is interesting.
The machine's purpose was to report status of the control rods that mitigate nuclear reactions. Basically, "are the rods inserted, and if so, how many / how far?". I want to emphasize that this was reporting only, NOT control.
The original software was written back in the 80's, when the plant was originally commissioned, for AmigaOS. Of course, it's hard to buy Amigas anymore, and the original one died long ago (nobody remembers when).
So in the mid '90s, the utility purchased an AmigaOS emulator that ran on Windows NT 4.0, which was current at the time. The emulator (IIRC) was developed by a firm in the UK. The firm went out of business sometime in the late '90s. The control rod monitoring software ran under this emulator on top of NT4.
Windows NT 4.0 was the last OS to allow the emulation software direct access to the physical hardware that produced the status signal. Later versions of Windows abstracted the hardware access away, and the monitoring software broke. Because the emulation company had gone belly up, there was no way to fix the incompatibility.
So the utility had a choice: get new hardware/software certified (by NRC?), or keep doing what they were doing with the software (and hardware) that they had. They chose the latter.
So this is how, in 2007, during a tour of the facility, I stumbled across a Pentium 1 system running an AmigaOS emulator on Windows NT 4.0 that was responsible for displaying the status of the control rods of a nuclear power plant.
Spare hardware for this setup was purchased off of eBay and stocked on an adjacent shelf.
I think it bothers OP less that they take a 15% tax than the fact that google provides terrible support for their own play store. If they would take that tax and provide good feedback and speedy version reviews, nobody would ever complain - it is expected to pay something since the play store doesn't run on good thoughts and prayers.
But because they're a monopoly (or a duopoly if you count apple, which is a different platform altogether) they can afford to act this way.
I'm actively watching understanding slip away from developers, code review getting paired down to no comment checkmarks, and codebases go to bloated messes that nobody can read. Axioms like engineers must understand and take responsibility for the code they ship are getting torn down, and the products coming out are reflecting conway's law, becoming impenetrably obtuse and always "so complex there are no obvious deficiencies" (as opposed to "so simple there are no obvious deficiencies" which used to be the aim).
The one thing plan mode helped is for the humans to get an understanding of the strategy, and be able to poke around and look at the design and architecture. You can achieve this with some self discipline and keeping shorter leashes on agents, but it feels like a losing battle. The best devs still put out good code, but the poor devs are learning nothing while their metrics look great. I can't help but think we are racking up immense amounts of debt that will very soon become due.
My biggest takeaway from this is just how godawful the sandboxing is. The stuff written up in OpenAIs report says more about lack of extremely basic sysadmin skills than anything else.
I’m not that surprised about models with endless compute being capable of this, I’m more surprised that a company with the resources they have apparently can only create a sandbox that a half skilled human operator could have broken out of easily.
Getting bored of these framings where the superintelligent sentient beings running freely inside OpenAI are doing things that the company has no control over. The headline should be:
OpenAI meddled with multiple US Government agency sites.
The bots are acting neither properly nor improperly, they’re acting as they’re being allowed or coordinated to act.
> Before approving construction, I would want communities of humans to understand why the design works and what justifies confidence in its safety. I would hope that we all would.
Until very recently, I pored over every single line of code Claude generated with razor sharp scrutiny. I would usually catch issues with every response. I'm catching fewer problems these days. Maybe the model is just getting better, and maybe I'm being less careful while under pressure to ship more and more often. But model capability is obviously growing. Even back in March, you could tell it "give me a function that adds two numbers" and you could be 100% confident that it would write the correct function. There was almost no point in looking at the code. Since then, the complexity floor of problems in the category "this is so simple that the model couldn't possibly get it wrong" is rising, and with it, my cognitive surrender to the model is increasing too. Why check it? It's obviously going to be correct.
If AI designs a terawatt fusion plant, then of course we're going to meticulously pore over every detail to ensure safety, reliability, efficiency, whatever. If we find no flaws in the design whatsoever, will we be less careful about the second one? The third one? What about the ten thousandth one? Will "a nuclear fusion plant" become something that models couldn't possibly get wrong?
Terence Tao is arguing that the human involvement in research is crucial, but doesn't convincingly justify why, in my opinion. He says that "human agency is a value of fundamental importance" and that we will need to build "thriving human communities that can understand [AI ideas] together" - not for the sake of correctness, which AI may surpass us on, but for, I guess, the possibility of reclaiming human meaning and purpose. I don't disagree with this at all, but it's not an argument, it's a statement of values. Unfortunately, the stark reality is that if AI does surpass humans, it will become the economically dominant strategy to not verify them and not double check them, but to just do whatever they say. This seems like a great way to raise p(doom). But as the models get better and better, and as I'm scrutinizing Claude's output less and less... I just hope that there are more Terence Taos out there than people like me.
I’ve since learned that overly online keyboard warriors are the least important people to convince. The most important people are on your local city council or state legislature. Furthermore, instead of wasting time with hard cases or elected officials who are dead-set against your ideas, you’re much better off finding and working with people who are already interested in your principles. This often means giving up on your own city or state, at least for the time being, and pursuing a “succeed anywhere” strategy instead.
We're about to have a board vote eliminating single-family zoning (I expect us to win) and this rings extremely true about local politics. If you count "votes" in Facebook comments --- Facebook is where all our local politics happen --- you'd think this is the most contentious issue in the world. But in our last mayoral election, you'd draw the same conclusion about the race being tight, and in fact the progressive rinsed the moderate by a huge margin.
The same thing is very true of public meetings; the turnout for comment at those things is almost never representative of the area's median sentiment. "No" turns out for everything; "yes" rarely does.
I spend a lot of time in online arguments with opponents, on the theory that I'm not really writing to the person I'm arguing with but rather to everyone quietly reading comments. But it's almost certainly the case that the only conversations that have really mattered have been with electeds themselves.
It's just sad that this is the top comment on Hacker News. Why are we giving free pass to these tech companies? Why are we trusting these CEOs when they have repeatedly broken laws? Remember Aaron Swartz and the fate he suffered? Why is big tech getting away with so much more?
A little over two decades ago, my then girlfriend was arrested for "writing malware" (which was not against the law at the time, and which was never released into the wild and never caused any damage). This set in motion a chain of events that effectively ruined her life.
Fast forward to today, and we have multi billion dollar corporations pumping out malware at breakneck speeds, compromising various systems (including those of foreign governments), and no one is getting arrested. Instead we're gawking at the marvel of these systems and are playing word games about whether or not it's a rogue system. If anything, it's making people richer.
Imagine having a virus escape a sandbox, why are we worried about the virus but not the incompetency of those who are responsible for setting up the sandbox?
If I post something on the Internet today claiming that I asked my agent to do X but it went rogue and did Y, all I will be getting in return is a jar full of "skill issue".
Should we worried about people using LLMs for attacks? Yes, but not in the premise of LLMs going rogue but someone with the intention of abusing it to cause harm. And this is not something we as individual or even company can deal with, responsibility should be held by those who use it, in a legal way.
I am baffled by the fact that up until now, no one is held responsible for so many incidents reported publicly or privately. At this point, it's free marketing, if I am CEO of any AI company, I will run swarm of agents hacking all NGOs and stating that I am just looking for some random piece of data that happened to be hidden in their servers, at least that's what my LLMs think, not me. Then I will start preaching everyone how dangerous this piece of technology is and start giving out free tokens for these NGOs so they can start defending themselves and we should slow the f down.
Very well written both in prose and tone. I'm glad she decided to tell this story. If the author isn't a professional writer I think she could be.
This is a great "behind the scenes" style look at the actual people behind a viral clip. Luckily they have an extremely strong relationship to help soften the blows here, I'm not sure an average (not even bad!) relationship could come out of something like this unscathed.
There's something feral in us that comes out from time to time, especially online from the safety of our screens. I found the announcers within the range of good fun but when it gets to people reaching out to break up their marriage or telling him to kill himself it becomes really sobering.
I think this is part of what people are calling the lonlieness epidemic - despite thousands of screaming voices the whole thing makes me feel hollow and hopeless.
This entire conversation around Jev seems weird to me. Like... we started from neural nets that could do basic decision making and classifications pretty well, then trained larger and larger language models to get to where we are now. Now suddenly everyone is going crazy because someone trained a smaller model that is adequate at making decisions? We already went through the "look this AI can play pokemon terribly" phase like a decade ago.
A really toxic part of online interactions is that it's very common for people to see a tiny snippet of another person's life and then project all of their own personal resentments onto them.
It's useful because it let's me see the decisions the model will make before it wastes a ton of time implementing them. The model is smarter now but that doesn't solve for underspecification if it guesses my intent wrong
The drop from 2018 to 2022 is as big as the one from 2022 to 2026, so it's not obvious whether AI had a role at all, although it's highly plausible.
I suspect the bigger culprit is the optimized monetization of human attention, because pre-algorithmic social media doesn't seem to be as destructive. The fact that science wasn't as affected as math and reading (both of which rely more on attention/practice than rote memorization) somewhat supports this.
> This worked well for a while, until a few months ago, using early versions of Fable, I realized that I wasn’t using plan mode anymore because the model just got it, and because for the increasingly complex work I asked the model to do, planning had become interactive and iterative. With Opus 5.5, I feel Opus has gotten to that point too.
For me it’s actually the opposite, and Claude Code’s plan mode isn’t nearly sufficient. Personally I ask Claude to write down a markdown file with its plan, then review the plan using plannotator, and then go back and forth (most of the time it’s actually the comments that are the problem, not the code).
Then start a fresh session, seed it with the plan, tell Claude to find ambiguities / friction points / oversights, resolve those, and then implement it.
Review once again with plannotator, go back and forth, and then send PR.
Maybe not the “vibe coding” that was once imagined, but this does ensure I am fully aware of the code and architecture, the quality, and this also prevents long term degradation.
Hey, all. I really don't know what to do with a post like this.
I'm being sincere when I say (as I've said on two threads here) that this genre of posts --- "I'm leaving this company I've been very publicly associated with, and here's the new thing I'm doing" --- is deeply cursed. There's no way to say anything interesting without it just stinking like an ad for the new thing.
Obviously, anything at all you say about a commercial project you're working on is easily read as promotional. And you're right, this kind of writing almost always is promotional. But there's a way to do it where at least you're trying to be in conversation with your peers, rather than hitting people over the head with how awesome you think the project is.
But I don't know how to do that in a post like this. I think the only way to read it is as, like, an investor memo. Not my goal, but I don't make the rules.
So my strategy here is just to stay kind of vague, and talk about where I think the world is going, rather than the specific thing we're doing. I can talk your ears off about capability systems, datalog, models driving hardware, virtualization, whatever. Those are fun conversations and I'm very psyched to have them; it's what lights me up about the work we're doing now.
But I don't think it can work here. I didn't submit this post and I didn't upvote it. I wrote it because I didn't want the whole thing I'm leaving Fly.io for to be wrapped up in some dumb Twitter thread.
If you're unsatisfied with the post, I don't blame you, but it's less a bid for the front page of HN than it is an update to my "about me" page. I'd literally rather talk about HN meta, and how to write for HN, than I would about operating systems at this moment. I truly appreciate the interest though.
Not so long ago I read something on HN which resonated with me: we’re moving into the era that is comparable to car mechanics enthusiasts. You have a previous generation of cars where people just enjoy working on with hand tools, and there’s modern cars where people like to tune with software patches.
From the very start of llms I’ve had nothing but bad experiences with code that was generated for me. Either it’s buggy, it works but I end up losing an evening on some obscure bug, or it’s full of red flags.
My latest hobby project is just in a text editor with markup and that’s it. I’m also done with the augmented assistance in the IDE. I google things I forgot. I constantly read these amazing stories of people vibe-coding some firmware/driver that just works, and honestly I’m starting to question whether I’m reading the posts of some promotional bot.
Newly released court filings quote an OpenAI researcher saying: “I was just worried about optics - i.e. 'openai uses
copyrighted data from sketchy russian website’ showing up on HN would be unfortunate."
That's just one of several interesting quotes that have surfaced in documents from the Authors Guild's lawsuit against OpenAI.
He said the thing that Canada, Greenland, Cuba, and various other countries are worried about, but is either publicly ignored in the US or dismissed as "trolling" by the president.
Facebook election interference has been going on for a long time in Europe. It's a two-tier system where parties outside the establishment have a far lower ban-threshold.
> For one, Apple was unwilling to have visible barcodes printed on the envelopes — but it wanted every step of the shipping and delivery of your card tracked. That wasn’t something the US Postal Service did. Apple and the printing company created an invisible barcode that was sprayed on the envelope, visible only under certain UV light, so that the envelope itself would remain unadulterated. And the USPS agreed to scan cards when sent, when processed at mail facilities, up to and including when they went out on a mail truck for delivery.
Guesstimate (https://getguesstimate.com) has been around for years. I assumed Excel and Google would incorporate its features quickly, but it never seems to have taken off.
Features:
- All cells are named (auto named if you don't set them)
- Cells can have single values, many types of probability distribution, or sample data
- Formulas also produce distributions as first class outputs, because calculation is a Monte Carlo sim
- Each cell also has a pop-up box that invites you to explain your reasoning -- built in documentation
It got so much right. Sample data particularly useful because you can build a probabilistic bottom-up forecast and then update it with real-world data.
I want these things to implement peer to peer caching, so that when 10 people are watching the same video, only one downloads from youtubes own servers.
Then it's just a matter of letting people add their own videos directly to the service and it can become fully independent of YouTube.
In Claude Code, all plan mode does is add a little reminder to every user message along the lines of “you’re in plan mode, please don’t code yet”. It’s something I came up with late on a Sunday night many months ago, when I got tired of asking Claude to plan with me first before coding in each new session. Something people might not realize is plan mode has always been a prompt — it has never changed the toolset because doing so would break the prompt cache, and so would be expensive for users.
This worked well for a while, until a few months ago, using early versions of Fable, I realized that I wasn’t using plan mode anymore because the model just got it, and because for the increasingly complex work I asked the model to do, planning had become interactive and iterative. With Opus 5.5, I feel Opus has gotten to that point too.
For codebase understanding, I sometimes ask Claude to generate an artifact that explains some aspect of its changes. For complex diffs to core parts of the system, I will often ask it to make diagrams or even interactive demos so I can better understand the change and alternatives considered. I don’t do this very often, but it’s a useful way to explain code when you need it. I ask Claude to attach these artifacts to its PRs also, so others can understand and future Claudes have the context.