Hacker Newsnew | past | comments | ask | show | jobs | submit | Catloafdev's commentslogin

Buddy you aren't addressing the comments that are replying to you.

You gotta stay on-topic. You keep coming up with random things to compare to that make no sense to anybody but you. Who said it was legal to say "kill the king"?


> Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5.

Sounds like they noticed the complaints. I'm curious to see what LLM-isms this one may have.


I don't think it's substantially different. I just pasted a random chunk of code and asked Opus 5.5 to comment on it:

> The Vercel target is hard-coded. That's common and not wrong, but it's opaque; nobody reading this later will know which Vercel project it belongs to, and if the project is recreated the target changes silently. A comment or a named variable would help.

> Pointing a DNS name at Vercel is only half the job. The domain also has to be added to the project in Vercel's dashboard, otherwise requests will arrive and Vercel will reject them. That step lives outside this code, so it's easy to forget.

> Finally, [CENSORED] existing only in production is slightly odd on the face of it. It may be perfectly deliberate (perhaps a single shared testing tool that only needs one public address), but if you're reviewing this rather than just reading it, that's worth confirming.

It has the same annoying cadence and writing style with slightly less prominent claudisms.


“It has the same annoying cadence and writing style with slightly less prominent claudisms.”

Seems like it based on my first session. It still does the whole “bury the important thing in a pile of words” coupled with the “it might actually be important” thing… so basically you never really know what it’s talking about.

Honestly I trust opus so little that the entire “opus” brand is completely tarnished. Its writing style is so god awful that it needs more than just a point release. Either dump the name and ship a different model entirely or at minimum call it “opus 6”. Calling it 5.5 makes it sound like it’s basically a continuation of the same garbage output that 5.1 had but with some minor adjustments. And based on my single first test, that is what it appears like to me.


Maybe if we had a single human we talk to 24/7 at scale, we would get annoyed at his cadence and style. You need variety to not pick up on known patterns I assume, which a single model can’t replicate?

No, it's just poor writing. Actionable points are buried inside the paragraphs and over-hedged, and one point is completely made up. Compare to a five second rewrite:

* Consider leaving a comment about the hard-coded Vercel target. It's not clear where does it come from.

* [This is just a bullshit point, because the domain is not "added to" Vercel, it's provided by Vercel]

* Are you sure that [CENSORED] is prod-only? The name suggests otherwise. [also, what "if you're reviewing this rather than just reading it" even means?]


> also, what "if you're reviewing this rather than just reading it" even means?

It means "I'm treating you as lay-person punter, not a developer working on this project." Opus 5 feels like it's constantly trying to reward-hack me into treating it as intellectually honest and epistemically humble, while in the same breath it talks down to me and tries to smuggle its own bullshit assumptions and assertions into the conversation unchallenged. No progress on this front apparently. Glad I cancelled.


Good points, but then even this little snippet is internally inconsistent. If I'm a lay-person, why should I care that "a comment would help"?

Claude is just comically bad nowadays.


Nah it's definitely a Claude thing. Other models even though they have their style are less annoying and less stereotypical.

It was difficult to not notice them. Opus 5 was unusable, most of my team went back to Opus 4.6 for most of their work. I hope we can move forward now.

It's unbearable but nothing that couldn't be fixed with postprocess.

How? Explicit instructions, memories and even skills have not been able to keep Claude from saying "genuinely" every two sentences and keep it from explaining heavily what something _isn't_.

"please repeat, ELI5 without analogies (I'm not a child, just ADHD)"

works quite well


Opus 5 was just incoherent - curious to see what improvements they have made here. Would love to see some kind of postmortem to better understand how writing styles change from model to model.

I wouldn’t be surprised if Opus 5 was trained on content written by other LLMs


I'm genuinely confused what's the relationship between LLMs improvements and them being so incoherent.

and it's not about the verboseness (even though it obviously contributes to the fatigue and loss of focus), I swear the vocabulary of the llms change working on the same task on the same codebase significantly.

I wonder if there are studies around this.


Remember when OpenAI models loved talking about goblins and whatnot due to the RL?

https://openai.com/index/where-the-goblins-came-from/

Small quirks can quickly add up in posttraining if not caught. Although TBH with how obvious Claude language is, I do feel like this is something Anthropic probably noticed and just assumed people would not care about. Now that people have obviously cared, they're probably actively looking to alleviate it


Can it be that now they are getting optimized against benchmarks that are valuing logics, rather than human appreciation? (I am not an expert at all, just an idea)

It's the reinforcement learning rather than supervised learning.

It is the switch from RLHF to RLVR. It benchmaxes better, but benchmarks don't cover human usability.

Maybe it's the time period we're in, maybe I'm just grumpy, but it bugs me that they release a new model every single week and the new one is just a fine-tuned version of the "old" one. If 5.5 performs similar to Fable and really does cost 40% less, then 5.5 really should've just been Opus 5. And they're essentially admitting that they are shipping slop.

You were right to notice the complaints. One decision remains, and it is yours, genuinely.

I appreciate humor here, but there are now a dozen of these comments on every thread about Claude. They no longer adding anything substantial and dilute the discussion.

I don't mean to pick on this comment in particular. The majority of my work day is now spent reading AI generated text, and I look at HN (too much!) because I want to read human commentary. Humans pretending to be obnoxious AI on repeat is net negative to say the least.


I agree. Hopefully Anthropic has fixed Opus' ridiculous communication style so that people - like me - no longer have any kind of weird impulse to imitate it.

It's a load bearing joke that was funny the first time but we're going to beat that dead horse until it starts getting funny again. If you beat it enough, it will get funny. Beatings will continue until morale improves.

Qwen3.8 Flash Next just released which hits that range.

Also, Deepseek V4 Flash can be run relatively well in hybrid 2-bit quantization on 128gb devices, with way better results than you'd expect for a typical 2-bit quant.

Those are currently the 'smartest' options for that memory level.


>until China solves the ASML problem

Which they are in progress on: https://www.reuters.com/world/china/how-china-built-its-manh...


Yes, but we don't know how well they work or what nm can they print or how machines they can make.

They've overcome every hurdle to go from nothing to having working machines, so all these questions are presumably a matter of how fast they will improve and ramp up production (initially 5 this year, 20 next year), not whether they will.

China have been really squeezing all they can out of DUV machines, but I'm not sure how much node size really matters for AI competition - more of a cost issue (more chips/power for same FLOPs) than anything, and TSMC & NVIDIA's healthy profit margins need to be considered too.


I'm well aware of how good China is at manufacturing. However, DUV/EUV machine manufacturing is unproven.

What is the timeline like? 1 year? 5 years? 10 years?


They already have working machines, and are slated to ship the first 5 early production ones right about now (before the end of 2026), and are projecting 20 for next year.

Small numbers perhaps, but this is happening right now.

What will the numbers be in 5-10 years time? Who knows, but ASML took about 5 years to go from 20/yr to 100+/yr.

Of course politically the world may well be quite different in 5 years time, as may be the AI market.


Yes, 5 years seems like an eternity in AI world.

We also don't know how small of nm chips they can manufacture. If it's 20nm, it's practically useless for advanced AI chips.


Intel shares are still up over 100% since the Trump admin invested in them and started going to bat for them.

I don't think your information is entirely accurate.


Such a great move by the government. Imagine what we could do for US industry if the government took a stake in all major corporations? Many people are saying that the healthcare industry should be next.

Not being able to have unrestricted access to the world's second biggest market does not relate to Intel being up over 100%.

Just logic.


Not making any GPUs worth buying makes you immune to export restrictions.

So we're basing winners and losers on vibes rather than money or performance?

A simple ChatGPT can explain it to you.

Intel being up 100% has nothing to do with China market being restricted. They could be up 200% instead if the China market is free.


>Intel shares are still up over 100% since the Trump admin invested in them and started going to bat for them.

Nokia shares were also up after the iPhone got announced. Current share prices don't really mean much in rational future outcomes, just FOMO, vibes and speculation of clueless masses of investors. They can also swing wildly up/down.


Especially since the market is hobbled and corrupt by now, owned by a handful of machines that humans barely have root control over anymore.

>the only reason ... is that current AI inference is unnecessarily RAM intensive

There's no indication that this will change, and every indication this will continue to grow. This is not some temporary thing. Demand is already exponentially higher than what is possible to produce, and there's no current reason to believe it has any ceiling.


I've actually done some work and was able to get the full 12GB Z-image Turbo to run on my 8Gb 3070, and throughout the generation only 600MB of VRAM was allocated (you can get a speed-up by double buffering, but the maximum VRAM was still only ~1GB through inference). It's very experimental and pretty much requires you to write all the compute kernels directly specifically against that particular weight and statically allocate the memory at compile time instead of using current ML frameworks like Pytorch/JAX, but I think in time somebody else would figure it out.

The hardware savings at the scale Anthropic or Google use would me immense, it makes me wonder why no big player has done more optimization already. When DeepSeek showed how inneficient were the models of its time, I'd expected each to create a permanent optimization team with all the talent they have hired.

I guess hardware is not that expensive to them in the grand scheme of things, at least not at this stage. OTOH, their propietary models might be thoughly optimized and we can't know, because they're still bound by supply contracts to buy the same amount of hardware nevertheless.


I fear that, whenever consumer-facing supply finally starts to return, the opportunity cost will be priced in. Consumer RAM may end up being nearly as pricey as HBM for a while, and I don't see any reason why it would ever go back down on it's own, unless there are entirely new fab processes that don't meaningfully overlap HBM.

It's not strictly compute, because this has noticeably improved in open-weight models too, such as Gemma and Qwen. I suspect they noticed this issue and adjusted their training to be better about it over time.

I built a toy news-summarizing agent with Gemma 4, and it was so frustrating, actually, because of the cut-off date.

The model wasted over half the token budget, each time, on internal debates over the current date.

When generating a World Cup summary, for example, it refused to believe qualification rounds were over and refused to even call the web searching tool to collect the data.

I injected the current datetime at the very beginning of the system prompt, but Gemma refused to believe it!

The m-effer insisted the timestamp was fake and hypothesized it was being evaluated in a synthetic lab test with simulated future dates!

No amount of system prompting could convince it to trust the clock.

That was the most frustrating and bizarre "bug" I ever faced!


Did you try other models (e.g. Qwen3.6 or others)? I'm curious how others fare.

I've noticed Qwen3.6 struggles a bit with today/date based logic.


Qwen 3.6 did the same thing for me.

Only after some cajolling it finally went to check the history I asked it to (I had been testing KoboldCPP's web search).


No, it doesn't. Nobody would go do this just to report it to Sony for a few pennies. That is a non-existent motivation for people in the know, who understand the opportunity cost.

Unless they didn't want to actually make money from anything nefarious, and the bounty was more than just "a few pennies"

This seems extraordinarily sketchy.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: