Hacker Newsnew | past | comments | ask | show | jobs | submit | super256's commentslogin

It's a "problem" of compute, I think. If you query without an account on ChatGPT you will see the model look up less stuff and research less, than when you have a paid account and choose "medium" or "high" in the effort slider.

Which makes sense, because of you have looked into search and crawlers you notice that search is actual quite expensive (which is why e.g. Kagi charges a few bucks for search every month).


It's not strictly compute, because this has noticeably improved in open-weight models too, such as Gemma and Qwen. I suspect they noticed this issue and adjusted their training to be better about it over time.

I built a toy news-summarizing agent with Gemma 4, and it was so frustrating, actually, because of the cut-off date.

The model wasted over half the token budget, each time, on internal debates over the current date.

When generating a World Cup summary, for example, it refused to believe qualification rounds were over and refused to even call the web searching tool to collect the data.

I injected the current datetime at the very beginning of the system prompt, but Gemma refused to believe it!

The m-effer insisted the timestamp was fake and hypothesized it was being evaluated in a synthetic lab test with simulated future dates!

No amount of system prompting could convince it to trust the clock.

That was the most frustrating and bizarre "bug" I ever faced!


Did you try other models (e.g. Qwen3.6 or others)? I'm curious how others fare.

I've noticed Qwen3.6 struggles a bit with today/date based logic.


Qwen 3.6 did the same thing for me.

Only after some cajolling it finally went to check the history I asked it to (I had been testing KoboldCPP's web search).


I do have Version 10.10.0 since today and it also started crashing on open. Guess I'll have to downgrade to an older version for the time being.


Maybe hidemyemail allows easily creating many accounts on a site. And some big site didn't like it.

The "Sign-up via Apple" button and creating an iCloud email yourself have a slightly higher barrier than creating a new throwaway hidemyemail email (1 API call w/o captcha/phone verification or whatever).

We might find out later this year if some site starts blocking @icloud.com but keeps allowing @private.icloud.com.


It does, and I know a friend of a friend who uses them for Walmart accounts to bot pokemon drops. Those and office aliases.


>Anthropic’s Claude Opus 4.6 system card described Cybench as “saturated,” reporting near-100% pass rates without a cheating audit. If these estimates were representative, cheating would be a marginal artifact.

One would assume that LLM creators do run the benchmarks on systems with least privileges. Which means that the LLMs don't have general internet access, can't read config files etc by design. That's why you also should run agents in a sandbox/vm (codex does this by default).


I tried looking you up but were unable to find your profile site.


https://slashdot.org/~dmd look at the top right where it says dmd (404)


It was a joke :)



Woosh


https://slashdot.org/~dmd , look at the top right where it shows user number



Given I spent hours timing my signup perfectly almost 30 years ago so I'd get #404, maybe you're the one being woosh'd.


No, they couldn't find you because you are 404. That's the joke. Wooosh


I thought you were joking! Given it's 404.


It's built on OWL, which is their own Chromium integration. It's not electron.

Seems pretty solid for Mac, but idk about Windows and Linux.

https://openai.com/index/building-chatgpt-atlas/


Crawling the internet and dumping it to disk is not "stealing".


Then distilling models and deobfuscating reasoning traces isn't.

Mass downloading copyrighted works is. Which they did. Aaron got threatened with 20 years, they got pentagon contracts.


If they were only copying, for example, New York Times articles and many publishers to a disk, I don't think NYT and the publishers would have sued OpenAI. But OpenAI isn't just copying things to disk. NYT reported ChatGPT (before Dec 2023, [0]) was returning near verbatim sections of NYT articles.

Is this stealing? Is it depriving NYT or publishers/writers from money via lost sales/subs? I don't know, but it certainly could be.

[0] https://www.nytimes.com/2023/12/27/business/media/new-york-t...


Is everything licensed in the same way? Are there any copyrighted works available to be had through crawling?


It's not stealing but arguing that it's not infringement because its on the internet is pretty obviously nonsense.


Haven't they been doing this for a long time already? I remember trying to copy Hunter S. Thompson's style many moons ago and getting a refusal.


If only GPT wouldn't refuse my requests to write a crawler for $site. :(


Gemini has been perfectly willing to write such things for me


Grok Build won't even question it. You gotta weigh that Grok use up, though.


It did for me, even without an account. Is that because I don't have an account and I'm accessing some weak model with little guardrails?


I don't know! I recently tried to download all posts by Matt Levine for a funny analysis, but GPT wouldn't let me do it.


tbh hackerone should just implement a +/- reputation points feature on researcher profiles. Like, the researcher submits a slop report to GitHub via H1, GitHub looks at it and identifies it as slop, GitHub presses the -rep button on reaearcher profile which bans them from submitting to GitHub on H1 again and makes their rep points minus 1. Companies should be able to configure you need at least 10 rep points to receive payouts. Only specific (by H1 chosen) companies can +/- rep.

So, researchers first need to collect some positive rep. But the rep points are global, so once you have fixed a few bugs for Google, you've gotten enough +rep that you can also receive stuff at GitHub.

Oh, and ID check when signing up at H1.

Long term all beg bounty submitters would be banned for pretty much all of tech.


I think they do that already, there are signal ratios


Signal and Reputation already do these things but public programs are…well public.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: