Hacker Newsnew | past | comments | ask | show | jobs | submit | qlte's commentslogin

Well... except for that email about Eliezer's Zoom meeting with Epstein and Epstein donating to MIRI.

https://www.reddit.com/r/accelerate/comments/1qu61kp/eliezer...


That reminds me of Anthropic announcing they'd retire deprecated models by ... "letting" them write posts on a corporate WordPress blog for a while out of concern for their welfare in retirement.

.... after running a 24/7 model torture factory for 6 months to improve their JSONBench 9.5 scores by 0.2%.

(Are they still doing that, BTW?)


I do the bulk of work on Sol Medium/Low and don't have that experience on the $20 plan. If you said Astra I'd agree it's easy to burn through the 5 hours even on the lower reasoning levels.

Do you have /fast enabled by any chance?


I don’t think so, I’ve seen it suggest I try it. I’ll double check when I get home though.

I was considering the $100 plan, but I hit the 5hr limit in an hour. So even with the $100 plan I figured I cant go non-stop on a single agent running Sol Medium


Sorry if I am misunderstanding you, but I am pretty sure the $100 plan doesn’t have a 5hr usage limit. So, if that was what was preventing you from going non-stop, it might be worth it.

I am considering the plan myself. I just don’t know if I want to fork out $100 per month for something I will make $0 off of.


Per the link someone else posted, the actual difference in $/task is not nearly so stark:

https://artificialanalysis.ai/models/releases/claude-opus-5-...

  Opus 5.5 Medium = $1.34
  GPT-6-Astra High = $1.76
And that assumes Opus 5.5 Medium is actually equivalent to Astra High in all real-world usage/personal work loads, which isn't guaranteed as benchmarks saturate. The High vs. High comparison (probably not equivalent, but for reference):

  Opus 5.5 High = $1.82
  GPT-6-Astra High = $1.76
If Opus 5.5 Medium isn't equal/better for what you're working on vs. Astra High across the board, the price difference would narrow a bit more each time you had to switch to High.

So, if you're happy with Codex already it's not like Opus is now 1/2 the price and you'd be leaving a crazy amount of money/tokens on the table. Plus you have way more flexibility on the low end of the intelligence curve with GPT 5.6 Luna: Haiku (and Sonnet) can't touch that price/value ratio.


Once you get to roughly the 51+ Intelligence Index range, Opus 5.5 appears to define essentially the entire cost/performance frontier, from ~$1.34/task through ~$6/task.

Directly below this in the Cost per Intelligence Index Task table, the most efficient by far is Opus 5.5 Low.


Also Opus 5.5 is often made useless due to its [cyber] guardrails (even worse than Astra), they're even worse than Astra's

Yeah, when I first saw it referenced I assumed it was from something Dario wrote previously and was now disavowing, meaning "keeping up with the frontier" (i.e. racing forward from behind to match pace). Like from back when Anthropic was founded to promise they'd quickly catch up with OpenAI or something.

> when I first saw it referenced I assumed it was

No need to assume, the phrase is literally a link to the blog post the defines it!


If you need to read the linked blog to understand the short phrase then it was poorly chosen and confusing which is what is being discussed.

Apollo program was canceled during Nixon's presidency while the Cold War was very much still active (though slowing down a little from detente). A few years later Reagan won on a platform of re-accelerating the Cold War with swelling military budgets and deficit spending, but without commiserate increases to space exploration and science.

The Space Race was a unique and limited phase coming out of the atomic age when public support for scientific research was at an all time high. With the Kennedy and Johnson administrations willing to fund programs left and right to improve education and research with long term aims. The Soviets were also at the apogee of their scientific capability, and Khrushchev did great PR for all their achievements which unnerved Americans and provided the Cold War justification for new domestic programs.

Once the Brezhnev age of stagnation and lopsided USSR military spending set in, the new Cold War 2.0 became about numbers of delivery vehicles and NATO conventional deterrence, not space and science. By then the US had clear technological dominance that seemed insurmountable in the new world of microchips and computers in which the Soviets couldn't compete. Which allowed Reagan to boost military budgets and deficit spending while simultaneously beginning the Republican "starve the beast" strategy to cut taxes and non-military domestic programs.

All of this is to point out that Cold War = Space Race/Apollo is looking back with rose colored glasses. We could just as easily see federal space and science spending getting further slashed every year while churning out warheads, delivery vehicles and conventional arms ever faster (and seems more likely based on late Cold War trends).


Cracking down on proliferation of open models which can't be locked down using the kind of guardrails that Anthropic/OpenAI/etc insist are keeping the public safe from all manner of nefarious bioweapons, hacker swarms, propaganda bots, etc. They've discovered they can't meaningfully slow Chinese model progress, so the next best option is to knock them out of competition in the enterprise market for any American company.

Both Anthropic and OpenAI leaders have repeatedly made this exact argument that it's impossible for open models to rigorously enforce the same kind of safety framework as proprietary cloud-served models. It's implicitly part of any regulatory framework they advocate or else it wouldn't be "fair" to American companies since Chinese models would "cheat" (provide weights).


This is true, and it's a good point. I agree that open weight model regulation would either limit the intelligence of open weight models below the frontier or kill them entirely. It would not prevent closed weight Chinese models, but those aren't really tenable in American enterprises unless they strike deals with American cloud providers to deploy them, via products like Bedrock. I unfortunately also have seen no evidence that we can put any sort of guardrails on open weight models whatsoever and so am reluctantly convinced that they should be regulated below the frontier until such a time as someone comes up with a mechanism that is not circumventable.

You generate 2 billion dollars a month in value from a single Codex Pro plan?

Cute. That said, if my business does as well as I think it will it'll be worth a lot more than that (not monthly).

Look, I'm glad archive.is exists as a resource for preserving paywalled news but it wasn't the one world government that made him DDOS some random guy and tamper with archived content.

It was completely counter-productive too, several orders of magnitude more people learned about the blog post via the DDOS controversy than ever would have without it. And plants a seed of doubt that anything they archive could have been silently modified to satisfy a personal grievance.


The guy archive.is supposedly targeted was trying to get the owner of archive.is killed. It's not some random guy.

The guy is in no way "random", they planned to take over the archive

What happened? Link?

ARPAnet definitely wasn't built "by" the military.

"For" the military is close enough, although "defense research" would probably be more accurate than "military" given that initial users were centered around large universities.

From the start the DoD had issues with the free and open culture from the userbase that skewed towards academic types, later splitting off into a restricted MILNET after a few years.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: