Because people want to participate in activities that cost money. It's easy to see use cases like an AI agent notifying someone that a band they like is coming to town and asking if they want to secure tickets to the show.
>The demo uses Claude Haiku rather than newer models like Opus or Fable. This is because I'm assuming that shopping agents of the future won't use newer, more expensive models,
Why are we assuming that shopping agents are going to be using the model that is easiest to fall for prompt injections? Only testing a year old, small model is going to lead to a misleading conclusion.
Also the author is deliberately chooses a model that is 1 year old and was supposed to be underpowered even then. If cost was deciding factor, they could’ve chosen Luna.
Model AA Index Cost/task
Claude 4.5 Haiku (thinking) 17 $0.21
GPT-6 Luna (max) 37 $0.07
GPT-6 Sol (low) 34 $0.13
GPT-6 Sol (medium) 40 $0.25
Luna (low) scores higher and costs about 47x less per task and performs better.
And Luna (max) scores 37 while still 30% of the cost of Haiku which scored 17.
Why deliberately choose a model that is known to be poor and also costly?
>The output of an LLM is useless without a human mind to comprehend it
This is like saying software is useless if the user doesn't read the source code of it. This what happens >99.99999% a person uses software. People want to be entertained or have their problems solved.
>the economically dominant strategy to not verify them and not double check them
In the big scheme of things is it really that expensive to verify it if a lean proof is generated? The agent itself will likely have already verified such Lean code before calling it "done".
Is it not possible that this would change the risk profile of depending on them. It makes sense to work with people who support you than oppose you for things which are critical.
My point is did everyone know all of those principals and how committed Anthropic was to them from the beginning or did new information come in and now they need to adjust. Just because something is a term in a contract that doesn't mean it's necessarily strongly held belief that will never change.
Who runs the government changes every 2 to 4 years… are you suggesting we completely swap every vendor in the US government to align with whatever political party is in office?
There’s nothing indicating Anthropic was not keeping their end of the contract. If they had, there would have been other recourse.
You don't think that there is competitive pressure between Anthropic, OpenAI, Google, the various Chinese model companies, and others, to advance the capabilities of their AIs?
Or you don't think that sufficiently advanced AI can cause (mass) harms?
I am saying that if they lost alignment and their new model started injecting cryptolockers, dropping all tables of productions databases into the software or if it stated writing poisonous cooking recipes there would be backlash, lawsuits, and more towards such a lab. Consumers don't want to trust such a dangerous model so they won't buy the tokens and the lab will not want to spend money on lawsuits.
>Or you don't think that sufficiently advanced AI can cause (mass) harms?
Even a feather can cause mass harm if it's used to sign a declaration of war. Something merely being capable of causing mass harm is not an issue and doesn't mean that the existence of feathers are an issue.
But this assumes an evil superintelligence. We don't even need that for terrible things to happen; we just need people doing people things and AI doing AI things at scale, and that scale is rapidly exploding as we scramble to deploy AI in the real world.
Case in point, that hallucination (which, if it was an evil superintelligence, could have been "strategic") that almost led to US boarding a Chinese ship over suspicions of nuclear weapons: https://www.msn.com/en-gb/news/other/us-military-ai-failure-...
It doesn't have SkyNet, it just has to be WOPR. And it doesn't have to be people who are motivated to cause harm, it just has to be people making consequential decisions who are careless, distracted, paranoid or anxious -- which everybody is at some point or another.
And I almost kill people when stopping at a cross walk. That doesn't mean I harmed a pedestrian. Society is set up to be very robust. Humans themselves make mistakes and do bad things and society has had to learn how to live with that truth.
>it just has to be WOPR
Then why argue for setting a pace for the frontier labs if we've already surpassed WOPR level integration / intelligence.
> And I almost kill people when stopping at a cross walk. That doesn't mean I harmed a pedestrian. Society is set up to be very robust. Humans themselves make mistakes and do bad things and society has had to learn how to live with that truth.
Maybe not you personally, but many, many other people have killed many, many pedestrians, and when they exhibit a pattern of bad driving -- or other deviant behavior -- we take them off the streets. That is an example of society being robust.
Except, over here, we're rushing to make AI, which we know has many deviant behaviors and we know caused harms in many circumstances (starting with AI-assisted suicides), even more powerful AND deploy it in more and more real world systems!
As history and, literally, current events show us again and again, society has failure modes that lead to widespread harm and destruction. We've had world wars and then literally had multiple close calls with nuclear war right after. And then we have all that's going on out there. (In related news, the Pentagon threw a hissy fit because Anthropic would not let them use Claude for autonomous killing machines. Do I have to even mention what they've been up to these days?)
Most of these failures are caused by misaligned incentives and socioeconomic forces. The incentives and forces around AI have hints of many brand new failure modes that we can't even foresee because things are moving so fast.
> Then why argue for setting a pace for the frontier labs if we've already surpassed WOPR level integration / intelligence.
"We're already going down this mountain pass at 200 miles an hour, why slow down now?"
I think we should not just slow down frontier AI development, we should also slow down where and how that AI gets deployed.
Which currently require interest payments of ~6-8% APR. Meaning that you need to be able to invest that money that is being borrowed back into the economy to hopefully get a return more than that. And if your investment fails you will have to realize a different investment. The interest being paid doesn't get hoarded either and is used to make other investments, pay employees, build products, etc.
The idea that a bunch of people are just hoarding their money and not reinvesting it back into the system is flawed. Taxes actually have the opposite effect to contributing to the system. Taxes are like if someone was to come and start hoarding money under their mattress for himself and not contribute back to society.
It's funny that their case study ignores how much wealth you lose from the interest on the loans. If you pay $140k of interest you can avoid $62k of taxes.
I don't know what point you were trying to get across with your link, so I gave my general thoughts on the article.
reply