I have found including this in my AGENTS.md to be quite transformative in this regard:
> Always use an aggressive red/green TDD-approach. It is critical to remember that in the red phase, things like module/exports import failures due to trying to import file paths that don't yet exist, exports that don't yet exist, etc. is not valid TDD. For valid TDD, the test cases must actually run. For this, you must create stubs of the expected modules and exports in the red phase, so that the test cases actually run and fail on the test case assertions themselves. In some cases, when using this approach, once in a while some of the red phase test cases might "incidentally" pass, and this is ok. Before running red phase tests you should always make predictions about the number of test cases you expect to fail/pass -- this count is not the number of test files or test suites, but rather the number of test cases. By performing these red phase expected counts of passing/failing test cases, it will help you catch errors in your prior reasoning quickly and efficiently.
> Always use a proof-driven, scientific method-based approach to validate hypotheses, assumptions, and conclusions: define the smallest falsifiable hypothesis, create or identify a reproducible failing case, gather direct evidence, make the smallest targeted change, and then re-run the same proof to confirm the issue is fixed. Avoid speculative fixes, broad rewrites, or changing multiple variables at once. When possible, preserve the reproduction as a regression test before implementing the fix. Consider that when gathering evidence, additional logging and durable files can be very helpful.
> The strict TDD and proof-driven approaches described above could be described as "proof-driven development". Try to internalize and generalize these concepts, as they are broadly applicable.
> It is critical to remember that in the red phase, things like module/exports import failures due to trying to import file paths that don't yet exist, exports that don't yet exist, etc. is not valid TDD. For valid TDD, the test cases must actually run.
Is this for llm purposes? In many forms of TDD a compile error is key to the red phase. And you are going from red to green before asserts even exist is part of the process.
It's a bit disheartening that this is generating so little discussion (this thread seems to be the one with the most comments, currently at 8).
99% of my max plan usage is non-interactive, and this post-June 15 pricing will far, far exceed what I can afford. I assume this applies to a great deal of us.
is using a Ralph loop for all tasks with "claude -p" and using my weekly limit up to 100% considered some kind of unprofitable outlier? It's a command line tool, it would be ridiculous to expect that a large number of users don't do this. I never launch "claude" interactively by itself.
My understanding is that now with this workflow I will pay the same amount but get much less usage before getting to 100% for the week. How is that not bad?
They've probably (rightfully) identified long horizon autonomous development isn't quite there yet, so most stuff created by forcing Claude to run until you hit 100% of your weekly limits is going to be relatively low value slop that will not convert to long-term sticky spend/usage vs someone actively steering Claude until they hit their limit.
If anything you might be an outlier if you do this and are actually producing something of reasonable value down the line.
They know if the value provided by subsidized usage is too low, they're literally burning dollars for nothing: people won't feel any great pain losing it one day and definitely won't convert.
Lots and lots of automation, along with long-running multi-instance headless tasks. My usage, which I was on the verge of increasing via a second Max plan, would cost 3 figures per day, rather than per month. I already know this because my initial tooling was prototyped against direct API usage, and that was the cost of my usage patterns.
That is not sustainable for me. I understand that Anthropic can't subsidize all of us forever, but this will force me away from claude which is unfortunate. A number of the people I work with are similarly affected.
Same. Lots of little tasks that add up over time. The economics might not be great for Anthropic, but the economics of this change definitely don’t work for me; I could build a home lab and recoup its cost in about three months for the same pricing. “Headless is expensive” is a weird tone to set given so many recent cc releases pushing for greater autonomy (agents view, auto mode).
I’ve been happily single-homed on Anthropic for the last year, but I’ll be spending the next month rewriting projects off it to take advantage of better value elsewhere. I suspect that’s a one way door.
I'm curious what you're doing so often and repeatedly that can't be done with more normal software. Or maybe write a good school, I've noticed Haiku+good instructions often matches Opus
Last week (Sunday to Sunday) I had a repo running a lot of cron workflows 24/7. After like 4 or 5 days I exceeded the free limits (Pro plan) and so set up self hosted runners.
After like day 2 my workflows would take 10-15 minutes past their trigger time to show up and be queued. And switching to the self hosted runners didn't change that. Happens every time with every workflow, whether the workflow takes 10 seconds or 10 minutes.
Odd that you would omit the part of the text you quoted that contradicts the impression your partial quote creates.
> The images were initially believed to have been obtained via a breach of Apple's cloud services suite iCloud, or a security issue in the iCloud API which allowed them to make unlimited attempts at guessing victims' passwords. Apple claimed in a press release that access was gained via spear phishing attacks.
I also found it notable that the source for the above unlimited password guessing password guessing is an Apple press release that states no such thing.
Also interesting was that all sources in that article suggesting anything about unlimited attempts describe to an app or script (unclear which) called iDar, which the only source to actual name iDar claims that it reports success 100% of the time, regardless of its actual success in guessing the password.
I've no love for Apple. Maybe it's true. But the evidence presented in this wiki article is weak.
Oh sorry, you're right. Here's an article[1] about Facebook, from the same outlet. With some cursory searching I couldn't find anything about other social networks hiring former three letter agency employees.
> an enormous amount of empirical data underscores the risk to an individual's health posed by the consumption of cigarettes and, therefore, it can be argued that a parent allowing their children to smoke could potentially amount to child endangerment or abuse, depending on the context.
1. Would you agree that reasonable people exist who, right or wrong, believe that the psychological toll of SM on children, while not in the same universe as the physical toll of cigarettes, still manages to cross the line of "this is sufficiently harmful that the government must make parenting decisions"?
2. Are you open to the possibility that, in a hypothetical future with sufficient empirical data, those beliefs might be shown to be accurate?
I don't have children / a horse in this race, I'm more exploring your position about the role of government.
1. I don't care what people believe. Human history is full of people who moved to restrict the rights of others because of their beliefs. Beliefs are irrelevant to me.
2. I'm open to the possibility that data could show this, but any proposed legislation would have to be weighed against fundamental human rights. There are a lot of "risky" activities that people, including children, can enter into and we don't legislate against that because people have the right (legally and morally IMO) to commit suicide by any means of their choosing.
Point taken on the fact that beliefs are subservient to data. But society is built on at least some shared beliefs. For example, based on your responses you may have some belief that there are certain inalienable rights endowed by nature/God/whatever. That's a shared belief that isn't necessarily rooted in data. We can't just hand-wave away the idea that some beliefs are necessary for society to function. We can, however, debate which beliefs need to be shared for society to function in a particular way.
> For example, based on your responses you may have some belief that there are certain inalienable rights endowed by nature/God/whatever.
I don't want to get into a philosophical debate here, but you can absolutely demonstrate that there are inalienable rights. I'm an atheist, I don't believe in any diety, but rights derive directly from your nature as a human being. That nature being what your requirements for survival are.
As a human you have the capacity to reason and this is your primary tool for survival. That's not a belief, that's an observable/demonstrable fact. As humans we can't fly, we can't run fast, we aren't particularly strong, we don't have venom ... but we can think rationally.
We also have material requirements. We must breathe, we must sleep, we must protect our body from the elements, we have temperature requirements, we are susceptible to viruses, we require nourishment etc.
The combination of our ability to reason with our material requirements for survival mean that in order to survive as a human being you need to think and you need to be able to act. I would argue that anyone that disputes those facts, or believes otherwise, is subscribing to a belief system not grounded in reality and reason.
Your rights, as a human being, are the direct corollary of those two facts of nature. The ability to think is a requirement of your survival and that is why you have the right to freedom of expression and freedom of conscience. Your material requirements for survival are the root from which your rights to acquire and own property, to associate and to travel derive.
Within a political context, these become moral sanctions on certain actions. But the "natural rights" theorists will point out that these are all things that are fundamental requirements for your survival even in an isolated context. You can't wish or "believe" these facts away.
Natural rights do not exist outside of the belief in a social contract. That social contract is also a fundamental part of human survival. They are an extension of human reason and do not exist a priori. History is rife with examples of how natural rights is a belief that can be eroded, unlike true natural laws. Besides that, human have the capacity for reason, but this should not be confused with saying humans are rational. That means we have irrational beliefs that can contribute to survival but do not, in fact, reflect reality. That in turn means "fitness for survival" is not a rational basis for determining facts.
As someone who has been doing the same thing recently, here's how I solved the issue where the page content has to be in the initial HTML.
The first thing I did was fall back to a headless browser. Let it sit for 5 seconds to let the page render, then snatch the innerText.
But 5-10% of sites do a good job of showing you the door for being a robot.
I wanted to try and solve those cases by taking a screenshot of the page and using GPT-4 visual inputs, but when I got access I realized that 1) visual inputs aren't available yet and 2) holy crap is GPT-4 expensive.
So instead what I do is give a screenshot service the url, get back a full-page PNG, then I hand that off to GCP Cloud Vision to OCR it. The OCRed text then gets fed into GPT-3.5 like normal.
I haven't tried this myself yet. But I'm surprised you didn't find it beneficial to pass the raw HTML to the chatbot (potentially after some filtering). Did `innerText` give better results than `innerHTML`?
My intuition is that the structure information in the HTML would be useful to extract structured data.
Heh, mostly as an experiment. I'd done a fair bit of scraping for some personal football apps over the past few years. Was curious about how GPT might be used when starting from first principles, as well as its abilities to solve specific challenges encountered with the traditional approach.
When this topic has come up with family and friends, folks often say that they aren't worried (yet) because while a human can be fooled, they can't yet fool readily available forensic tools, and perhaps can't ever.
I can't speak to the veracity of that claim, but as the post points out, the past several years has shown us that it doesn't matter, not in the least.
The author goes on to say how it feels inevitable that we'll see a Bannon or Stone type use this technology to create fake scandals.
I'm more worried about the grass roots efforts. Crowdsourced conspiracies like QAnon. Now they'll have more capable tools to radicalize people.
What I found particularly striking about them was how much they reminded me of both neurons and larger brain structures, as well as some of those newer, ML-assisted FMRI imagery.
Probably just coincidence and wishful thinking, but it instills a sense of daydream-like wonder all the same.
I have found including this in my AGENTS.md to be quite transformative in this regard:
> Always use an aggressive red/green TDD-approach. It is critical to remember that in the red phase, things like module/exports import failures due to trying to import file paths that don't yet exist, exports that don't yet exist, etc. is not valid TDD. For valid TDD, the test cases must actually run. For this, you must create stubs of the expected modules and exports in the red phase, so that the test cases actually run and fail on the test case assertions themselves. In some cases, when using this approach, once in a while some of the red phase test cases might "incidentally" pass, and this is ok. Before running red phase tests you should always make predictions about the number of test cases you expect to fail/pass -- this count is not the number of test files or test suites, but rather the number of test cases. By performing these red phase expected counts of passing/failing test cases, it will help you catch errors in your prior reasoning quickly and efficiently.
> Always use a proof-driven, scientific method-based approach to validate hypotheses, assumptions, and conclusions: define the smallest falsifiable hypothesis, create or identify a reproducible failing case, gather direct evidence, make the smallest targeted change, and then re-run the same proof to confirm the issue is fixed. Avoid speculative fixes, broad rewrites, or changing multiple variables at once. When possible, preserve the reproduction as a regression test before implementing the fix. Consider that when gathering evidence, additional logging and durable files can be very helpful.
> The strict TDD and proof-driven approaches described above could be described as "proof-driven development". Try to internalize and generalize these concepts, as they are broadly applicable.