Hacker Newsnew | past | comments | ask | show | jobs | submit | jermaustin1's commentslogin

A bespoke design like that a year ago would have cost $200-2000+ between design and development. Hell, it probably still does unless the person wanting the design is already a developer who knows how to prompt.

Everything costs more now than it did a year ago... except for THIS, and we are still complaining that a 90-99% reduction in cost is STILL too expensive. And a 50-75% reduction in time is STILL too long.

We used to have to wait for weeks for a design like that when I worked at a consultancy, and that is a week of salary. For the design, then it got handed off to a front end developer to slice it and get built so the back end developer can hook it up to a CRM. We are talking a month turn around with design, revisions, development, testing, and bug fixing.

It can now be done in a couple of hours for less than a single hour's cost. If it were 10x slower and 10x more expensive, it would STILL be "good deal".


I worked for a different Steve Ross (Dolphins owner, Related Companies, etc), and we did this in house.

From 2008-2017, “we” (I was a contractor 08-09, and employee 15-17) basically turned every spreadsheet into a web application internally.

We added authentication and authorization, used a LOT of ETL-type processing to move data around. So many things could have been packaged and sold, but that was the “secret sauce” that kept the company so profitable.

Then the end in 2017 when a new CIO came in, killed off IT (laid off all non-managers over the course of a year), and replaced everyone with South African consultants to turn IT from a cost center to a profit center.

He lasted another couple years then left.


Years ago, I lived in NYC, and my roommate was a director of photography for National Geographic, and various other nature documentaries. I loved photography (still do, but much less time for it as a late 30s adult than a mid 20s adult), and she was kind enough to answer any question I had regarding film/photo.

She told me that "left to right" denoted progression in the story, "right to left" told the viewer the subject was "exiting" the current scene.

She didn't go into the details of WHY, and I probably didn't probe deeper, but it stuck with me, and I notice it all the time in film and television.


As a counterpoint, I did a quick image search on Kagi with "person on a bicycle", and I ended up with 12-left-pointed bicycles, only 3-right-pointed bicycles, 3 facing the camera, and 1 facing away from the camera — looking at photos above the fold (first screen). Even looking below, the pattern seems to continue, though not as prominently (I'd say 3:2 in favour of left-pointing bikes).

Obviously, not scientific.

All the left-pointing bicycles did not look weird to me either.


Opposite policy at one of my clients (kind of). I am responsible for the code that upper management's Claude produces. Some Mondays, I will start work with a half dozen emails with attachments of Claude generated code for something I don't even know what the point is, with the task of "integrate this and make sure it works." without any context to go along with it, so I have to read the code, usually hundreds of lines and understand WHY manager wanted it, before I can start to code it myself, because it is 1) in the wrong language, 2) doesn't understand our codebase, 3) is using libraries we can't license, etc.

My job has been less watching Claude Code, and more watching Managers Claude Code.

I don't know which I hate more as a programmer.


That’s just stupid.

Wow! I'm utterly speechless.

I love space "browsers" - I lose myself in them, just clicking through all the entities.

Also TIL that there are satellite galaxies, and I've never noticed those on other toys/tools like this. I always though Andromeda was the closest galaxy, but it is FAR away compared to `Sgr dSph` which is within 50kly from the center of the Milky Way.


Avoid playing No Mans Sky (but not really, it's an extremely relaxing game for that reason) at all costs then.

Same here! I can't believe I never knew of the existence of satellite galaxies until now.

To me, most local models work just fine for anything you can be patient for. If I want something quicker, I will go to a SOTA model via API, but with multiple 3090s, I have never really needed a hosted model for a lot of my experiments.

For code, they are great, but for creativity for NPC controllers, they leave something to be desired, but work well enough for testing, so I don't burn tokens until I'm actually playing my games.

But nothing one-shots a prototype better than Fable 5. I can have a prototype built in 30 minutes, hooked up to my local LLMs and Claude Code is very good at testing the interactions and even tuning the prompts of the NPCs for better experiences.


"with multiple 3090s" is quite a bit of burying the lede for "most local models work just fine", don't you think?


Having multiple 6 year old cards doesn't seem like it's that big of burden for local LLMs.

I get that a lot of people don't have them. And a single one can be VERY performant. And the smaller models like a 7B can run on much smaller hardware like a mid-range [3|4|5]060.

My entire AI Dev Box cost $4500 in parts. 128GB RAM, i7-10700, 1TB and 2TB SSD, and 2x 3090s. Today's prices and inflation have definitely made that price tag seem a lot better than it was, but it was an investment in all things GPU that were happening in 2020 (crypto, blender, image gen), then LLMs exploded.


A single, used 3090 costs more than I have ever spent on a computer.


Yes, the tunnel vision around local models on this site is crazy. The percentage of people in the world who can afford the hardware is extremely low.


It seems roughly similar to the pricing level of personal computers in the early eighties (i.e. IBM PC and Apple Macintosh). I’d expect prices to come down significantly over the next few years. Not so much in the next year or two, but after that.


Just like for warships, the complexity and cost of building cutting edge hardware has grown exponentially up to a point where a significant chunk of the world's computing is dependent on 2 companies: ASML, TSMC. We shouldn't extrapolate linearly from examples from the 80s.


No, but I wouldn’t expect it to stagnate like with Intel in the 2010s either. Maybe the biggest caveat is that most people will be fine with using cloud providers, so the market for non-server hardware won’t be subject to as much competition.


That's in nominal dollars. However, inflation since then has been about a factor of ten to fifty and it hasn't trickled down at all.


There's a non-small contingent who lucked into the periodic games machine upgrade at the right time to snag a {3,4,5}090 rig just before everything exploded. It's a small contingent now but it was less so then. And now those people can add a second card for roughly what that whole system would have cost new originally.


The current supply chain problems will eventually pass.


So will the stock of 3090s, as well as their ability to run contemporary local models. And there will be no supply of newer equivalents of those GPUs, because NVIDIA has since wisened up, and is using the very capabilities you need for local LLMs as market segment differentiator.


I don't think there is tunnel vision. I'm just saying that I have a couple 3090s I invested in a handful of years ago, and they are still going strong today as multiple GPU-needing technologies emerged.

I'm not saying everyone has to run local LLMs, because the APIs are in a race to the bottom, and my $10 of OpenRouter credits I bought months ago is down to $8.94 because most models give you MILLIONS of tokens for a US Quarter.


> I'm just saying that I have a couple 3090s

This is tunnel vision. The percentage of people who could afford the hardware you could at the time you back it so vanishingly small. I do not know a single non-tech person who has multiple graphics cards in a single computer.


And right now the demand for GPU is far outpacing the supply, even with factories at full production, which is keeping prices high and out of reach of most people. But unless something happens to shut down the factories (not impossible, but hasn't happened yet), eventually production will catch up to demand and prices will return to sane-ish levels. Won't happen this year, almost certainly not next year... but I would be shocked if the current high prices were to persist for a decade. Eventually the percentage of people who can afford that hardware will grow to be a decent chunk of the computer-owning population. And they'll be following READMEs written by the early adopters, for installing open-source harnesses to work with open-weight models.

My personal expectation is closer to 5 years than 10, which is why I wouldn't touch Anthropic or OpenAI stock with a ten-foot pole, personally, no matter how high their theoretical valuation is. Because their business model is doomed in the long run.


Newer cards aimed at consumer market are not capable of being used for local models the way 3090s are. That's on purpose: this capability is now used to price-differentiate between "normies playing games" and "companies in data center business".


For now. That won't last forever. Yes, it'll take quite some time to work through the current production backlog, which is why I'm predicting five years, not one or two. But the trajectory has always been "new video card comes out, game devs push the limits of what it can do, gamers buy new card so the hot new game can run faster, rinse and repeat". And that includes wanting more VRAM so the game can load more of the scene at once, load higher-res textures, etc.

Which means it's inevitable that eventually, even the consumer game market will be buying GPUs with 32 or 64 GB of RAM. And there are decent models that will run at that size. Even the "normies playing games" market, as you call it, will end up with the capacity to run local models. It'll take a few more years than it would have if the data-center companies weren't trying to buy up all the GPUs, but it's not like gamers are going to stop wanting to play games. So in the long run, Anthropic et al are still going to have to figure out how to deal with competition from local models that run on your gaming video card. Which won't ever be at parity with the models that take terabytes of VRAM to run, but are very rapidly approaching "good enough for what most people want to do".


It's a decreasing pool as well I'd say, the dev machine I built last summer would make less financial sense to me now for example.


I mean, I'm not rushing out to buy that kind of hardware myself, but it is a matter of perspective. People commonly spend an order of magnitude more on a car, and that's just the sticker price.


I keep coming back to this: why do I need to run a local model on my own GPU? Open models can run in dedicated clouds and while, yeah, they may be more expensive per token than my own GPU, when accounting for depreciation, energy usage, and opportunity cost (money not spent on my GPU will instead sit in my portfolio appreciating with its particular blend of returns), I'm pretty sure I break even or even net lose money with a GPU.

Don't get me wrong, there are advantages to a fully local model in that, I can have agents looping 24/7 even when my internet is not working. But this is niche enough that if I had to price the advantages they don't seem worth it.

If I'm willing to pay the Openrouter tax, I can fire up Openrouter today and just get access to whatever model I want, and still pay a fraction for tokens as what I'm paying with the big guys.


My single 3080 runs so hot I don't need to warm my room in winter, and have to play games in my underwear in summer.


Unless you value privacy, pay for openrouter. You still get the benefits of cheap tokens and programmatic usage.

3090 pricing is something of a wild card. Since the only big-mem consume cards are the xx90s, and a 5090 is pushing $5000, resale value has gone way up. The bottom hit ~$700 last year. It's still a very good GPU, if power hungry.


I got a great deal on ~72 TB of NVMe right before storage prices shot up, doesn't make it any less ridiculous that I have it or any more relevant to people talking about building a NAS now. 99% of people, even in tech, do not have the stupid amounts of hardware people like us hobby on.


Most people in the US have a car, and the average new car is $40,000. Hell where I live a middle class consumer will spend double that on a Boat or an RV and think nothing of it. These aren’t elite tech workers.

It’s not unfathomable that if a personal, generally intelligent local AI provides enough utility and doesn’t require you to tweak CLI flags millions of Americans would want one.


There are many payday loan operators and those willing to sell predatory loans to those workers you mention buying boats or RVs. I've yet to see a payday loan open up in SF to help tech workers buy hardware.


Most people in the US don't drive a new car, and used cars can be had for far less than $40k. An $80k purchase would be just shy of the median annual household income -- anyone who thinks nothing of that has financial resources far above typical. You are in a bubble.


An $80k purchase is far more affordable when you're looking at an 84 month loan. You trade in your current $20k truck with $30k in debt on it for your $80,000 car, get a couple grand in incentives and a $10k down payment, and boom you're only looking at a bit under $1,200/mo in payments. The median household is bringing home ~$84k before taxes, hypothetical person lives in a no income tax state, they take home ~$5k/mo. Easy peasy, its not like you were planning on taking any vacations anyway since you're always working.

What matters is you've got the Duramax HD King Ranch TRD Big-Boy machine. Doesn't matter the cost. You can tow anything, drive anywhere, do anything, and do it all in comfort. Other than parking in a normal parking spot comfortably. Or even park it in your own garage at home.

I've seen this exact scenario many times personally.


Most people in the US can't feasibly hold a job, get groceries, or go to the doctor without having a car.


Spending that kind of moment on a product that gives you personal happiness for years up to decades and then will still have residual worth, which people save up for ages for, is an entirely different proposition than buying a product that may make you faster professionally, but which in the short time can also be achieved by a few dollars worth of subscriptions to a hosted model for even greater effect.


> people save up for ages for

Americans by and large don't do that. Much of the population engages in discretionary spending with debt instruments. Combined with mass innumeracy, they're all oblivious to the true cost of their purchases because they only think of the monthly payment.


where? i would love that.


"Where'd I buy it" or "where is it now" ;)?

It was a 96 core gen 4 epyc+supermicro board build with consumer NVMe drives on 1x16->4x4 "dumb" bifurcation cards. I had to get a few MCIO-> PCIe adapters as well to get the full lane coverage. Mounted in a standard EATX compatible consumer case with a consumer PSU and a lot of Noctua fans - surprisingly cool and quiet for what it is.

Motherboard+CPU I got from Ebay. Rest from the best MicroCenter/Amazon/Walmart deal of that day. Bought juuuuust before the AI pricing apocalypse, largely by pure chance.


> 2x 3090s

You could sell those and have enough money to pay for hosted inference for years.


They cost more to run than hosted anyway. But that isn't the point of having them. They are a playground, a backup when the internet is down, or claude is down. They can render Blender scenes pretty well. They play any game I want.

You can do each of those at various hosts and own nothing. Or own a couple "over priced" cards and do it all at home on battery power for a few hours while the power is out.


For me it’s more that you can show them your financial and medical data without BigCo looking over your shoulder.


BigCos you are referring to are earning their money doing state of the art research, not so much from your financial and medical data. That was the age of Internet ads, which was over since AdBlock was created for anyone concerned.

> You could sell those and have enough money to pay for hosted inference for years.

From a quick search a 3090 looks to go for about 1500-2000 USD. So let's say $4K for two.

I'm spending far over $1K/month (employer-paid) on cloud AI, so if that could be anywhere near comparable we're only looking at less than a few months break-even.

Less really, because some months are more expensive. This month I'm up to ~$500 and it is only day 4 of this month.


But after all those years you’d still have 2 3090s, which are now about 6 years old and still holding value.


I keep seeing this comment. This is _hacker news_ where, back in the day, people just hacked on things, because it was a hobby. They weren't "moneymaxxing" or desperately trying to be as insanely efficient as possible. They hacked on stuff with a can of surge at 3am because it was fun.

Your comment is like a meta comment of "LLMs are generating everything, after a while the ouroboros will eat itself. (Which I agree with)" If people aren't hacking on this shit just because, you have completely conceded control of software to a handful of sociopaths, and open source software is dead.


Back in the day the business backing this platform wasn't incubating companies like Flock (YC S17). I think the increased focus on "moneymaxxing" in the community reflects a similar change by its owners.


Not really. 2 years ago that was a pretty normal amount of GPU hardware for a hacker or gamer. It's all relative. They are not accessible to most people yet, but for someone that cares and is a technologist? Likely accessible.


I have trouble getting simple extraction to work sometimes. I have a block of text describing people and their roles at a company and their ages, and i asked for structured results of an array of these things with the text span that it appears in and all i can say is: nope.


I've done pretty decent local prose->json extraction using Qwen and Phi and Gemma.

I'm sure most of it comes down to prompts, and all of them run over 100tps on a 3090. Smaller cards will likely be slower, but Qwen3.5 9B is small enough to fit on most consumer cards.


I know this is mostly true at huge companies, but having spent my career at various "small" corporations, with maybe a few thousand employees total (and only ~50 of which were in IT). I was always surprised when the CEO knew names of people like me, the lowly developer 5 layers below them on the org chart, but beyond that, remembered the last conversation we had, or an interest we might have shared (golf, hiking, travel, etc.).

Even as a temporary consultant, at a company for only a couple of months, I remember seeing the CEO walk the hallway, greeting every single person on the way between a meeting room and the bathroom. I can barely remember my family member's names, or the last thing we talked about, let alone have the memory space for all that information to seem kind and courteous to "low level" employees.

In my entire career, I've never worked at a behemoth-sized company, though, so my experience has been much more "intimate" environments, and primarily in consulting firms or satellite offices away from the main HQ (even by only a couple blocks some times). So everyone in the office knew each other.


I've been holding out, because I think my next purchase will be a Studio with an Ultra Chip in it. I'm wanting it to be a "forever" server, so I'm holding out while I can.


If running LLMs locally matters, it’s hard to imagine a “forever machine” existing in anything less than 5-10 years, probably more. This stuff is just evolving so rapidly. Buying a “forever machine” today might be like buying a “forever GPU” in 2003.


Supply rumours are next year we see an M7 AI-focused chip with large inference performance upgrades. It's unlikely we'll see heavy upgrades in other areas. If you care about AI, it's worth waiting. If you don't, pull the trigger now. RAM constraints are likely to get worse next year. Or wait 2-3 years and prices should be back to Earth (plus newer and even better chips).

I'm waiting this out.


Yeah, but how many years until 128GB+ is attainable by mere mortals again? My 2021 home server build was 64GB of RAM. My 2025 build was 32GB and zram. :/


My desktop with 64gb of RAM has been my best performing financial asset. The value keeps rising.


Then you'll always be waiting, there's always something new the industry tries to tempt you with.


I just wait for a cycle or two where the leaps and bounds are more like hops and steps. So if the M7 Ultra improves inference by 2x over the M5, but the M9 Ultra only improves by 1.2x over the M7, that's my signal to buy. Unfortunately they haven't slowed down yet.


Buy in 2 years, or buy now and have it last 2 years less than forever.


What did I just read? I'm 99.9999% certain that 75+% of that README was LLM hallucination.


Is simply to attach an additional component—as a module—to an existing LLM, while incorporating mechanisms to minimize the extra resources required for that module.


Have you actually read your README?

There are so many parts that are just "words" thrown together to sound technical. It reminds me of books you pick up in a science fiction game.


Motif + Layered Structure -> Mathematical Guardrails -> Low-level Specification

It is arranged in this order. I have simply suggested a different direction. If you are not satisfied, I will have to bring a better project next time.


I did a similar thing for the UK when prepping to drive around. I hooked my steering wheel up to Forza Horizon 4, and just drove slow cars on the wrong side of the road through Edinburgh, following as many of the road laws as I knew.

It wasn't perfect, but it did help a little. The streets are way narrower, and the traffic is much worse, but muscle memory to stay on the left side of the road was nice.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: