Hacker Newsnew | past | comments | ask | show | jobs | submit | Spacecosmonaut's commentslogin

Isnt the sensory grounding even in abstract cases some (limited) intuition that simulates in a mental world model?


This seems temporary. AI generated images will become the norm for many applications. Today we wonder whether a model in an advertisement is real or generated, 5 years from know this wont even be an interesting thing to consider.

In many contexts, images are meant to illustrate a point. The human becomes the gatekeeper of generated content. The question becomes: does this picture illustrate the point I want to make. If so, you are adding your human approved stamp to the genrated output.


I wonder, how does distillation deal with unprobed spaces in the knowledge landscape? Is a distilled model worse in some niche area that was not probed? Presumably, this is why frontier labs dont distill their own models internally to release them to the public as a servicable frontier model.


Do frontier labs not distill their own bigger models into the smaller/cheaper variants? I thought that’s been the case for a while


I wonder how much of the sloppy feeling is created by the knowledge that its AI generated.


I wish I could see it with fresh eyes without the accumulated knowledge I have of how AI looks. What does this video look like to a person clueless about AI?

What do they think of the constant repetition of slow panning shows where barely anything is happening, or a guy spinning on the floor but his head morphs into his feet and his feet morph into his head, or a person kissing a mirror that isn't quite mirroring the person and it slowly becomes a person kissing what appears to be their identical twin standing on the other side a floating empty mirror frame?


Same. I think there's a lot of AI prejudice influencing people's opinions on AI output.

I wonder what would happen in a study where some art is critiqued by one group that is informed that the art was created by AI and another group that is uninformed.



Accelerationists may argue that the eroding of proper attribution and proof verification by humans is a meaningless short term struggle of a dying field.

Mathematics seems to be entering an era where human + machine maximizes performance, much like chess in the 1990s. However, imagine a future where even talented mathematicians are nothing but noise in the machine (as is the case in chess now). A future where AI generates and verifies proofs without humans in the loop. Where the mathematics may be beyond human comprehension.

In that future, does it matter that early career mathematicians are inhibited by these developments? Perhaps not. Programming faces the same issue. As AI crawls up the competence ladder, does it matter that fewer people have opportunities to develop the skillset of a senior engineer? Perhaps not.


Chess has set rules and is a closed world with a set objective. In maths, you make up the rules. For every clearly defined problem that everybody cares about and is not yet proven, there's someone who first who recognised the importance of that given problem and managed to define it clearly enough for it to be recognised as such.

There's also the separate, less glamorous issue that people don't want to talk about, which is proof reliability. [0] If you have systems to help you formalise the problems and leave an algorithm or AI or whatever solve it in a verifiable way, that's a win for both the mathematicians and the rest of the world.

The deeper question is whether AI can replace the human role in deciding what mathematics should be done and what concepts matter. If that's automated, then yeah, we're screwed.

[0] https://lamport.azurewebsites.net/tla/proof-statistics.html


> However, imagine a future where even talented mathematicians are nothing but noise in the machine (as is the case in chess now).

Isn't chess more popular than ever? Ai dominating the game didnt seen to matter


Chess is a game where we play and observe for the sheer joy of mastery. Mathematics has some of that “done for its own sake” aspect but not as cleanly. There’s always the chance that some applied mathematics might help something in science or engineering. Whether we care more about process vs outcome is now more important than ever before.


Sure, but if we continue that analogy it does mean that there will be no human contributions to frontier mathematics.


(This doesn't disprove your point because it seems like it should be trivial to make but interestingly)

I don't think there's any computer system which autonomously come up with new directions in openings. As far as I know; a GM looks at stockfish's evaluation of the top x moves and analyzes one that hasn't been played a lot etc.


If the output is better who cares?


> imagine a future where even talented mathematicians are nothing but noise

That would be AGI. My conjecture is that LLMs alone are not enough for that future. They are incredible, but AGI needs other breakthroughs.

In that sense, I think math is very different from chess or Go. Chess and Go are complete-information games with fixed rules and a fixed board. Math is open-ended.


Keep dreaming


Accelerationists may argue that the eroding of proper attribution and proof verification by humans is a meaningless short term struggle of a dying field.

Mathematics seems to be entering an era where human + machine maximizes performance, much like chess in the 1990s. However, imagine a future where even talented mathematicians are nothing but noise in the machine (as is the case in chess now). A future where AI generates and verifies proofs without humans in the loop. Where the mathematics may be beyond human comprehension.

In that future, does it matter that early career mathematicians are inhibited by these developments? Perhaps not. Programming faces the same issue. As AI crawls up the competence ladder, does it matter that fewer people have opportunities to develop the skillset of a senior engineer? Perhaps not.


Much like for many the point of chess is that it's played by humans, with truly superhuman AI relegated to a training aid, mathematics is in many ways about human comprehension. You can use AI to find and proof new theorems. But if you get to the point where humans can't understand it, is it even still math?


Neural networks are already systems of linear algebra that are beyond human understanding. Most humans could probably grok a 1 or 2 dimensional slice of a network, but the latent vector space is completely beyond the human brain. We have to use tools to analyze neural networks piecemeal in exactly the same way that we analyze any other higher-dimensional construct. Few humans are truly capable of reasoning in 4+ dimensions, that doesn't make string theory "not math". Nor does a trillion-dimension vector space of an LLM make it "not programming".

Humans by themselves invented mathematical concepts beyond human understanding a long time before we invented neural networks.


Perhaps P=NP. The new algorithms are handed down to us. We can apply them without fundamentally understanding why P=NP.


> Much like for many the point of chess is that it's played by humans, with truly superhuman AI relegated to a training aid

It's much more than just training. Humans use the engines to prepare openings and find promising novelties. Over time these novelties unearthed by engines fill out theory. It's easy to fine elite games where neither player is out of book for dozens of moves. Modern players are full hybrids in that sense. Looking back at chess, it seems natural that Mathematics will go the same way.


I think there would still be a place for it if it's beyond human comprehension. For instance, really complex lemmas to solve human-tractable problems. If you can pose a question in a proof assistant language like Lean, have an AI write a Lean program that solves it, you can use that as a Lemma for some other problem. There's quite a bit of math out there that is "correct assuming conjecture X is correct", maybe AI could fill that gap and "still be math".


Is it still chess, if humans cannot understand it? Because that's the point we are at in chess. Engines making moves, that humans cannot understand, but somehow they work out to be best or seemingly best. Look at the Leela Zero games, when it came out. These engines play kind of other-worldly chess.


Exactly. If we wanted "the best chess" we'd watch Stockfish against Leela Zero. Far better mechanically than human chess. But people are much more interested in Magnus Carlsen playing Gukesh. Both train with chess engines, but the thing that makes the chess game interesting are the human beings that understand their own moves and try to understand those of their opponent


An issue I see is who controls the information. The next generation may not recieve the knowledge, it may be gatekept by industry who *will* own the gate.

The future may not have access unless we fight to ensure they do. This is how I read the article.


Surely such AI would also be able to reduce and simplify the math for human understanding. Which is what mathematicians do all the time, from turning base 60 cuneiform into modern number systems to simplifying Maxwell's equations for the students.


In general, most humans top out at 3D visualization, and instead rely on crude mathematical tools to work with higher dimensions. Every so often, people like Euler or Leibnitz pops up to give people new methods for blind men with a cane to explore the unseen yet knowable world(s).

Scientific work is not normally naturally statistically salient for LLM observational data inferences. =3


On the other hand, it could stall out at: good enough to take the easy problems, not good enough to take over the field, but damaging enough to erode the quality of new entrants. (Which incidentally is the scenario I think plays out for software)


I think the OpenAI model that resolved the Unit Distance Problem would be capable of solving a significant proportion of mathematics PhD thesis problems.


AI (in this form) will never be able to solve things we truly cannot solve yet. It might catch things that we didn't project properly or brute force things no human can , but it will never unify general relativity with quantum mechanics. It's amazing at finding hidden truths in large datasets, but won't win a Nobel unassisted.


> AI (in this form) will never be able to solve things we truly cannot solve yet.

Argument?


The strongest argument for this is structural: what LLMs are.

In a brutal simplistic way: each token is represented in a high dimensional vector. LLMs operate on them. They are the true, underlying meaning of the token for the LLM. Think of it as 1000+ ways to think of that word/token. Those meanings are baked in at training time. So, LLMs might be able to cross-reference them and solve a class of problems that flew under our radar, but can't come up with revolutionary theories that were never in the training set.

Of course, they will help winning a Nobel in the years to come, no doubt, but can't speak mathematics we can't understand (beyond simple obfuscation) and won't discover anything substantial on their own.


> but can't come up with revolutionary theories that were never in the training set.

Can you elaborate? I don't think the solution to the unit distance problem was in the training set, but I'm guessing you mean there's some higher bar for revolutionary theories LLMs cant reach? If so where do you expect the limit will be?


Instead of going into a long technical argument of why your description of LLMs is flawed, I'll go straight to the point, because people keep moving the goal posts.

What exact problem would need to be solved by LLMs to convince you that they DO discover novel solutions?


I'm more interested why you think my understanding is flawed honestly. I thought I distilled it decently well in two sentences. The bottom line is, in this hyperdimensional space you can find relationships that are not easily distinguished by human minds, but the corpus is still fixed, a llm can't truly know anything beyond its training data.


> Think of it as 1000+ ways to think of that word/token

I assume you used 1000 because that's in the ballpark of the vector size. But these are not independent scalars, like each might store a certain property. Just like in 2D you can have 4 quadrants (or subdivide further), with a vector of size 1000 you can encode an insane amount of meaning.

> Those meanings are baked in at training time. So, LLMs might be able to cross-reference them and solve a class of problems that flew under our radar, but can't come up with revolutionary theories that were never in the training set.

There's a lot of jumping to conclusions here, but I'll try to answer more generally.

This idea of how LLMs work is mostly to build an intuition, like with a CNN you'd say imagine a layer does edge detection, and so on. And to some degree you can detect those kinds of behavior, but a NN is a VERY general architecture. It needn't work like you say, it can calculate any function and running under a loop and a scratchpad (basically an agent) is turing complete.

Even ignoring that, this part is misleading

> Those meanings are baked in at training time.

Being baked in at training time does not mean it didn't build novel meanings at training time.

This is even more significant when you take into account post training RL.

A simple proof that transformers can generate novel, superhuman solutions, is that you can build a transformer based chess bot, feed it 0 human games, and train it with RL until it can beat any human, completely novel and unconstrained by human gameplay (because it would've never seen it).

You can do that with any task that's verifiable, like coding or math.

(Also as a separate fact, as long as a task is easier to verify than solve (basically always), you have somewhat of a million monkeys with a typewriter, and with temperature sampling the model might eventually stumble it's way onto a solution.)


unify general relativity with quantum mechanics. The continuum hypothesis. The traveling salesman problem in polynomial time.


I think it's cool how in a decade we went from

"Neural networks will never be able to understand this sentence that's obvious to humans"

to

"LLMs must be able to solve problems that humanity hasn't been able to after almost a century, and that might even be unsolvable"


So that is kind of the point of studying maths right?

Why something in unsolvable or undecidable can be as important as the output of a theorem.

Questions like these, fields medal level problems or Karp’s 21 NP-complete problem are problems working mathematicians are interested in.

Will LLMs help as an human assistant in the future? Probably.

Will LLMs answer these questions themselves, provide insights and bounds to these new mathematics and teach other mathematicians why this new math they create is true?

Will these models have phds and take candidates teaching them how to apply and think about the maths problems they are interested in?


it can operate at the level of a mere mathematics professor, who everyone knows are barely conscious, basically automatons. wake me up when it's Einstein


The continuum hypothesis was proven independent of ZFC over sixty years ago, I think even GPT2 could have told you that much.


I don't see how any of this follow. Yes, the LLMs will learn the "meaning" (here narrowly defined as relative configuration in the embedding space) of vectors that correspond to tokens in whatever tokenizer is used to feed into them. But that vector space is not discrete, and nothing precludes the model from internally operating on other vectors that it never saw in training, based on how they relate to those vectors which it did see.


We have yet to see evidence of proper generalization AFAIK. Examples such as this proof are the closest I'm aware of. I haven't read this one in detail yet but the other examples I've seen have been (upon examination) much closer to an (absurdly) deep literature search than to novel thought.

Obviously that doesn't mean we won't eventually achieve novel thought, or even that the current form is fundamentally incapable of it, merely that we've yet to see evidence of it and thus the default assumption is that we aren't there yet.


The burden of proof is the other way


In general, most researchers already incorporate LLM into their workflows, as it is quite good at context search. However, the relevant training data is based on the collective works of the field of experts. Collecting current data on that work is what makes the LLM sound relevant, and any improvement of the LLM model requires frequent new data from both researchers and the chat bot users themselves. LLM are not real "AI", and anyone that says otherwise is selling people something.

To phrase this differently, LLM companies conduct unauthorized targeted intelligence gathering on peoples work, codify that act of plagiarism or theft as MoE documentation, and sell unaccountable token output to other users.

There is a reason output becomes more nonsensical as "AI" companies try to use dynamic weight granularity and conceptual compaction. It is not necessarily "AI" hallucinations, but rather people fooling themselves into believing smart people are no longer needed if they willingly become a hapless exploited data source caste. This simply isn't true, as people will leave the field for awhile.

The LLM business model regularly requires copyright theft and plagiarism to persist. It will not magically become sentient/AGI/less-stupid, as these algorithms have been operating for over 40 years. What has changed is the scale of the deployment, data pool size, and the energy consumed.

Scientists are still necessary, as they create the world models LLM try to guess at by statistical inference. Hype and FUD ahead of an IPO for a highly dubious revenue company is expected. We look forward to the low cost liquidated GPU hardware in the near future. =3


Reading this invoked the image of ouroboros in my mind.


The Ouroboros in western mythology is a cautionary tale about the uselessness of the first perfect immortal being, and why humans should suffer our imperfections with insightful grace. The concept also made a great Red Dwarf episode.

LLM are more like the Mechanical Turk trick, but the persons inside the machine running the con is unaware of how their actions affect the confounded observers.

Have a wonderful day =3


Thank you, human being! Have you looked at the price of a RTX 5090? I can get a used car for that.


Indeed, just paid $3k more for the same workstation we purchased last year at this time. Just the DDR5 sticks and NVMe drive cost more than most parts right now including a rtx 5070 Ti 16G card. For h265 hardware encoding, the performance differences on higher-end cards benchmarks was negligible.

Building systems based on application specific benchmarks rather than general what-if use-case scenarios will sometimes show you something interesting. ymmv. =3


"Generative AI is art. It’s irredeemably shit art; end of conversation."

I think most people cannot destinguish between "genuine" creativity and an artificial almalgamation of training data and human provided context. For one, I do not know what already exsists. Some work created by AI may be an obvious rip off of the style of a particular artist, but I wouldnt know. To me it might look awesome and fresh.

I think many of the more human centric thinkers will be disappointed at how many people just wont care.


Further I'd argue we KNOW people don't care if you look at the music industry.

Pop music is often composed by dozens of people who specialize in a thin sliver of the track - lyrics, vocals, drums, &c. - and then it's given a pretty face and makes the charts. That's really no different than something like Suno.

I think AI is forcing people who thought that THEIR thing was too nuanced or too complex to be replaced by technology to reckon with what makes them special.


The question is how subtle AI can be. I feel like art sometimes seems to communicate A, and the artist intended to communicate A and perhaps some B, but clearly, it also hints at another C (and maybe also D, E, ..), which was not intended by the artist or recognised by many viewers, while to some people it's clearly there. Now where did that come from?

And can or will AI create it?


People have been having this debate with popular art forever. Some people do not even believe in taste, and that everyone's artistic opinions have equal merit.


most people are just utilitarian and do not care for "art" (in the high art sense).

AI is perfect for that. It reveal, perhaps to the dismay of those who revel in high art, that it might be an illusion that art has genuine creativity, if most people find ai to produce acceptable output.


There is genuine value in a wisdom of the crowds assessment of future events that is only sharpened up by the requirement to put money on the line.


The often-repeated "wisdom of the crowds" justification is misapplied to online betting markets. Like people, crowds can either be wise or unwise depending on the situation. Famous experiments like guessing how many gumballs are in a jar work because each person who can see the jar has a source of valid information, and in aggregate that can be surprisingly accurate.

You can't assume that the majority of individuals participating in betting markets have a source of valid information. Given the destructiveness of these markets to both individuals and society, the aggregate wisdom of the individuals participating in these markets is highly doubtful. Any meager value above more traditional forecasting does not justify the cost, corruption and a loss of trust in institutions.



Please show the dollar/realized benefit to society VS (in response to OPs statement) the results don't "justify the cost, corruption and a loss of trust in institutions" along with a breakdown of the cost/negatives to society that result from those factors.

This isn't big oil (yet) you can't just externalize all the downside and say the product is a net benefit.


Pish tosh, my dear sir, it's simply common-sense that there are oodles of people out there with secrets that would be completely ethical to distribute and would undeniably better all humankind, but they're sitting on them purely because they haven't figured out how to make a profit from it. /s

In other words, the overlap between these is too small to justify the idea that prediction markets are a net-benefit by default:

1. Is valuable

2. Not already known

3. No current reward mechanism exists (e.g. patents)

4. Not criminal or unethical to disclose


Well-evidenced by marketing from those same betting markets, got it.


I'd be interested in learning in what ways it is factually incorrect.


There is definitely value in that, but that value is outweighed - dominated, even - by the incentive produced to fix outcomes, incentivized by that money put on the line.


How dare you impugn the mystical powers of the wisdom of the crowd. Why would anybody fix outcomes?

Wait a minute, you can bet on pro wrestling? OK I'm out of ideas.

https://www.betus.com.pa/sportsbook/entertainment/wwe/


It would be useful to predict things like earthquakes and tornados. Gambling on what politicians and celebrities will do is not science it's degenerate court gossip.


Game theory can be applied to these problem, and if you can correctly model reward, cost and motives, you can predict decisions of politicians more often than not.


> you can predict decisions of politicians more often than not

What makes you believe this? The performance of economist/sociology experts using game theory to make predictions has been worse than a coin-flip up to this point. It also has done enormous damage.


Even granting the idea that game theory can be applied successfully here; that does not really help with one-off events. Consider, knowing the odds of a coin flip does not grant you any real help in knowing what the next coin flip will be.

This is also ignoring that game theory of partisan games breaks if any of the participants knows what the other will do. Is one of the more famous ideas.

To that end, if you want to predict what someone will do, more often than not you are best looking at their experience doing said thing.


I don't think Nash envisioned politicians and their aides being able to profit from making decisions. Do you launch an attack on Eastasia? Well the market right now says there's a 40% chance, so I guess it's a good idea to grab your crypto keys and make some bets before you call the Joint Chiefs.


I think you're right broadly, but when that wisdom is applied to e.g. how many dildos will be thrown at WNBA players, I don't know how much actual value is created.


That… that actually seems like something which society should want to predict with more precision so that we can hire someone with a net on a stick.


It's also a real world event that can be influenced by the existence of the bet.


There's value in that only if assessments made by people using the prediction for something aren't better than crowd wisdom. I would guess that a large part of prediction market participants are simple gamblers whose assessment is worse than the prediction of someone doing it because they need the information for something.

So if people who need predictions for decision-making start relying on prediction markets with a lot of low value gambler predictions and opinion manipulators (wealthy participants can use the platform to mislead people as well) the value of these things might be negative.


Or if decision makers start using the gambling markets to drive their decisions, the value of these things will go extremely negative.

The headline bet in Polymarket right now is on when US troops will invade Iran. If some unscrupulous official can pull some strings to get boots in the ground in the next few days, they stand to win a significant amount of money.


Genuinely asking: has there been a case of the prediction markets being "right" or valuable about something outside of the norm?


That's why the metaverse happened. Lots of money bet on it's success


My understanding of the world is enhanced immensely because suckers can bet on how many times the president will reference the lips of his press secretary in a speech.


Wisdom such as the mass purchase of GameStop shares, temporarily inflating the value of an ultimately doomed company?


Regarding point 2: I think most people cannot destinguish between "genuine" creativity and artificial almalgamation of training data and human provided context. For one, I do not know what already exsists. Some work created by AI may be an obvious rip off of the style of a particular artist, but I wouldnt know. To me it might look awesome and fresh.

Furthermore, I think many of the more human centric thinkers will be disappointed at how many people just wont care.


AI tools are here to stay. They will start to creep into everything, everywhere, all the time. Either you recognize the moment at which it becomes a significant disadvantage not to use them (I agree that moment is not now), or get left behind.


The metaverse is here to stay! Blockchain is the future!

Without integrating metaverse and blockchain features into Firefox, Mozilla is at a significant disadvantage compared to other browsers. Don't get left behind!


They did actually jump on metaverse with Firefox reality and Mozilla hubs. Both weren't bad products at all. Both are now cancelled and they have done basically nothing for Mozilla's market position.

Edit: so I mean I agree here in case that wasn't clear


Many thing are "here to stay", should Mozilla also implement a "share with TikTok" functionality into their browser?

> or get left behind.

Last time I heard this phrase it was about VR, and before that it was NFTs. I wished the tech community wasn't so susceptible to FOMO sentiments.


Indeed. I never understood, let alone bought into, the NFT hype, but I think VR is a good reference point for AI:

There was a real, genuine product in the Oculus Rift. It did something that was an incremental improvement over the previous state of the art which enabled new consumer experiences for low cost.

The Metaverse was laughable, and VR got glued to a lot of things where it added zero value, or worse negative value, for example my attempts to watch pre-recorded 3D video gave me nausea because the camera can only rotate, not displace, with my head movements.

Compare and contrast with AI:

LLMs and Diffusion models are also real, genuine products, that are incremental improvement over the previous state of the art which enabled new consumer experiences for low cost.

A lot of the attempts to integrate these AI have been laughable, and have added zero-to-negative value.


Non corporate VR is actually doing some interesting things - but yeah, what Meta did with it was pure garbage.


I didn't mean it as VR being useless - I'm sure it can be useful for some applications or fun for gaming - my point was that you shouldn't fear getting left behind just for not having an Apple Vision Pro app or a land in the Metaverse :)

Another way to see this: Hammers can be useful, the Internet can be useful, but this doesn't mean that as a hammer manufacturer you should make your next hammer an IoT product ASAP or you will be left behind.


Well stated, agreed. :)

Just wanted to note that even after the bad publicity that companies like Meta (ugly avatars, unusable bland virtual spaces) or Apple (overpriced device with no software or content) have given to VR, some people tend to regard it as dead even though there is quite a vibrant user and creator community doing some incredible things (even just what people do with VRChat is amazing!). And there are even companies that seem to get it (Valve).


I don't think it's quite that simple. A great deal of work has nothing to do with computers, and even more human activity has nothing to do with economic advantage. The scope of your statement is a bit too broad in that regard but for computer based work I think you are a) more or less right but b) if you are right it's not clear how much economic benefit LLMs will actually provide on balance, long term.

Does it make the world a better place, and more prosperous? Does it just move economic activity around a bit in regards to who is doing what? We'll find out in ten years when the retrospective economic studies are done.


People wonder why there's a backlash when the pro-AI side sounds like the Borg.


I disagree and I think the moment is now. Gemini 2.5 and now 3.0 is incredible. People that don’t recognize that and use AI tools now are as silly as a craftsman that uses a hammer as a screwdriver when he has a screwdriver in his toolbox. A good craftsman uses the right tool for the job to save time and do a better job and knows the limitations of each tool.

I can spend hours learning photoshop and then trying out color schemes for my new intricately detailed historic house or removing a car from the driveway or I can use Nano Banana and be done in a prompt. There is dignity in learning all that minutiae but I don’t care, I’m not a Photoshop artist, I just want the result and to just move on with my life and get the house painted.


AI crap has already been crammed into everything for months now, and nobody like it nor wants it. There is no proof that AI will continue to improve and no certainty that it will become a disadvantage not to use them. In fact, we are seeing the improvements slow down and it looks like the model will plateau sooner rather then later.


> There is no proof that AI will continue to improve and no certainty that it will become a disadvantage not to use them. In fact, we are seeing the improvements slow down and it looks like the model will plateau sooner rather then later.

While I expect the improvements to slow down and stop, due to the money running out, there's definitely evidence that the models can keep improving until that point.

"Sooner or later", given the trend lines, is sill enough to do to SWEng what Wikipedia did to Encyclopædia Britannica: Still exists, but utterly changed.


> While I expect the improvements to slow down and stop, due to the money running out

This will certainly happen with the models that use weird proprietary licenses, which people only contribute to if they're being paid, but open ones can continue beyond that point.


The hyperscalers are buying enough compute and energy to distort the market, enough money thrown at them to distort the US economy.

Open models, even if 100% of internet users joined an AI@HOME kind of project, don't have enough resources to do that.


You are right, machine learning models usually improve with more data and more parameters. Open model will never have enough resources to reach a compatible quality.

However, this technology is not magic, it is still just statistics, and during inference it looks very much like your run of the mill curve fitting (just in a billion parameter space). An inherent problem in regression analysis is that at some point you have too many parameters, and you are actually fitting the random errors in your sample (called overfitting). I think this puts an upper limit on the capabilities of LLMs, just like it does for the rest of our known tools of statistics.

There is a way to prevent that though, you can reduce the number of parameters and train a specialized model. I would actually argue that this is the future of AI algorithms and LLMs are kind of a dead end with usefulness limited to entertainment (as a very expensive toy). And that current applications of LLMs will be replaced by specialized models with fewer parameters, and hence much much much cheaper to train and run. Specialized models predate LLMs and we have a good idea of how an open source model fares in that market.

And it turns out, open source specialized models have proven them selves quite nicely actually. In go we have KataGo which is one of the best models on the market. Similarly in chess we have Stockfish.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: