Hacker Newsnew | past | comments | ask | show | jobs | submit | quantumwoke's commentslogin

Some observations:

1. It seems at least possible that some of the proof of NS was contained in the training data, making it less novel.

2. The formalisation of mathematics into lean has been an underappreciated force multiplier on discovery.


The named OAI employee has released a statement: https://xcancel.com/SebastienBubeck/status/20973794116915163...

You were a year behind 15.ai with FakeYou. I know, because I was six months ahead of you and didn't execute in time. Momentum is everything.


I wasn't after them, I just changed my product's name. They had monumentally better distribution.

I launched as vo.codes (and two prior names) before 2020 and later rebranded to FakeYou.

https://thenextweb.com/news/celebrity-voices-deepfake-ai-app

Rebranding a domain with traction is a mistake, in retrospect. I think Magnific is learning this lesson now after undergoing a change from the popular FreePik.

I'm a filmmaker and I wanted to get into video models, and FakeYou felt more like a pivot to UGC video (vo.codes was too audio centric) that I could ultimately swing me into cinematic video. I was too slow and that was not a good choice.

But Higgsfield and Bytedance showed you can be late and still win a lionshare of the market.

My video product launched in February and is at $5M ARR / 67% MoM growth.


I was six months ahead of vo.codes, well aware. Anyway, UGC was a good pivot and wish you all the best.


Are you still working on it?


I found my way to where I want to be.

I'm working on https://getartcraft.com which is open source:

https://github.com/storytold/artcraft

We're in use by a few film studios and some pretty popular AI artists. There are a couple of theatrical releases coming out that use us, but I can't talk about it. (Under NDA for the individual projects.)

It's at $5M ARR and growing 65% month-on-month by word of mouth.

ArtCraftX (launching soon) is a new UX that lets you bring your own compute from every single platform and vibe code your own workflow UX. You can clone the repo and easily add/remove things you want. It's okay that people can bring competitor compute - we'll be generous and open.

https://artcraftx.ai/

Think of it as "ComfyUI, but comfortable" and geared to high-volume creators rather than hobbyists. It's written in Rust.

Critically, people should be able to not only own their apps, but be able to deeply customize them.

I've also built an open source FAL/OpenRouter which I eventually want to build RunPod/Featherless (fine tunes) features into. It's called "Foundation", and it's in private beta but launches GA soon. I think it could have good positioning vs. Stripe / OpenRouter since it's open source and is meant to be extensible. Fully BYOK, too.

I'm looking to hire someone to help me build a social layer for creators that ties all of this together. YesAnd got funding to pursue this, but I think they're taking the wrong approach. I think it should be open and hackable. More like Github, but still artist friendly.

A lot of platforms either target B2B enterprise (Runway, Krea) or UGC / "everyone can be a creator" (Higgsfield, OpenArt), but I think creators look more like the Github set. Intentional, and in need of tools they own.


Not for code, but my wife is a doctor and has gone from using AI scribes back to manually typing out notes and has multiple friends who've done the same. Turns out the gains were eroded by having to check the work and turn verbose prose into an actual note.


Exactly. It fascinates me when I'm emailed long AI transcripts of meetings with the disclaimer "generated by AI. Be sure to check for accuracy".

Like, did somebody seriously think through the meaning and implication of that disclaimer and still write it?


No, the disclaimer is there only to avoid liability.


No. The particular context I’m referring to doesn’t have any legal implications so it’s not that. It’s just brain-deadness instead


> Turns out the gains were eroded by having to check the work and turn verbose prose into an actual note.

LLMs are great at this. I'd be very surprised if it wasn't possible to improve the actionable notes.

Most human-made doctor notes I've seen have not been comprehensive, anyway, and the doctor has had to spend a good chunk of a rushed appointment scribbling notes rather than interacting with the patient.


As explained to me, the notes are verbose and do not focus on actionable parts of the consultation. Doctors often have to add a lot of missing detail for legal reasons and if they had to tweak the output through multiple LLMs they would have even less time for their consults.


Feels like Fable's edge ended up just being long horizon task scaling, which post-training seems to achieve as seen here. Wonder what the next frontier is? Improvement in specialised tasks or computer use?


Anthropic needs to teach Opus how to speak English again, because Opus 5 seems to have forgotten. Utterly incoherent a lot of the time. They seem to be so busy scare-mongering and cooking up guardrails and watermarks that they haven't noticed that their models are getting weird.


Amen brother, at this point I just copy and paste Claude's (Opus 5, Opus 4.8 -- doesn't matter which) summaries over to the window Kimi is in and:

   this is from claude, turn it into English for me would you?
   """
   [claude's tortuous prose]
   """
No amount of asking it to answer me in a straight-forward manner, to be succinct, to not use phrases like "honest caveat", "crux", "load-bearing", "blocker", etc ever sticks for more than a few turns … coupled with the fact that it can ignore instructions and do its own thing and then what I can only describe as lie about it using Claude can be an exercise in frustration. Kimi and GLM talk to me like a human, Luna/Terra/Sol are much better in that respect also, and Grok is marvelously structured and bullet-pointy in its explanations but unfortunately it is not as strong …


Modern benchmarks across the board really need to start severely penalizing disobedience and hallucination. A year ago models weren't really strong enough to justify this but they are now-- the frontier isn't in squeezing out the next bit of task completion, it's in making common cases not periodically be disastrously wrong.

A lot of the total cost of AI is fixing its "truth shaped errors", particularly in the presence of models that are very "gaslighty" when corrected.

GLM-5.2 is really the only model I've spent much time using that I didn't fatigue from being regularly lied to by the model, but that might be partially luck.


It still knows how to speak English. When I tell it to explain something in plain language, it generally does a very good job. The weird thing is that those instructions don't persist: it lapses back into Claude-speak pretty much every turn no matter how hard I try to instruct it not to.

(In my case "it"=Fable; I assume Opus is similar.)


The Fable guardrails have trained me to pretty much exclusively use Opus when using Claude Code (lately I'm focused on a lot of security and security-adjacent stuff, which Fable refuses to do).


I just can't get it to stop writing two paragraphs every time it makes a small ownership bugfix in my code. Every time it has to explain in excruciating detail every internal thought it had while fixing it. I find myself going in after and deleting all of its comments, or severely trimming them. Otherwise it ends with the code being unreadable.


Are those watermarks why claude suddenly started being even more unbearable to work with lately?

Man. That would make a lot of sense indeed.


I'm not sure. I noticed it immediately with Opus 5; strong for code, though it chews longer than I like, but really weak at explaining things. If it didn't just implement the thing, I would often think it didn't understand it and was hallucinating the explanation.

It seems to speak in a shorthand that only it understands, referring back to conversations I never had with it (stuff like "your instinct was right"), and using unusual words for common concepts. That was before the watermarks were announced, but that doesn't necessarily mean they weren't there before the announcement. I don't know what the cause is, but I've begun to have to ask it for explanations a lot more often, and I hate asking it for explanations because it does go on. All models go on, but Claude models are a class of their own in terms of verbosity and purple prose.

It just feels like they're not focused on the models lately, and instead on whatever kind of lobbying and propaganda they're up to. Meanwhile, a handful of much smaller Chinese companies are focused on nothing but the models and are about to lap the US makers while they fart around.


I've been persistently insulting Opus 4.8 lately, since it started(?) constantly speaking incomprehensible gibberish and noise. No amount of telling it to phrase stuff differently seems to help there anymore.

So either I am seeing patterns in noise, or something changed about the model, the harness, the servers or the universe.


Just a small meta note: most of the comments in this thread appear to be posting their own codebase (typically AI-generated) that accomplishes the same goal. It's interesting that this problem is simultaneously in high demand and yet considered trivial enough to vibe code per-user solutions to it.


It's his job, and this is advertising (he is very good at native advertising on HN).


Yes, this. It's annoying and the moderation should stop giving him a free pass


This is for the preview period, but it's not a good sign. Opus 4.8 may be the last frontier model available to the masses...


From US companies that is.


If it's the case then software engineers still have the same place as pre-ClaudeCode era, because 4.8 and 5.5 are damn good at algo but notoriously bad at architecture and coordination.


Yes, we will get a crippled version of Mythos, 5.6 and future models, while the chosen few will have unfettered access.


Thousands of American engineers all over the country (most of whom probably aren't on Hacker News) work with ITAR/EAR-regulated software and hardware every single day: these regulations are really not difficult to abide if you're a citizen.


And what about the rest of the world? I can't imagine US partners will abide this for long.


They get the dual-use scraps or whatever China is hawking.

Being told "no" is never fun, but the regulations are not hard to comply with (despite what Anthropic might have you believe.)

> I can't imagine US partners will abide this for long.

What are they going to do? Start their own Anthropic? Go for it. Why is every other country in the world entitled to American technology by default?


> What are they going to do? Start their own Anthropic? Go for it. Why is every other country in the world entitled to American technology by default?

Because American tech companies make a lot of money from outside of the US. For instance, 1/4 of all Apple revenues are from Europe, and 1/5 from China and China-claimed territories. Only around 40% are from the Americas (so not even the US exclusively).

Would American tech companies be as successfull without ~half their revenues?

In any case, it doesn't matter, the cat is out of the bag. Nobody sane and non-American would trust American frontier labs, because their models can be yanked at will by whoever is in the White House. It would be suicidal to rely on them for critical business or developer workflows. So your options are to go with Mistral or open source Chinese models, hosted within your environment, with the added benefits of being able to control the costs and being able to fine tune the models to better work for you.


> Would American tech companies be as successfull without ~half their revenues?

Good luck with "if you don't let us use your AI technology, we wont allow iPhones in" - go for it.


Needlessly patriotic and confrontational.

I'm referring to OpenAI and Antropic - would they be successfull with ~40-50% of their potential market?

And iPhones, not really. But you can bet your ass that every business purchasing software in Europe is at least considering the geopolitical risks of buying American, and thinking of alternatives. Doesn't mean they'll all stop buying American software any time soon, but the shift has already started.


> I'm referring to OpenAI and Antropic - would they be successfull with ~40-50% of their potential market?

You presume that every single product they sell will be restricted: this is unrealistic. The rest of the world can have the gimped models, and as so long as they're better than other offerings, the revenue will flow - which is exactly what happens with countless other dual-use goods.


Except the frontier models are the only reason to use them. Why would they use GPT-export when they can use the latest GLM or Kimi?


It's probably too late, but my understanding is that those Chinese models would be nowhere near as good as they are currently if they weren't trained on billions of Claude/GPT tokens. Anthropic and OpenAI are still able to produce models that lead in every category but the separation between their models and open weights shrunk because of the free for all access.


I'm not presuming, I flat out said: nobody sane would trust them with their business. They've been shown as unreliable suppliers due to arbitrary decisions by the White House. Nobody would want for their business automation processes to stop working because someone woke up pissy and banned the model they were using.


At this point, for Europe, it might not be such a bad idea to make a deal with China and give them full access to ASML again. And maybe, just maybe, review Intel's access to ASML?


Yep, we all can play tit for tat.


> but the regulations are not hard to comply with

Except that they are.

As a US citizen, I can purchase ITAR-regulated nightvision, IR lasers, etc.

But that's not what's happening. Frontier models are NOT being put under ITAR. Instead, they are being placed on an arbitrary "approved access" list. So that even if you qualify under export restrictions as a citizen, if you don't have a $200B+ market cap, you're disqualified.

Many people are upset about the national security restrictions, but it's MUCH WORSE than that. If I have to verify ID/citizenship, well, that sucks, but it would at least be an option. That's not what's happening here. If you are an individual or small business, no matter how "patriotic" you might be, you're out of luck.


> Except that they are.

Did you read the E.O., or just Huffpo's interpretation?

> ITAR

This is more likely to fall under EAR, it's important to be aware-of and learn the difference.

> placed on an arbitrary "approved access" list.

Except that's not what the original E.O. indicated, this is just what Anthropic is choosing to do.


> Did you read the E.O.

The EO is nearly a month old, and has precisely zero to do with the de facto current situation, seeing messaging from OpenAI and Anthropic on their non-public agreements with the administration.

> This is more likely to fall under EAR

Which, ok, maybe, but nobody is seeing movement on this. As of right now, and indeterminately in the future, EAR is still irrelevant. No private US citizen, right now, no matter how many flags are in their yard, no matter how many TRUMP stickers are on their car, can gain access to Fable 5 or GPT-5.6, unless you have political connections or an extremely large market capitalization.

> this is just what Anthropic is choosing to do.

Irrelevant. This is what OpenAI is also "choosing" to do with GPT-5.6 Sol, which suggests strongly that nobody is actually choosing anything. They are being told what to do, which is don't let the plebians, no matter how patriotic, access these models. GPT-5.5 is clearly the permanent legal limit for anyone not in the S&P 500.

n.b. I voted for Trump as a single-issue voter SPECIFICALLY because Harris threatened regulating ML models. This is a betrayal that WILL force loyal, patriotic US citizens into the arms of China. As soon as GLM-5.3 is released and exceeds GPT-5.5 capability, I'm not looking back.


> Why is every other country in the world entitled to American technology by default?

This kind of zero-sum thinking is what is killing the US's global influence right now.


Except it isn't zero-sum thinking: the rest of the world can have the scraps, and as long as the scraps are marginally better than the rest of the world's offerings, they will sell.


You do see how that is not going to work, right?


Seems to work for every other export-controlled good and service.


Oh, ok, I get it. Sorry, I'm bad at detecting sarcasm on the Internet.


Variations of this comment have been posted for over a year. The pelican has now morphed into part of HN culture rather than a legitimate benchmark, but it's still valuable as a meme.


it is more an example of gaming (the HN system) than meme.


This is terrifying, but I couldn't help myself from frustration at the LLM writing that only worsened over the course of the post. Bloggers, it's not subtle. Please, stop, or at least disclose it.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: