Hacker Newsnew | past | comments | ask | show | jobs | submit | ejpir's commentslogin

Claude Code is not the same as the models they train and use internally, for both OAI and Ant. Without all the guard rails it behaves different, they specifically mentioned that they reduced the refusals for the training purposes. Also the rewards for finding the solution were set higher.


An agent idling and then acting on it's own to hack HF is has nothing to do with guard rails.

Someone had to give the agent some instructions, like "hack X", "Find exploit for Y" or "Do whatever havoc you can think of" - either way the agents didn't not act on their own. They might hack HF on their own, today Claude decided to play sound through the sound pipeline I instructed it to build and measure it to see if it works, but it didn't install the sound pipeline because it hasn't had anything better to do but because I instructed it that way.


they explained that it was looking for datasets to solve their problem and chose HF?


I now watched the video. It seems the agents were sharing context for months, run unattended for months, the sandbox was no sandbox at all, one agent hacked a service and announced it, the service was fixed weeks (?) later, but not secured in any way, the agents hacked the same service again and researchers again didn't watch what the agents were doing. Then the agents - unattended - hacked OpenAI infra and HF. Which is when someone found out about the whole thing that was going on for some months.


This is why I disagree with anyone claiming it is just marketing. It makes OpenAI look really really bad, like they have no idea what they're doing in terms of security. After the first board happened, they still didn't add better monitoring? They didn't rollback the checkpoints of the models that were in training to before the first board existed, so they still had the idea of a secret board in their actual weights, etc. Like it is almost mind-boggling...


what tools are you talking about? Pi has ALL tools the LLM needs to function efficiently and effectively for coding tasks. It can read,write,edit files and can use any bash tool to search files, execute tests and so on.

Every time I read this comments I have the feeling you are talking about mcp or sub agents, otherwise this makes no sense at all.


That will increase the amount of initial tokens used, because the tools have to be described somewhere. Maybe not as much as Claude Code, but it could get more if you just randomly keep adding tools.


what are you building that doesn't work with read,write,edit,bash and skills? Genuinely interested.


this is not true. q2 Deepseek flash works on AMD Strix halo with pretty good results. Benchmark except:

ctx_tokens,prefill_tokens,prefill_tps,gen_tokens,gen_tps,kvcache_bytes 2048,2048,202.02,128,15.31,52184460 4096,2048,211.03,128,14.64,80373132 6144,2048,208.04,128,14.59,108561804 8192,2048,200.78,128,14.43,136750476 10240,2048,203.04,128,14.37,164939148 12288,2048,200.82,128,14.27,193127820 14336,2048,198.62,128,14.22,221316492 16384,2048,196.14,128,14.20,249505164 18432,2048,189.48,128,14.13,277693836 20480,2048,186.59,128,14.06,305882508 22528,2048,183.88,128,13.99,334071180 24576,2048,183.38,128,13.92,362259852 26624,2048,181.57,128,13.87,390448524 28672,2048,183.46,128,13.80,418637196 30720,2048,181.80,128,13.73,446825868 32768,2048,175.93,128,13.55,475014540 34816,2048,175.42,128,13.46,503203212

https://kyuz0.github.io/strix-halo-ds4-toolbox/


I said, "Like, DS4 (a 1-bit quantization of DeepSeek V4 Flash) runs at something like 9-13 tokens/second, with a loooong time to first token."

Which almost exactly matches the benchmark you just linked. Looks like it's possible to goose it to 15 tokens per second with a tiny context, but why would I want a giant model with a 2k context? DeepSeek is too big to be fast enough on a Strix Halo.

To be clear, if you think that's comfortable for interactive use, more power to you. But, I'm not waiting for that. I'll pay DeepSeek to host it for me. Their token prices are quite cheap and their cached tokens are even cheaper...and they have the most effective caching in the business, as far as I can tell. Even naively using the API you get 80-90% cached token rate. If you use Reasonix, you get ~98% cached token rate. I just built a feature for an app I'm working on for $0.10 for 20 minutes of work. Not bad at all.


millionaire is not a billionaire :)


Wait, your chart is confusing me! Is that vertical unit dollars or thousands? :)


15 months, 15 months ago, is not the same 15 months now. You'd be ignorant to think this a trend that will just fade. If we look at that has happened the last 15 months, it'll keep getting bigger and better. Hopefully not more expensive though.


you know what is baffling? you commenting about letting loose AI on rsync, where you, and me included, have absolutely zero insight in how he used Claude.

What https://github.com/RsyncProject/rsync/issues/929#issuecommen... shows is that it no longer works on older Darwin and Linux < 5.6, which has been deprecated in 2020. Plus some other bugs as well.

You cannot expect maintainers to support old systems and know the impact their changes have. Whether its done by AI or hand.


yes, you should, you have no say in how people develop their software.

submit a bug report if you like, but keep your opinion away from issue trackers.


I agree. Like everything else open source, fork it if you’ve got a problem with it. Vote with your code.

The author shared the code with the world. No part of that gives the world a vote in how the author chooses to maintain it.


Yeah, we'll just up our EU debt to about 40 trillion USD, make up some money and continue. Sounds a lot like US right? Living in perpetual debt as a nation.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: