Hacker Newsnew | past | comments | ask | show | jobs | submit | aszen's commentslogin

Their zdr is not a concrete promise though, there's no way to verify that the provider is not storing the logs

Okay, true, but I'm not sure how they could verify that? Isn't that like proving a negative?

I was thinking the same, obvious suspicion is they benchmaxed it on older bench.

Google has always done quite a lot of benchmaxing for Gemini.

Because stronger models are harder to jailbreak from the paper it says haiku was easily fooled into giving us thinking contents


Won't surprise me if the llm just calls sleep after it's convinced it knows all


And does it actually use it


Is your plugin just skills or does it have custom agents?

My own experience is that Claude will use your skills but will ignore your agents or custom search tools.


It adheres to the full scope of skills, agents, and my custom gates. I’d hate Claude without it.

https://devarch.ai/.


Claude code already fans out and sandboxes context by calling sub agents so I'm not sure this approach brings much benefit there. A complex search strategy only makes sense if the search is slow and compute intensive.


i find claude subagents are incredibly slow and the vast majority of that is inference time between tool calls


Coding agents prefer to do iterative search, I have yet to see them create a complex search script. They try different search cmds in parallel, evaluate their results and then refine or dive deeper.

This approach usually works great but I can see many use cases where a smarter search strategy may make sense especially to optimize context.


Wasn’t there work being done where a model could remove irrelevant context? Maybe these complex search scripts can help the model get the revenant files and then it can remove it.


By slowing down engineers with ai agents adding multiple code reviews on top. Also encouraging engineers to engage in manual testing themselves to better understand the product.


Claude code is not infra, the model is the infra. They changed settings to make their models faster and probably cheaper to run too. Honestly with adaptive thinking it no longer matters what model it is if you can dynamically make it do less or more work.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: