Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Well... , yeah, but, these models already produce errors for other unrelated reasons, and like...

Well, what exactly would we be showing that these models can’t do? Quines exist, so there’s no general principle preventing reflection in general. We can certainly write poems (etc.) which describe their own composition. A computer can store specifications (and circuit diagrams, chip designs, etc.) for all its parts, and interactively describe how they all work.

If we are just saying “ML models can’t solve the halting problem”, then ok, duh. If we want to say “they don’t prove their own consistency” then also duh, they aren’t formal systems in a sense where “are they consistent (as a formal system)?” even makes sense as a question.

I don’t see a reason why either Gödel or Turing’s results would be any obstacle for some mechanism modeling/describing how it works. They do pose limits on how well they can describe “what they will do” in a sense of like, “what will it ‘eventually’ do, on any arbitrary topic”. But as for something describing how it itself works, there appears to be no issue.

If the task to give it was something like “is there any input which you could be given which would result in an output such that P(input,output)” for arbitrary P, then yeah I would expect such diagonalization problems to pop-up.

But a system having a kind of introspection about how it works, rather than answering arbitrary questions about its final outputs (such as, program output, or whether a statement has a proof), seems totally fine.

Side note: One funny thing: (aiui) it is theoretically possible for a oracle that can have random behavior, to act (in a certain sense) as a halting-oracle for Turing machines with access to the same oracle.

That’s not to say that we can irl construct such a thing, as we can’t even make a halting oracle for normal Turing machines. But, if you add in some random behavior for the oracles, you can kinda evade the problems that come from the diagonalization.



I don't have a rigorous argument but as I recall Rice's theorem says that computers cannot compute nontrivial properties about programs, in general. So we cannot have virus checkers, etc. (Well we actually do have imperfect virus checkers, of course). So following that, a nontrivial property such as self-explanation, e.g. correctly explain what a neural network is doing, for all possible inputs. (But again this is not fully clear to me, someone might point out that for a specific program this is different and Rice's theorem is irrelevant.)




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: