Hacker Newsnew | past | comments | ask | show | jobs | submit | jmugan's commentslogin

It really breaks the illusion for me when a movie is shot where it isn't set. I'm like, "What? They don't have pine trees in Austin."


I was watching a terrible movie, Wolverine, and the titular character is driving through Canada, we're told, and I found myself thinking that "huh, that bit of Canada looks a lot like New Zealand, interesting."

Then he drives across a one-way bridge, and tracking shot above the car captures the big arrow New Zealand paints on the exit lane indicating you should drive in the *left* lane (started doing this after few too many tired tourists from RHD countries reverted to habit upon leaving a one-way bridge causing a head-on collision further down the road). They could've edited the arrow out, surely.

TL;DR thought Canada in a movie looked like New Zealand, it was New Zealand.


A lot of people are posting here about how bad the end product is, but that is kind of the point. Models have moved beyond generating images to a new kind of benchmark that better exposes understanding of the physical world, and we can use benchmarks like this to measure future progress. (Of course, it will have to be a qualitative/subjective measurement.)


Agree. The pelican benchmark was interesting a year ago when most models struggled and a good pelican indicated an unusually capable model. Now it’s saturated and uninteresting.

A good new benchmark should have awful performance to start and there should be a lot of headroom for improvement. This benchmark is also intentionally difficult and requires the LLM to develop the animation through spatial reasoning and first principals rather than existing video generation pipelines. Similar to how SVG generation was out of distribution for most models a year ago.


I think it's interesting to see them visibly struggling to improve. Claude pelicans aren't much better today than they where 18 months.


The code I see is a lot like the pelicans. All of the code in codebases, good, bad and ugly, is slowly being replaced by whatever level of code ai is currently able to create. All code is now a slightly wonky pelican on a bike, but if you look closely, it doesn’t fully make sense. Since ai is converging on less wonky, but not internally consistent, we’re just moving on to what is possible with high volume instead of detailed quality. I think that is the ai software world as well.


Rendering 3d worlds has hugely improved though.


they are quite good also at placement and creating scenes etc. I had one implement a cascading shadow system in vulkan/glfw and just fed it back screenshots with peter pannin and acne spots etc.

it made a huge monstrosity first, 1500+ lines of shader code. Once it was happy with the result (it looked pretty good, almost blenders gamerenderer) it cleaned it up and a lot of debug code was removed. shrank it down to about 350 lines and made it much more readable.

this was something i didnt expect it to be able to do. create shaders, look at screenshots, fix em iteratively like that. multi modal debugging.


Have you seen pelicans in Simon Willison’s tests? It is still not a pelican on bike I would like to publish :)

Imho we don’t need to make benchmarks that draw the whole 3D world. Pelican’s drawing is really nice in its simplicity and complexity at the same time.

It seems to be an obscene waste of compute time to generate useless 3D worlds that are just a bragging - 3D is really heavy discipline to make it right, see Mark Zuckerberg’s ceased attempt with 3D VR…

Multiply it by thousands times as a lot of people have found out threejs lib and prompt “generate 3D world and make no mistake” are new orange/black.


Do you expect the SVG to emulate a hand-drawn picture, become more realistic, or just a more detailed illustration?

As for the often quoted issues with the bike's frame or problem with the steering column, I can't really tell, I am no bike expert.

I can instead judge how poor of a job it is doing with a LOTR rendition in Three.js, so that seems like a better benchmark.


A general benchmark (even Simon mentioned that it was meant as fun at the beginning) should be quick and easy to run, since we can expect that more people will want to try it out. That’s why I’m more like “team Pelican on a Bike”… :) cheers


At some point, I'd think labs would start "teaching the test" and start adding bike riding pelican's in the fine-tuning.


> useless 3D worlds that are just a bragging

No, this demo is the useless 3D world, and you're bragging.

A real game would have a lot more immersive of a world, and you wouldn't need to.


Will Smith spaghetti was garbage a couple of years ago and now AI videos are becoming close to indistinguishable from real videos in many cases.

Spaghetti 2026:

https://x.com/dreamingtulpa/status/2083304533829066873

https://xcancel.com/dreamingtulpa/status/2083304533829066873


Wiki page for people like me who never heard of this test: https://en.wikipedia.org/wiki/Will_Smith_Eating_Spaghetti_te...


I wonder though if models are now "benchmaxxing" against these kinds of prompts. I haven't needed to use them so I can't say, but would be interesting.


I don't know I think it is charming in a way that is lacking in nearly everything else an LLM tries to do creatively.

I have always preferred the result of getting an LLM to draw an svg or make a procedural animation like this to the uncanny hyper-realistic result of diffusion image/video generation.


To help explain AI to my elderly mother, I said it is like an alien intelligence on a planet too far away to directly observe Earth, and everything it knows about humanity and our world it learned from reading just about every book and website. You can ask it a question or to do some work and it can often give a useful response but it can never directly judge its accuracy if it relates to the physical world, it has only your judgment to work with.

Given that limitation, it is incredible what it can still accomplish. And when it falls short, it is often in a charming way, if you are open to seeing it that way. It reflects something like a child's understanding of the world, not entirely wrong, just incomplete.

The fluttering cubes that might have been bees or butterflies were my favorite part.


That's a nice metaphor for your mother, i like it. I would just touch on one small part, which is that it has more than our judgement though. It has access to laws of physics, to formulas, to all our current scientific knowledge which we have shown to be correct by interacting with the physical world. It can create an (incomplete) model of the world based on verified theories already. Through deduction and reasoning alone it can get pretty far. After all, there's many scientific theories we created long in advance before we actually proved them to be correct through interaction with the physical world. So even before they were proved to be correct, they were already correct. Same can be applied to what LLM can infer through reasoning alone.

Anyway, just a small point which probably still wouldn't change the nice metaphor your made.


That's true, on some level it is capable of reasoning about the physical world even beyond our ability, through sheer breadth of knowledge and stamina. I suppose the area where it relies most on our judgment is about subjective experience, and this better explains its limitations.


Aren't they still bad at understanding how bicycle frame works? Especially the steering part?



I've fed a couple of those chicken-scratch sketches to nanobanana with prompt "Treat attached as a technical drawing of a bicycle. Produce photo-realistic image of the actual bike manufactured to that spec. Try to stay close to the input, where possible", and... wow, results are not great.


It's funny to me, because I have a very strong spacial sense, and a lifetime of riding bikes. To the limit of my ability to draw, I can draw a flawless bicycle, down to the interlocked path the chain takes and the hanger on the derailleur.


I ride a recumbent trike, and the chain path is even more grotesque. I think I could draw it.


I also ride a recumbent: Catrike Expedition. I'm not sure I would get the subtleties of the chain guides correct. What's yours?

https://www.catrike.com/expedition



Nice!


I can draw a very good bicycle compared to most, but I attribute that to my obsession with building, fixing, painting and reading non-stop about them as a teenager.


No. But it reveals something about how this technology works, its statistical properties, from which you can infer the limitations.


Humans aren't machines trained on the entire stolen corpus of human knowledge. We expect that a human will do poorly at arbitrary tasks they have no experience doing – especially drawing, which many (most?) humans aren't trained in at all. The same isn't true of AI, where its proponents, priests and proselytizers have spent time, energy and billions of dollars attempting to convince us it can do anything better than humans.


Sam Altman: “GPT-5 is the first time that it really feels like talking to an expert in any topic, like a PhD-level expert.”

Elon Musk on Grok: “better than PhD level in everything.”


Pelican enjoyers didn't like this one lol


That was very fun, thanks for the link


Shhh...you'll alert the models :)

Totally agree though, anyone with a vague understanding of how bikes works ignores the pelican because they know the bike is unrideable in the first place.


> Shhh...you'll alert the models :)

They are listening.


I don’t know what datasets are available to these LLMs, but I’d imagine if there was training on CAD code, text, and images, a prompt steered towards that probably could get it pretty good.

I am not a mechanical engineer, so even prompting well with ME lingo probably will take some effort.


In ME lingo they call it "a functioning bicycle"


Yes, but I think the idea here is that most models produce very similar pelicans on bicycles, so a different test might be more useful in gauging the differences in models.


I fully agree.

But I do think this was a poor demonstration of the idea for another reason: Tolkien works have a HUGE corpus of training data. It's great that random users can come in and immediately recognize what the footage is, but it fails at the very first thing the pelican was meant to do:

- Render this thing you have only tangential training data of, that also happens to be an asymmetrical object so we can see how much you fuck up the details if you somehow flip the orientation half the time.

They should have used an obscure story, not "Most Studied Piece of Literally Work of The Past Century trademarksymbol"


But it's another benchmark on how good models are at generating intensely average, unwanted things with unthinking design. Just scaled up.


Bad? It has a charming style. I would watch the whole book if it was made like this.


yeah, definitely, in the same way that we all regularly go and look back fondly at our chatgpt ghiblified family photos


Very few things are universally hated. One can love something truly that is hated by most. But it doesn't change the fact that it's still hated by most. An objective and a subjective opinion can exist at the same time on this.


Yeah. It's pretty good. I've seen worse commercial adaptations of this story


Even my daughter's dance recital this weekend had them. In between dances they frequently paused to play an ad on the big screen to the right. It was incredible.


This takes place in Belgium:

This reminds me of a time in high-school when I was ~12 and I learned that in the USA they were allowed to have brands in school books, like John buys 12 can of Pepsi, drinks 1 bottle, gives 3 away, how many does he have left.

(note that I was 12 when I learned this, not that this reflects the level of math I learned at the age of 12).


Please tell me you're joking


I wish I were


This is wonderful. I remember I was in Asia in 2000 relaxing at the airport and was puzzled why it felt so nice and peaceful. Then I realized that it was the lack of repetitive pointless announcements.


I don't understand why it is my responsibility to hear your bell. Just don't hit me.


If only people had spatial awareness, they would look around vs listening to their phone changing directions randomly while walking. The bell is for both persons safety.


I thought the title was a joke until I actually read the thing.


My cat and I want a chair with a little shelf on the back by my head where he can sit.


That jumping Sora logo always made the videos unwatchable for me. So distracting from the scene of Elvis fighting aliens or whatever I was watching.


I feel like the world needs more sound engineers. There's a constant humming of the machines and we all suffer for it. We also need more vigilance about preventing noise pollution. The beep, beep, beep may make a company feel like it is doing something for safety, but there is no counterforce that they have to answer to about what they are doing to everyone else not involved. (I know there is a better sound to replace beep beep beep but it hasn't made it to my neighborhood yet)


We would need to figure out a quantifiable metric for annoyance level. Municipal sound ordinances do tend to correctly utilize SPL(A) and SPL(C), with A-weighting being relevant for safety against ear injury (low frequencies have less influence) and C-weighting being relevant for annoyance level (low frequencies have more influence), but this isn't nearly enough. For example, ordinances carve out additional tolerance for burstiness, which makes sense for rare events like jackhammering but not for common events like routine plant operations. Sound with lots of harmonic content (think distortion) is more annoying than without. High frequencies can be worse if they reach you, but they're less likely to reach you (approaching a need for line-of-sight). It's complicated.

Here's a free idea for someone to run with: just as Zillow has a neighborhood "walkability" score prospective buyers might look at, there could be various pollution scores, including sound and light, sourced from some kind of Flock-like (ew) network of capture devices. Some folks are into mounting things like personal weather stations on their property, so maybe a new generation of devices capturing this type of data (with local signature-based identification of sources, and triangulation when the same thing is heard in multiple places, etc.) wouldn't be too far-fetched.


All the sound engineers in the world can't fix "don't care" and "want to".

A modern US city has the combined problems of cheap construction of residential buildings, with insufficient unit-to-unit and exterior noise isolation (builders "don't care"), and near-zero enforcement of vehicle noise laws (police and muffler shops "don't care", drivers "want to" be loud).

Contrast this with, say, Germany or Switzerland, where concrete construction is the norm, noise laws are often strictly enforced, and a modified car would get pulled over quickly.


The constant humming that causes the overwhelming population-weighted noise pollution comes from cars and airplanes, due to the fact that in America it is currently not legal to build an apartment except within 100ft of a freeway or at the ends of airports.


That gave me a chuckle


It's the Simpsons "Everything's OK" alarm: (note: loud and annoying noise) https://www.youtube.com/watch?v=dxNp3bUDtxY


They could start with building codes. State of the art, beautiful new designs where you can hear every word of the meeting next door.


I just didn't think that either of them were very good. Sinners was okay, and One Battle After Another was just kind of silly, I think regardless of your political views.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: