I feel like we read different articles, because without context, the pictured dish looked like overcooked pork chops served next to a type of carbonara tagliatelle topped with a fried egg & chives. I can't imagine anyone has ever set out to make spamsilog and ended up with the generated dish.
It's different in kind to a "my Big Mac is more squashed than the advert on TV" type situation.
>But is that a widespread feeling, are the vibes bad for regular people, or is this effectively just old man shouting at clouds?
A cursory glance at YouGov strongly suggests that people have bad vibes about AI in basically every territory where they're asked about it, and for basically every use-case
This touches on the heart of it. AI generates bullshit. By definition. Bullshit is something that is produced without any care for whether it is true or authentic. The bullshitter's only concern is its value to the bullshitter. AI has no notion of truth or honesty. It's all just prompted, predictive streams of output. When we recognize it, we have the same reaction as we do to human-produced bullshit. Maybe even a little worse, because human bullshit at least may have some evidence of humanity in it. AI bullshit doesn't even have that.
I agree with others saying that this doesn't disqualify complaints.
In fact, it can be precisely because they can't tell the difference that they would be concerned about consuming AI outputs when they don't know they are. If someone wants to take a stand and only read human text and human art then we may need regulation to add notices when text and images are LLM outputs.
I think that many people are against toxic pesticides in food too, or meat that was inhumanely slaughtered, but there are few signals in the packaging to tell how things are made (besides labels that are dubiously trustworthy). So people might be against everything LLMs represents (bad for local and global environment, made with slave-labour, made by stealing material, removing jobs), without being able to tell something is LLM-generated. However, when they can tell, the gut reaction is disgust and the brand takes a big hit. I know when I notice something as LLM-generated, which I know I don't always do, I'm immediately disgusted by it and avoid it if I can.
Almost none, sadly. Most of the population just doesn't have that skill, which is why places like Facebook are just reams of AI Jesus upvoted to the Moon.
I saw this menu online the other day, and it's not even discernable from the photo what the food products are; it even sets off trypophobia:
I agree. For most people I've talked to, when this topic comes up, they'll tell me they don't like AI, but then cannot really identify it. Especially the elderly folks in my neighborhood. You look at their facebook feeds and it's 100% slop. No joke. Every. Single. Post. Is. AI. Slop. And for the most part they can't tell.
The problem is in 2026 I work in this space and it's getting to the point where I can no longer tell sometimes. By next year we're probably all cooked. My first go-to now on an image is TinEye.. if it first appeared before AI then I know it's safe.
Most people I encounter who complain about AI use it as a shorthand for 'anything I don't like or disagree with'. When asked why they think something was created using genAI, they cannot point out reasons.
What's the argument here? People don't like spit in their food, how many people can tell their food got spat in? Does that mean they're somehow wrong or to be belittled for not liking spit in their food?
It's not an argument, it's an open question: do people who express a dislike of AI know when they see a poster that was generated with assistance from AI?
I meant to ask what the point of the question is basically. What would change, what would it matter for, if the answer was yes or no?
Besides it being so general a question to really be useless as anything but an insinuation IMO, which is why I'm trying to tease that insinuation out. E.g. are people who use Apple left-handed? Are they gay? The answer is "yes, some of them, of course, but also many aren't... why are you asking?"
There is no thing as "just a question, nothing in mind really". That does not exist. The mind that can ask a question cannot arrive at it in a vacuum.
OK, my point here is that a YouGov poll about people's attitude to AI may not tell us anything actionable about how people feel about AI posters - it depends on how good they are at detecting them.
They're still in the stage of trying to get uptake on the sub. Ads come later. When they've got sufficient uptake, they'll start to add different sub levels.
In the UK, farm vehicles use a specially dyed Red Diesel, which is the same liquid with a chemical added to it to turn it red. It’s taxed at a much lower rate, so you’ll sometimes pass petrol stations selling two Diesel from two different pumps, one of which is around half the price of the other.
If you’re caught running red diesel on a normal carriageway, the penalty can be steep.
In Germany that red stuff is applied to all diesel, not only agricultural use.
And it's not the same liquid. From first principles maybe, but not after it has gone through the supply chain until it reaches you.
Anecdotical evidence from decades ago, when I had an oil-fired oven for heating.
With a big barrel in the cellar, which you let fill by hose from a truck maybe once a year.
And hand pumped into a can every few days, and carried that up, to fill the tank of the oven. One year I was rather broke, and that barrel empty. And then came winter.
Thought I'd be smart, because it's just the 'same stuff', right?
Walked to the Aral maybe 200 meters away, filled a canister.
That 'same stuff' made the fire drum clog from the inside with the finest carbon, like those mushrooms which grow like disks from the side of trees.
That never happened with just oil. And that was normal Aral diesel, nothing special.
One can also notice a difference in smell with older cars. Just oil stinks more.
Respectfully, the pelicans used to be an unrecognisable mess and now they’re unquestionably pelicans on bicycles, rendered poorly, from every model.
In the same timescale, model capabilities across the board have only meaningfully improved in places where the labs are focusing their training efforts.
Moreover, they have a uniform style, even though your prompt doesn’t ask for one. There's no model going rogue and producing a watercolour of a pelican. They’re all rendered in an approximately uniform style, even though the svg format has a basically unlimited possibility space.
You know what, that's actually something I hadn't considered before. There's definitely a bias towards a pelican cycling from left to right on a red bicycle against a blue sky and green grass.
Blue sky and green grass aren't that surprising, but the color and direction are interesting.
When I finally build the proper gallery I'll throw in a few other creature-vehicle combinations, and track some characteristics like which direction, color of bicycle, general pelican geometry etc. It will be interesting to see if other creatures end up with coincidentally similar design choices or if that's unique to the pelican-bicycle combination.
The other thing to consider (as someone who frequently take a photos of their bike) the common direction has the drive side out! In cycling forums it is sacrilegious to post a photo of your bicycle without showing the drive side.
Took some searching and sleuthing to actually figure out what "drive side out" means, as I'm just a casual "from A to B" cyclist: apparently this is referring to the side the chainset, chainwheels and all those things are on.
Beat me to it - but I had the same thought. Most amateur and nearly all professional studio photographs of a bicycle will have it drive side out so I expect this plays some role in it.
It is. All over the Arab world, imagery in ads is “backwards” and I believe several companies will flip their ads horizontally, and UI localization involves flipping graphics.
Do the gears and stuff sit on the other side of the bike in the Arab world? Otherwise I'd expect cycling ads to still show a bike going from left to the right, considering https://news.ycombinator.com/item?id=48951828
That was my first thought too, I wonder if it works the same in countries speaking arabic (as that's the first one i could think of that's a language with truly no-buts right to left writing).
Yes, people will usually post or draw a bicycle right to left which is going to ve opposite of what normally is drawn. I tried the prompt in arabic for many models and I don't recall any adjusting it based on that difference at least culturally speaking.
There was a glorious moment when I thought that the Chinese models were more likely to produce right-to-left cycling pelicans, but sadly that trend didn't seem to hold up.
What's interesting is that given the fairly general and short in length prompt for the test, none of the models are attempting things like more discrete details of the bike. Such as showing V-brakes or dual 160mm disc rotors, rear derailleur, water bottle in a bottle cage, panniers, lights, saddlebag, the rider wearing a helmet, or other details that might be found on as vague a description as "a bicycle".
The models are already brilliant at that. My own todo app generates 128x128 pixel art icons for my todo items. They are mind blowingly creative and funny.
There's a bias in the direction all things face. You can ask these models to generate a thing animal, car etc and you will notice that 90% of them will converge towards the same sort of results. If you ask for something rotating, 90% of them will rotate right and a few odd ones will rotate left.
I have done some variation of the other animals, also for something more tricky where they need to calculate things, I ask them to draw an SVG at a certain angle.
For example: "generate an SVG of a chessboard seen from a 45 degree angle slightly higher POV" or "generate an SVG of a basketball court from a TV broadcast perspective".
I haven't seen many AI works that produces a pelican on a bicycle done in a "Ligne Claire" style, for example.
I guess AI's narrows down the output probability space drastically and converge on some agreed upon aesthetics. Works great for computer programs but bad for art.
Bicycle color, grass color and sky color are all part of the prompt.
>Cartoon illustration of a white pelican wearing a red scarf, riding a red bicycle along a gray road with white dashed lines; the pelican has a large orange beak and webbed orange feet pedaling, with white motion lines behind it; the background shows a light blue sky with white clouds, a yellow sun, two small black birds in flight, and green grass with tiny white flowers in the foreground
That wasn't the prompt. That text was generated by asking the model to describe an image and feeding it a rendering of the SVG it had previously generated.
> the pelicans used to be an unrecognisable mess and now they’re unquestionably pelicans on bicycles, rendered poorly, from every model
You would not expect that to happen if the models trained on the unrecognizable mess, right?
> model capabilities across the board have only meaningfully improved in places where the labs are focusing their training efforts
And the labs clearly did focus on improving image rendering.
> they have a uniform style
SVG output from LLMs always looks like that. It looked that way from the beginning; no LLM ever produced a watercolor when asked for SVG output. They all render the prompted element centered in the picture. They all tend to draw things going from left to right, and so on.
I’m not suggesting Simon’s pelicans in the dataset are having a meaningful impact. I’m expecting that a company like ScaleAI has a product along the lines of “benchmax dataset: SimonW’s Pelican on Bikes test” which is a private curated series of well-drawn SVGs of animals riding vehicles for training and RL.
If they're benchmaxed on SVG pelicans then the outcome of that has still produced a surprisingly good generic SVG image generator.
Go invent your own random alternatives and the AI models have across the board gotten better over time. Insects playing sports, anthropomorphic fruits performing martial arts, wizards conjuring weapons of WWII, whatever you can imagine. I've tried a lot of these, well beyond what I think would be a reasonable thing to specifically train as combinations. If they have given it a corpus of SVG drawings it has learned to extrapolate.
(note: wizards conjuring a tank got me a surprise animated SVG with my Qwen 3.6 35B model)
If you’ve been keeping track of all of the pelicans, there is actually stylistic differences - sometimes pretty big differences as far as watercolors go. It’s an SVG so I’m not sure what you’re looking for there. Most look the same because the prompt is to make a pelican on a bicycle as an SVG. It’s not some giant image prompt.
> Moreover, they have a uniform style, even though your prompt doesn’t ask for one.
This shouldn’t really come as a surprise, particularly to anyone who’s used diffusion models. The same thing happens when you ask an LLM for a short story [1] without providing any specific details.
Even cranking up the temperature or top_p values is no panacea. The more generic your prompt, the more pedestrian the response.
> model capabilities across the board have only meaningfully improved in places where the labs are focusing their training efforts
That doesn't seem right. I use these models as research assistants when writing lots of random blog posts (including in economically ~useless areas like the history of contra dance) and Fable 5 is a serious improvement (when I don't get downgraded!) over Opus 4.6-4.8 which was a serious improvement over Opus 4.
being able to draw a picture of a pelican is really cool and it requires intelligence but i don't think it's a good measure of improving capabilities of these models nor AGI. we don't have to spend so much breath on it.
Also, I'd assume the ideal output for an underspecified, generic prompt is the most expected, generic result. Not something that defaults off the rails with creative license.
In part because model vendors specifically prefer when people think that lots of content is produced by their model. The more Claude-like writing appears on the internet, the more signal there is to investors that people are using Claude for a greater number tasks.
Decent summary of it here[0]. The “space” part of “SpaceX” is valued by market analysts and money managers at around 5% of the company’s entire value. Almost all of the rest is “AI stuff”, and Twitter is a rounding error.
That is, if SpaceX went back to being a space-only entity, and dropped the AI stuff, its share price should be expected to fall from $130/share to around $7/share.
> to demonstrate that the tech industry isn't just here to extract wealth from the poor/many and transfer it to rich/few?
I think the problem is that the tech industry in large is just here to extract wealth from the many and transfer it to the few. That's why it's focused on scale.
People aren't dumb, and most of the time they can see when they're on the receiving end of an extractive relationship - even if there's lots of PR work going on to hide that reality from them.
It's unlikely that AI will get to the point where it makes handwritten coders redundant, and then not immediately be at the point where vibe coders are redundant too. So if you earnestly take the position that handwriting code is a "ngmi" type activity, you also need to take the position that the vibe coder (or agent- assisted-developer/loop-architect, or whatever its nom de guerre is this week) is "ngmi".
> I can do projects in 3 days what would take 6 months.
The hyperbole on this keeps growing every time I see it. Soon we’ll be having people claiming they can do in 12 seconds what used to take them 17 years. What is never presented is proof. People (and programmers are no exception) are notoriously bad at estimating. We already did studies where people thought they were being faster with LLMs when they were in fact being slower.
As companies begin to rehire to fix the mess made by LLMs, it’s clear that just getting something out the door isn’t enough. It never was. Maintenance is an important part of any long-standing system.
That one does sound like hyperbole, or maybe he just works slowly as a human? People are different.
Likewise, I think they're having wildly different results. Look at how differently humans drive vehicles, and realize they're doing the same with compute. Some people probably are working at light speed, and some people are actually slower like in the study.
> Look at how differently humans drive vehicles, and realize they're doing the same with compute.
I’m not sure that comparison is evoking the image you intended. Hands down the best driver I know—the one I’m sure won’t get me car sick, won’t ever have me worrying for my safety, the most fuel efficient, the smoothest rider—is by no means the fastest but the most thoughtful and methodical.
I don’t care how fast you develop your software. Is it good? Is it carefully considered? Will it not bite me in the ass? Those are the things that matter.
>4x improvement on geospatial tasks with map in the loop.
The graph shows a baseline 2% task success rate improving to to 8% task success rate, but the evals section details 100% success rates across the board.
I'm not sure what the effectiveness of this skill is from the readme. Is it 8% success, or 100% success?
It's different in kind to a "my Big Mac is more squashed than the advert on TV" type situation.
reply