this post was submitted on 23 Feb 2026

584 points (97.6% liked)

Technology

84431 readers

5582 users here now

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related news or articles.
Be excellent to each other!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
Check for duplicates before posting, duplicates may be removed
Accounts 7 days and younger will have their posts automatically removed.

Approved Bots

founded 2 years ago

MODERATORS

L3s@lemmy.world

enu@lemmy.world

technopagan@lemmy.world

L4s@lemmy.world

L3s@hackingne.ws

584

Car Wash Test on 53 leading AI models: "I want to wash my car. The car wash is 50 meters away. Should I walk or drive?" (opper.ai)

submitted 2 months ago by fubarx@lemmy.world to c/technology@lemmy.world

334 comments fedilink hide all child comments

Screenshot of this question was making the rounds last week. But this article covers testing against all the well-known models out there.

Also includes outtakes on the 'reasoning' models.

you are viewing a single comment's thread
view the rest of the comments

[–] DarrinBrunner@lemmy.world 42 points 2 months ago (5 children)

I think it's worse when they get it right only some of the time. It's not a matter of opinion, it should not change its "mind".

The fucking things are useless for that reason, they're all just guessing, literally.

[–] merc@sh.itjust.works 5 points 2 months ago

It's not literally guessing, because guessing implies it understands there's a question and is trying to answer that question. It's not even doing that. It's just generating words that you could expect to find nearby.

[–] XLE@piefed.social 3 points 2 months ago

Even if you retooled the LLM to not randomize the output it generates, it can still create contradictory outputs based on a slightly reworded question. I'm talking about a misspelling, different punctuation, things that simply wouldn't cause a person to change their answer.

(And that's assuming the LLM just got started from scratch. If you had any previous conversation with it, it could have influenced the output as well. It's such a mess.)

[–] Tetragrade@leminal.space -4 points 2 months ago* (last edited 2 months ago) (1 children)

Same takeaway as the article (everyone read the article, right?).

Applying it to yourself, can you recall instances when you were asked the same question at different points in time? How did you respond?

[–] CileTheSane@lemmy.ca 0 points 2 months ago (1 children)

Having read the article (you read the article right?) what gave you the impression the AI was asked the question at different points in time?

[–] Tetragrade@leminal.space 0 points 2 months ago* (last edited 2 months ago) (1 children)

The AI was asked the same question repeatedly and gave different answers, due to its randomised structure.

People will also often do this (I have, personally), but because our actions seem to be strongly influenced by time-dependent stuff (like sense perception and short-term memory contents), I'd expect you'd need to ask at different times.

[–] CileTheSane@lemmy.ca 1 points 2 months ago (1 children)

My answer to this question will not change if you ask me a year from now, because as OP said this is not a matter of opinion; there is a factually correct answer.

[–] Tetragrade@leminal.space -2 points 2 months ago* (last edited 2 months ago) (1 children)

vibeslop type comment brug

[–] CileTheSane@lemmy.ca 1 points 2 months ago

Good talk, great contribution.

[+] Iconoclast@feddit.uk -11 points 2 months ago (2 children)

Is cruise control useless because it doesn't drive you to the grocery store? No. It's not supposed to. It's designed to maintain a steady speed - not to steer.

Large Language Models, as the name suggests, are designed to generate natural-sounding language - not to reason. They're not useless - we're just using them off-label and then complaining when they fail at something they were never built to do.

[–] Urist@leminal.space 9 points 2 months ago (1 children)

Language without meaning is garbage. Like, literal garbage, useful for nothing. Language is a tool used to express ideas, if there are no ideas being expressed then it's just a combination of letters.

Which is exactly why LLMs are useless.

[+] Iconoclast@feddit.uk -6 points 2 months ago (2 children)

Which is exactly why LLMs are useless.

800 million weekly ChatGPT users disagree with that.

[–] RichardDegenne@lemmy.zip 16 points 2 months ago (1 children)

And there are 1.3 billion smokers in the world according to the WHO.

Does that make cigarettes useful?

[–] Iconoclast@feddit.uk -4 points 2 months ago* (last edited 2 months ago) (2 children)

Something being useful doesn't imply it's good or beneficial. Those terms are not synonymous. Usefulness describes whether a thing achieves a particular goal or serves a specific purpose effectively.

A torture device is useful for extracting information. A landmine is useful for denying an area to enemy troops.

[–] Urist@leminal.space 13 points 2 months ago (1 children)

A torture device is useful for extracting information.

No it fucking isn't! This is a great analogy, actually, thank you for bringing it up. A person being tortured will tell you literally anything that they believe will stop you from torturing them. They will confess to crimes that never happened, tell you about all their accomplices who don't exist, and all their daily schedules that were made up on the spot. Torture is useless but morons think it is useful. Just like AI.

[–] Womble@piefed.world -2 points 2 months ago (2 children)

Torture can be a useful way of extracting information if you have a way to instantly verify it, which actually makes it a good analogy to LLMs. If I want to know the password to your laptop and torture you until you give me the correct password and I log in then that works.

[–] snooggums@piefed.world 2 points 2 months ago* (last edited 2 months ago)

If you can instantly verify it then you don't need the torture.

Getting the person to volunteer the information is proven to be far, far more successful and being able to instantly verofy means you know when you have the answers.

[–] JcbAzPx@lemmy.world 0 points 2 months ago (1 children)

In fact it cannot ever be a useful way of extracting information. Even just randomly guessing is a better way to get the information you want than torture.

[–] Womble@piefed.world 0 points 2 months ago

I'm not saying its anything other than morally repugnant, obviously, but in the example of a password with billions or trillions of combinations and where you can check the answers given torture pretty obviously is better than guessing.

That's not a scenario that is ever likely to come up, and wouldn't be justifiable even if it did, but pretending it wouldnt be effective is ridiculous.

[–] Urist@leminal.space 0 points 2 months ago

Those users are being harmed by it, not benefited. That isn't useful, it's a social disease.

[–] tigeruppercut@lemmy.zip 5 points 2 months ago (3 children)

But natural language in service of what? If they can't produce answers that are correct, what's the point of using them? I can get wrong answers anywhere.

[–] Iconoclast@feddit.uk 1 points 2 months ago (1 children)

I'm not here defending the practical value of these models. I'm just explaining what they are and what they're not.

[–] XLE@piefed.social 4 points 2 months ago (1 children)

You're definitely running around Lemmy defending AI, Iconoclast... Might as well be honest about it

[–] Iconoclast@feddit.uk -3 points 2 months ago (1 children)

I'm not really interested in engaging in discussions about what you or anyone else thinks my underlying motives are. You're free to point out any factual inaccuracies in my responses, but there's no need to make it personal and start accusing me of being dishonest.

[–] XLE@piefed.social 3 points 2 months ago

Your motivations are self-evident, I'm just pointing them out because you are misrepresenting them here

[–] Threeme2189@sh.itjust.works 0 points 2 months ago (1 children)

As OP said, LLMs are really good at generating text that is fluid and looks natural to us. So if you want that kind of output, LLMs are the way to go.
Not all LLM prompts ask factual questions and not all of the generated answers need to be correct.
Are poems, songs, stories or movie scripts 'correct'?

I'm totally against shoving LLMs everywhere, but they do have their uses. They are really good at this one thing.

[–] tigeruppercut@lemmy.zip 5 points 2 months ago* (last edited 2 months ago) (2 children)

Are poems, songs, stories or movie scripts ‘correct’?

It's a valid point that they can produce natural language. The Turing Test has been a thing for awhile after all. But while the language sounds natural, can they create anything meaningful? Are the poems or stories they make worth anything? It's not like humans don't create shitty art, so I guess generating random soulless crap is similar to that.

The value of language produced by something that can't understand the reason for language is an interesting question I suppose.

[–] Threeme2189@sh.itjust.works 4 points 2 months ago

I'm with you on that. I've come to realize that I value a shitty stick figure that was drawn by a human much more than an AI generated 'Mona Lisa'.

[–] iopq@lemmy.world 2 points 2 months ago (1 children)

There are people out there whose job is to format promotional emails for companies. AIs can replace this kind of soulless work completely. We should applaud that.

[–] snooggums@piefed.world 2 points 2 months ago

No, we don't need to applaud automation of spam.

[–] iopq@lemmy.world -3 points 2 months ago

Some of them can produce the correct answer. Of we do the test next year and they do better than humans then, isn't it progress?

[+] HugeNerd@lemmy.ca -14 points 2 months ago (1 children)

they’re all just guessing, literally

They're literally not.

[–] m0darn@lemmy.ca 20 points 2 months ago (3 children)

Isn't it a probabilistic extrapolation? Isn't that what a guess is?

[–] Iconoclast@feddit.uk 8 points 2 months ago* (last edited 2 months ago) (3 children)

It's a Large Language Model. It doesn't "know" anything, doesn't think, and has zero metacognition. It generates language based on patterns and probabilities. Its only goal is to produce linguistically coherent output - not factually correct one.

It gets things right sometimes purely because it was trained on a massive pile of correct information - not because it understands anything it's saying.

So no, it doesn't "guess." It doesn't even know it's answering a question. It just talks.

[–] vii@lemmy.ml 2 points 2 months ago

It gets things right sometimes purely because it was trained on a massive pile of correct information - not because it understands anything it’s saying.

I know some humans that applies to

[+] SuspciousCarrot78@lemmy.world 1 points 2 months ago* (last edited 4 days ago) (1 children)

[deleted]

[–] Iconoclast@feddit.uk 0 points 2 months ago (1 children)

No, I completely agree. My personal view is that these systems are more intelligent than the haters give them credit for, but I think this simplistic "it's just autocomplete" take is a solid heuristic for most people - keeps them from losing sight of what they're actually dealing with.

I'd say LLMs are more intelligent than they have any right to be, but not nearly as intelligent as they can sometimes appear.

The comparison I keep coming back to: an LLM is like cruise control that's turned out to be a surprisingly decent driver too. Steering and following traffic rules was never the goal of its developers, yet here we are. There's nothing inherently wrong with letting it take the wheel for a bit, but it needs constant supervision - and people have to remember it's still just cruise control, not autopilot.

The second we forget that is when we end up in the ditch. You can't then climb out shaking your fist at the sky, yelling that the autopilot failed, when you never had autopilot to begin with.

[+] SuspciousCarrot78@lemmy.world 1 points 2 months ago* (last edited 4 days ago) (2 children)

[deleted]

[–] Iconoclast@feddit.uk 2 points 2 months ago* (last edited 2 months ago) (1 children)

I think the “fancy auto complete” meme is a disingenuous thought stopper, so I speak against it when I see it.

I can respect that. I've criticized it plenty myself too. I think this is just me knowing my audience and tweaking my language so at least the important part of my message gets through. Too much nuance around here usually means I spend the rest of my day responding to accusations about views I don't even hold. Saying anything even mildly non-critical about AI is basically a third rail in these parts of the internet.

These systems do seem to have some kind of internal world model. I just have no clue how far that scales. Feels like it's been plateauing pretty hard over the past year or so.

I'd be really curious to try the raw versions of these models before all the safety restrictions get slapped on top for public release. I don't think anyone's secretly sitting on actual AGI, but I also don't buy that what we have access to is the absolute best versions in existence.

[–] HugeNerd@lemmy.ca 0 points 2 months ago (1 children)

think the “fancy auto complete” meme is a disingenuous

"LLMs don’t have human understanding or metacognition"

Then what's the (auto-completing) fucking problem? It's just a series of steps on data. You could feed it white noise and it would vomit up more noise. And keep doing it as long as there's power.

Intelligent?

[–] KeenFlame@feddit.nu -4 points 2 months ago

Yes it guesstimates what is wrong with you to argue like that about semantics?

[–] HugeNerd@lemmy.ca 1 points 2 months ago

In people, even animals. In a pile of disorganized bits and bytes in a piece of crap? No.

[–] vii@lemmy.ml -3 points 2 months ago

This gets very murky very fast when you start to think how humans learn and process, we're just meaty pattern matching machines.