r/artificial 14h ago

News Three hikers got rescued off a mountain this week after following Gemini's advice. The same week OpenAI launched what it's calling the AGI era. I keep thinking about both together.

The hikers story happened September 1st. Three guys from Roseville used Gemini to plan a Mount Shasta summit. The AI told them to bring far less food and water than they needed. They summited at 7pm, four hours after the recommended turnaround time, descended in the dark, one of them hurt his knee, and they spent the night stranded in a canyon until rangers found them the next morning.

Google says they can't replicate the bad answers Gemini gave. Maybe the prompts were vague. Maybe the AI was overconfident. Doesn't really matter which. What matters is that three people trusted a model's output as expert advice in a context where being wrong had serious consequences.

Two days later OpenAI launched GPT-6 Astra. 99.9% on ARC-AGI-3, 97.6% on FrontierMath Tier 4, 100% on ExploitBench. OpenAI is calling this the start of the AGI era. Independent benchmarks from Artificial Analysis are more cautious and show Anthropic's Fable 5.1 still ahead on the broader intelligence index.

But here's what I can't stop thinking about. The hikers story and the capability story are not separate things. Every time a model gets more capable, more people trust it in higher stakes situations. That gap between what the model can do and what the person using it understands about its limits doesn't close automatically when capability improves. If anything it gets harder to manage because the outputs get more convincing.

I work with organizations on AI adoption and the single most common thing I see is not people being too skeptical of AI. It's people not knowing when to stop trusting it.

What's your take? Does more capability make the trust calibration problem better or worse?

Upvotes

24 comments sorted by

u/Willing-Gazelle4591 13h ago

This isn't an AI problem. It's an idiot problem. We've had idiots forever. We've had overly confidant humans with AGI forever. Many of those humans have given out terrible over confidant advice since the beginning of time.

u/Bill-Maxwell 12h ago

Agreed, only idiots start from a place of trust with AI. It’ll get you part of the way there but if you think it’s infallible, well then I’ve got a virtual bridge to sell you.

u/curious_astronauts 12h ago

Its the oracle dilemma. People who outsource all their thinking, then treat LLMs like an oracle, are susceptible to the consequences of that blind faith.

u/DeliciousArcher8704 6h ago

We've had overly confidant humans with AGI forever.

We certainly have not..

u/ygg_studios 12h ago

that is the problem, most people are stupid. that's why you don't unleash a powerful tool on a public that lacks the critical thinking skills to use it appropriately. hope this helps!

u/Double_Suggestion385 13h ago

The thing is, even without AI, people routinely take on hikes without sufficient resources or equipment.

I wouldn't be surprised if those guys are just larping for some kind of anti-AI stunt. Can anyone replicate the answers they supposedly were given?

u/justin107d 13h ago edited 13h ago

They are getting absolutely roasted over on r/mountaineering. They clearly had no idea what they were doing. I think the simpler explanation is that they were trying to blame AI to save face. It is kind of impressive the way they managed to butcher, what appears to be, an overnight hike so bad.

Edit: to answer OP's question, terrifyingly the same could be applied to people. Currently you have a junior analyst do the work, a senior review, and maybe a third review if absolutely critical. Their mistakes will become more and more human like.

u/Colorful_Monk_3467 12h ago

Article says it's 8 miles and 6500' of gain. That's a long day but not an overnight. I'd attribute it mostly to poor fitness/inexperience.

u/justin107d 12h ago

I saw that they reached the peak at 7 pm and assumed that it was an overnight hike from the start. The article is way worse.

SF Gate article

u/ygg_studios 12h ago

cope harder

u/Gr8deadon 13h ago

This whole story has morphed from what actually happened. I love close by to this. They used the alltrails app. That is powered by a.i. they found out that they were not good enough to do the hike too late. They got stuck on the way back down cause the dude using the app phone died. So when they got stuck they couldn't look at their navigation.

u/GrizzyGramBag 13h ago

They are dumb for trusting Gemini. Good for creative story telling, bad for anything else. That being said, it's like asking Reddit, alot of people take some pretty bad advice from randoms

u/Few_Worker4102 12h ago

This is the Dunning-Kruger Effect. Most people don’t have the ability to see they’re out of their depth and cannot make accurate judgments about a topic. 

In other words, AI tells you “do X” and you can’t tell if that is reasonable or not…but you think you can. 

u/presentofai 12h ago

honestly this is an idiot problem cosplaying as an AI problem. anyone summiting at 7pm was gonna find a way to get stranded with or without gemini

u/PeakNader 11h ago

I see lots of doomer posts in this sub where the op ghosts the comments… odd

u/InventoryOfCheese 10h ago

"hey i know, let's say AI planned our trip so we can get out of paying any rescue fees" 

u/AkindaGood_programer 10h ago

While I do agree that, with proper research, the hikers would've never needed to be rescued, there were, and always will be, dumb people. Do you really think that if Gemini didn't exist, they'd be doing a lot of research on the subject? I'd claim no.

u/Nangangbangbang 1h ago

Did you also post this to your LinkedIn?

u/EightyNineMillion 1h ago

Idiots will be idiots. Highly recommend "I shouldn't be alive" on YouTube if you want to see more idiots being idiots. AI is not the problem. People who don't know how to use AI are the problem. It's no different than performing a manual web search and believing everything you read from the top result.

u/ThomasToIndia 40m ago

The problem is that humans can be wrong too, and even experts can miscalculate. I think I would argue in many cases on certain topics the failure rate is lower than a human.

So you can't take one story like this without taking in all the other stories where AI was right and the human was wrong. Focusing on one story is whataboutism.

Just think about malpractice lawsuits for people.

u/jbano 1m ago

Don't plan a life threatening trip with a flash model on Google. Choose a better frontier model

u/katoptronophile 13h ago

AGI is here.

Some people are just in denial.

That's OK, let them cope.

The future begins now!

u/DoorPsychological833 7h ago

Stupid people follow ASI over the cliff like lemmings. Thankfully we have our the smartest, most righteous people in power positions

u/Lordofderp33 5h ago

That is clearly not you.