r/BetterOffline • u/r77anderson • 5h ago
r/BetterOffline • u/ezitron • Jun 23 '26
Ragebait Videos/Clips/Pieces Will Now Get Removed
Hey all! As part of the ongoing success of the show, it appears that a coterie of people have started making videos with the intent of using my name to get clout/traffic/views. Please do not engage with or share these pieces! They exist entirely to piss you off and get you to post them here so they can siphon off traffic.
These posts are not a violation of any given rule and won't get you banned, I get that many of you want to fight for my honor! But I also want to make sure that we don't fall for obvious trolls. It's far funnier watching people get in a tizzy for no reason.
r/BetterOffline • u/ezitron • Jun 10 '26
Low Effort Posts Now Get A 7 Day Ban
Hi all,
I hate to do this, but people - including users who have been here for over a year - seem to not be taking the low effort post rule seriously, even when I remove 3 to 5 of their posts in the space of a month. As a result, any and all low effort posts will now get a 7 day ban. I didn't want to do this, but it's become apparent that people don't read the rules, or the pinned threads, so I'm going to have to get serious. I really do not want this place to turn into a selection of links and single-line posts or web comics. Please read the rules.
r/BetterOffline • u/creaturefeature16 • 21h ago
I think Astra is the industry's equivalent of the “Mission Accomplished” banner from the Iraq War.
IMHO, It's a coordinated effort between OAI and NVIDIA to get past the term entirely. They need to diminish the definition of it, and then stake their claim on it.
This is also why Altman recently been disparaging of it, calling AGI an "irrelevant marketing term". It's become albatross around their necks (even though it is OpenAI's reason for existing), one that has become quite clear that they will certainly not reach with LLMs, so they just declare it has finally happened, and try to never bring it up again.
Also good to note that Dario isn't saying anything of the sort; dig into the contention between Dario Amodei and Jensen Huang to see Jensen's OTHER motivation here.
As Gary Marcus brilliantly said: if/when AGI arrives, "you won't have to squint to see it".
r/BetterOffline • u/Disastrous_Room_927 • 16h ago
Y’all ever notice how well-established theories of intelligence never come up in discussions on AI/AGI/ASI?
123 years ago, Charles Spearman published a paper titled “General Intelligence: Objectively Determined and Measures”. This paper basically formalized the loose concept of intelligence so that it could be studied empirically. Modern theories of intelligence grew out of this, as did the popular notion of it we have today.
It doesn’t just bug me that the people working in the AI “field” seem completely unaware that over a century worth if research on intelligence is out there, it bugs me that there’s seemingly no interest in giving the concept of AI an empirical treatment.
I can’t say I’d expect anything else from the tech industry, it’s just incredibly problematic that some half baked notion of intelligence is being used to justify shoving LLMs into everything. Or that the people who don’t see an issue with this have the world’s economy in a stranglehold.
r/BetterOffline • u/No_Practice_745 • 1d ago
OpenAI’s latest math breakthroughs commit research misconduct, experts say
I don’t have much to add in commentary, quite frankly most of the math and concepts are beyond my arts-degree brain. However, the article does a great job of showing, once again, how disingenuous OpenAI and other AI labs are in describing what their products are doing.
They basically want the headlines that their new models are “solving math,” but what they’re doing is plagiarizing other peoples’ research and claiming things have been moved forward. There is clearly a use case here for researchers in using an LLM to catalogue and evaluate large swaths of data from over periods of time, but these things don’t think, they aren’t creating anything, and OpenAI doesn’t give a shit as long as people see the headline and bow their heads to their new scI-fi god.
r/BetterOffline • u/Icy-Recognition-7453 • 1d ago
Cal Newport exposes the Rationalists/EAs in the New York Times
Key takeaway
"It helps to ask how things would have been different if generative A.I. had emerged from companies without connections to Rationalism. We can glimpse this counterfactual in the example of China, whose booming A.I. industry is independent of Silicon Valley futurist movements. As Ross Douthat recently explained, whereas the United States is “behaving as though frontier A.I. models could be the equivalent of nuclear weapons,” the Chinese see A.I. as a “lower-risk technology to be shared and commercialized to win friends and influence the world.”
We can find a similar moderation in the American A.I. leaders who have minimal connections to Rationalism. Consider Nvidia’s chief executive, Jensen Huang, who co-founded the company years before Mr. Yudkowsky became an Extropian, and never seemed to have mixed much with Rationalist circles. His frustration with the more extreme talk from figures like Mr. Altman and Mr. Amodei is often palpable, as when, during a podcast interview, he made the following plea: “I appreciate that many of us grew up and enjoyed science fiction, but it’s not helpful. It’s not helpful to people. It’s not helpful to the industry. It’s not helpful to society. It’s not helpful to the governments.”"
We should all be thankful to Cal Newport (friend of the show) for this article. He is one of the good guys.
Disclosure: I subscribe to his newsletter and you should too.
r/BetterOffline • u/Otterz4Life • 1d ago
Another banger from Gamers Nexus. Your smart tv is spying on you.
You thought smart tvs spying on you was a conspiracy theory? Think again.
Do they even sell dumb tvs anymore?
r/BetterOffline • u/Some-Ad7901 • 1d ago
Should I care about this Astra thing? What does it do differently?
I've been staying offline -and am better for it, pun intended- for a while now. It's been amazing for my health, but every once in a while, the Nerd-Reich class manages to create something so putrid as to cross into my bubble and disturb the peace.
Clammy Sammy has managed to return with this GPT 6 Astra thing, so I hopped onto twitter (big mistake), and I've noticed there is a lot less hype online for it than has historically been the case for these things, I think the luster of AI has almost worn off (and it cannot come sooner).
However, and correct if I'm wrong, I've seen certain highly impressive demos of the thing hijacking people's computers and generating paintings using the mouse, or sculptures in Blender and Nomad sculpt (there are also many fake demos stolen from talented artists' timelapses).
I'd be lying if I say this didn't freak me out a bit. How does this work any different? How does it know how to make that shitty bat that it made? It's nothing special if a human does it, and I can make it in about 12 minutes, but it freaked me out that a machine can seemingly navigate a complex UI and use features like material nodes and rigging. I may be wrong here, but it seems much more capable than LLMs have been historically at interacting with UI's like a human does, is this capability new?
r/BetterOffline • u/Aphid_red • 1d ago
Even the latest thinking models get things confidently wrong
Note: In a way that no honest human ever would (because you'd get caught instantly by anyone bothering to check)
Sometimes I've read things that practically amount to 'hallucinations are passé', but this is still not the case.
Perhaps a very expensive model with an advanced prompt could have caught this, or perhaps random chance could get you lucky. (Or maybe there's errors in proof frameworks or whatnot)
Either way, you still can't trust it with the unverifiable.
Background: I found a hidden base64-encoded message in Korean in a game I was playing (some lore hidden by the developers). Intrigued, I tried forwarding a screenshot to AI to decode it. Then I did the same manually.
I tried decoding the text copy the AI agent (Gemini) helpfully provided as an in-between step -- I instructed it to be careful and provide details, by giving it the plan of action; "I have reason to believe this to be base64. OCR first, then translate."
A couple minutes later...
It was invalid UTF-8. Yet the model went and "translated" that anyway, by making up an entirely different output that only matched the first few characters, because that base-64 didn't include the symbols i, I, l, and 1 (the 'confusables' in most fonts). Carefully inspecting the image myself under 4x magnification and re-sharpening I could tell these apart very cleanly, properly OCR it, then feed it into a translator.
Still not sure about the actual message (I'm trusting an AI translation here after all); I wished they'd included an English version of it and showed a different image depending on the language you set the game to, a little translation oversight.
r/BetterOffline • u/brevenbreven • 1d ago
There are too many ai boosters treating this subreddit as a testing ground for bs
Every week there is another 'moderate' ai fan who just wants to point out a small error made by Ed or there is a 'big' change in the latest ai update. I find it disappointing that thos place gets used as a workshop for bad faith Ai arguments. This is not all the time but its usually very unproductive and uses a lot of lawyer words to get their point across.
i find this insincere engagement with Ed's content really lousey. I dont think this is being written by Ai directly but it feels lame. If no one else feels this is happening on the subreddit, ill take a step back admit my mistake and take a break.
<EDIT: I want to thank the comment section for sharing that folk are experiencing a lot of bad faith arguments, I really appreciated that people spoke from the heart. As well as the other people in the comments who just kind of proved that there are a lot of boosters in the subreddit>
r/BetterOffline • u/HilarityJester • 1d ago
Where are all the optimizations ? The OS rewrites ?
There's is constant appraisal for LLMs when it comes to software development, to the point where people are claiming that the ai is writing all the code. If that were even remotely true, we would be getting complete software rewrites in the blink of an eye. If I was a software company, the very first task I would assign AI is to optimize the current software.
Do exactly the same thing the current code is doing, but use less ram, make it faster, less disk writes, less network queries, smaller file size and so on. I would also immediately rewrite my software so that it is compatible with Windows, Mac, Linux, so that my market is bigger overnight.
Where are the news about software or games now being highly optimized ? Repeating the exact same task, but better should be a trivial task for "intelligence".
r/BetterOffline • u/maccodemonkey • 1d ago
“Cory Doctorow on the Big AI Lie”
Nice interview with Cory in a forum that’s somewhat skeptical of the bear case. Ed gets some mentions.
r/BetterOffline • u/aaron11144 • 1d ago
Little bit annoyed about Anthropic
A lot of people seem to be saying things like, “Claude is the best for coding,” “Anthropic is the ethical AI company,” “Anthropic is the one whose IPO will succeed,” and “Anthropic is the one that will survive.” They also portray Anthropic as being far more ethical than OpenAI. I find this narrative pretty annoying because, at the end of the day, Claude is also just an LLM. I’ve used multiple LLMs, and honestly, they all feel broadly similar to me, even for coding. Capabilities can also be copied, distilled, or reproduced across models, so I don’t really understand why Claude is being treated as uniquely special or uniquely ethical. And then you have Dario, who is arguably one of the biggest fearmongers in the AI industry. So I’m not really sure where this perception of Claude as the uniquely “ethical” AI company comes from—it feels more like a narrative people have collectively bought into than something that is obvious from the actual technology or the company.
r/BetterOffline • u/usefulservant03 • 1d ago
The "but you're using the trash unpaid model, what did you expect?" nonsense.
I've never paid for an LLM's "better self", whatever that even means. When one inevitably disappoints me and I use it to remind the AI bros that they're living in a colossal lie that only diminishes and makes a mockery of its own gullible supporters, they frequently say things like "oh you were on the unpaid trash plan, what did you expect?? The paid one is the smart thinking one!".
There are several glaring issues with this that help expose the LLM bubble's massive lies.
First of all, why would I pay for something that, even during the "free trial", can't demonstrate to be able to reliably and consistently perform well even for simple tasks? Instead relies on a big fat "trust me bro" pop up to get me to pay, despite nobody having given me even the slightest guarantee that it will give me correct output? This argument alone gets orders of magnitude harder to refute when you remember it's is being marketed as massively paradigm-shifting tech that is supposedly gonna change the world forever?
Furthermore, what exactly does "the better model that you gotta pay to use" even mean? I ask this because it feels like even the companies pushing that stinky garbage in our faces seem to not know what exactly that means. When they temporarily lock you out of the free plan on Claude's website or ChatGPT's website, etc, you get a pop up that lets you know. Problem is, I've noticed them completely change what that pop up says multiple times over the years and it's always something extremely vague. First it was "to unlock higher limits". Then it was "Your time with our strongest model is up, you can use this dumber one for the time being, or pay to keep using the strongest one". Then it was "You're on the free plan, as you use it more, its memory will weaken and it may start getting details wrong" (saw that one yesterday on Gemini iirc). Can anyone even explain to me what any of these deliberately misleading versions of "it will get smarter if you pay for it" mean??
Lastly, it's the lack of warranty. This one is huge. When I go buy a new car or dishwasher, it has a warranty period of several years, during which the seller GUARANTEES to me that what I've paid for will serve me well, or I get my money back, or repair it for free, etc. When exactly do you guys remember seeing any of the LLM providers give such a guarantee for the paid plans? When do you remember anybody saying "I GUARANTEE that it will give you correct output, OR your money back"? Cuz I don't remember seeing any of them do this, ever. All they do is say vague things like what I already mentioned above and that's just another one of the many bad signs. Because, you know what? Even the LLM providers THEMSELVES have no idea how any of this trash works!
This is so obviously a gargantuan con and lie, I feel sorry for anybody who still sees any credulity whatsoever in the narrative that LLMs are somehow ( any day now...? ) gonna change the world.
r/BetterOffline • u/JustDoIt52 • 1d ago
Hikers got stuck after taking Gemini's advice
This shit is so funny and sad at the same time. There are people who already have done these trails and their information is available online. Still people would go for the artificial imitating machine to get this info. I don't put all of the blame on these guys instead I think it's fault of the clueless CEOs. Dennis Hannabis is going to come out on stage, spew put some AGI bullshit meanwhile his own AI gives blatant false information.
r/BetterOffline • u/falken_1983 • 1d ago
Prozac and AI Agents - HyperNormalisation (2016)
Enable HLS to view with audio, or disable this notification
What Eliza showed was that in an age of individualism, what made people feel secure was having themselves reflected back to them, just like in a mirror.
Ed referenced HyperNormalisation in a recently news letter, and there was a thread about the doc as a whole, but after re-watching it, this particular scene jumped out at me as it seems to be describing something close to AI psychosis and it doesn't even present it as a new thing.
r/BetterOffline • u/kekllkek • 1d ago
AI is getting better at SWE according to benchmarks. Haters would call them laughable.
DeepSWE is a prominent benchmark that measures coding agents’ ability to solve the corpus of 113 human-authored realistic tasks in the real projects’ codebases.
The benchmark prides itself on the following points which dumb haters would argue (I’ll play the hater):
Contamination free: Tasks are written from scratch, not adapted from existing commits or PRs, so no model has seen the solution during pretraining.
This benchmark was published in May. Every model released since could ingest both tasks and solutions. This barely if even includes Opus 4.8 and later, so “no model has seen” should’ve been changed to “no model had seen” the moment after publication.
High diversity: Tasks span a broad pool of 91 repositories across 5 languages.
The tasks are all very technical and well-defined. None of them are solutions to business problems how a product manager or a customer could describe them. All of the target codebases belong to the projects that are, as Opus 5 would call it, load-bearing. They’re not end-user products.
Real-world complexity: Prompts are ~half the length of SWE-bench Pro's, yet solutions require 5.5x more code and ~2x more output tokens.
I implore everyone, including people not familiar with coding, to look at one of the tasks. Would you say it’s trivial to describe a solution with this level of fidelity? Would it take anyone a couple of minutes or does it require pre-existing knowledge and research?
Most engineers wouldn’t formulate the task this clearly. They’d hand-code the solution or hand-off to an AI a half-assed prompt that’s nowhere near as detailed.
“Real-world complexity” is when you have to figure out the solution or even what the problem is, not when you already have a clear definition of what should be done.
So in the end of the day frontier models successfully solve this benchmark’s tasks. Not all of them every time, but practically speaking they do. Inspired, I decided to come up with my own benchmark to push DeepSWE’s mission even further. Example task and solution:
- Task: Do a massive number one followed by a massive number two.
- Success criteria: toilet is clogged.
- Solution: PISS AND SHIT.
You can run this benchmark at home, and you don’t even need LLMs. See if you can beat Fable.
Let’s move on to the next benchmark. SWE Atlas - Codebase QnA measures LLM’s ability to answer questions about the behaviour of complex codebases.
As the benchmark’s page describes target codebases:
They are also contamination-resistant, using strong copyleft licenses (e.g., GPL).
That’s a huge relief. We all know how AI labs hold intellectual property laws in the highest regard. There’s no possible way that models could’ve ingested the projects in question and still remained closed-source. They might’ve scraped all other protected works like books in violation of their licenses by claiming “fair use”, but they would never do it for code.
Now take a wild guess: have all tasks and solutions for this benchmark been available publicly since February? The answer is yes. Could it be since that time models have been trained on them? The answer is no, as the models are getting so intelligent that benchmaxxing is not needed and doesn’t happen. Surely.
Conclusion
It cannot be denied that models are getting better. Just yesterday Fable helped me draft a divorce settlement with my wife (soon an ex-wife). I’ve had enough with her complaining about me “not being physically and emotionally available” and other junk. I gave up on trying to explain that I had to tokenmaxx not to be left behind, and she was in a better position to do house chores and shit after her 9-5 job while I catch that sweet Claude quota reset to get some proompting in.
I hope when she leaves and takes the kids, I’ll have enough spare time to finish my own benchmark. I feel like the existing ones are not complimentary enough, and I hope to create the one that every model would score a hundred at. That’s when we’ll get to AGI and replace all jobs, remember my words.
r/BetterOffline • u/makersfark • 1d ago
RE: The Diary of a CEO question
There was a point in the interview where the grifter kept insisting that the process isn't important and that it didn't matter because the "output" was the only thing that was needed at the end of the day.
Besides being just a terrible sociopathic take (because it was trying to create a "gotcha" out of gloating you can make a little more money if you just betray everyone you know without remorse... until someone does it to you), I got what Ed was saying, and totally agree, but I think it can be well summarized as "the process informs the output, you dipshit".
"All I care about is the output" is something you can only say in vacuum. You hire people who are experts because you're not and don't fully understand the output as well as one. If all you care about is that the output "looks right", then sure, LLMs are great at that and your output and product will be completely outclassed by any output that "is right". It will be safer, more stable, and easier to maintain. Their job as experts is to care about or convince you to care about the stuff you don't know about or think isn't a big deal. If you think the output is impressive, it's clear you're a novice at whatever your assessing, because the entire tech is to create an average. If you don't care and just want to pump out a ton of below-average to mid-tier shit, then you're creating a world that you yourself don't want to live in.
"I have experts use it". Cool, your experts are getting dumber and lazier as time goes on, you've cut off the pipeline so that no new experts are created, and most importantly, THE OUTPUT IS STILL WORSE.
Example: If you have a world class composer write a song from start to finish, they will provide you with something of good quality. If you have a world class composer walk in at the end of a writing session being held with a bunch of highschool music theory students, and have them tweak whatever is there, the output definitely is not going to be the same. Not to mention that these days the expert is under new stricter time constraints.
The process informs the output, blowhole.
EDIT: Just heard this but also really feel it's pretty accurate. "These are people who have no empathy, and when suddenly told they have a competitive edge by way of staying exactly the same, not learning or changing at all, they latch on to it, thinking they will outperform others by being more okay with collateral damage, when that's really all they end up with."
r/BetterOffline • u/Ok_Display_3159 • 2d ago
Interesting thoughts from a former OpenAI employee on X
"You know, I might be one of the few people in this world who has worked at a frontier lab, observed the nature and pace of research progress, and actually become more bearish on AGI timelines (relative to pre-employment baseline) as a result."
- Despite the seemingly magical nature of LLMs, reflection over a >3 month timescale suggests my total productivity hasn’t increased by over 100%, or perhaps even by over 50%, and a lot of time is actually wasted because LLMs enable me to spend time on gratifying but low-productivity tasks that in the future turn out to not be useful
- Also, capabilities are incredibly spiky and highly correlated with the degree of investment poured into them, which my earlier tweet about math benchmarks implicitly points out
- From the above, it seems that the nature of LLM intelligence is wildly dissimilar to that of human intelligence and we won’t trivially get to something superior to human intelligence in all important respects just by scaling up existing approaches with various tweaks; even if AGI Is eventually achievable, this implies a significantly longer timeline
- Benchmark progress is almost definitionally guaranteed to happen because the process of constructing a benchmark is a direct precursor to the process of constructing a training dataset used for hill climbing that benchmark, but the scope of what can be captured in a benchmark is (at least for now) grossly lacking in terms of its relevance to real-world work, with maybe several limited exceptions
- Progress seems highly gated by data but the nature of model training means that each “next dataset” is significantly harder to assemble than what preceded it; some wins are possible through synthetic methods but those feel more like “patching up gaps” than “pushing the frontier forward”
At a higher level, I guess I’d say there’s a sort of refusal to think carefully about what models are or are not useful for in a rigorous way which I find personally quite annoying, and instead a reliance on some nebulous notion of being “AGI pilled” as a replacement for serious thought. I think people are very quick to anthropomorphize LLM intelligence because humans communicate through words and we infer the intelligence of human counterparties through comprehension of their language, but this leads them to wrong conclusions; for example if we observe that a new model proved some incredible mathematical theorem, some will say, “well, don’t we have AGI now, huh?” But to me, it’s actually more like, “well, given how hard it would have been for a human to do these mathematics, and given the limited economic effect of LLMs upon the world so far, isn’t it actually a negative datapoint vis-a-vis the generality of LLM intelligence?”
The ability of LLMs to make you waste time doing stuff that is not actually useful is massively underrated imo; I spent for example 10s of hours earlier in the year preparing random legal documents with ChatGPT but in retrospect I was too hasty and none of that work was useful. Of course it’s saved me time in some other respects and I think the net balance is positive, but it’s a quite significant countervailing factor.
r/BetterOffline • u/magick_bandit • 1d ago
One Datapoint
I’ve been planning on setting up a new website for one of my businesses, just never got around to it because the existing site is mostly good enough.
Mrs Bandit is an engineer, zero dev experience, and she said “well everyone is vibe coding, let me take a stab at it”.
I told her to go for it, I’m a professional dev, and I do AI literacy training, so that right there is a fun experiment because I get to see a really smart person try to use AI in an unfamiliar domain.
Claude design + Claude Code, some fable, mostly opus.
—
Result: 6 page started, visually they look nice. Some responsive design issues, but sure, at a glance just fine.
Then the code review…
4,000 lines of CSS. Many dead styles, duplication all over the place. Lots of small consistency issues across pages (like headers not being the same font/size/padding/margins, etc.)
If she kept going, it would become an unmaintainable nightmare. (It already is).
Expertise still matters a lot.
r/BetterOffline • u/deadpanrobo • 2d ago
With the release of GPT-6 Astra, im still struggling to see why anyone cares?
So GPT-6 Astra came out 2 days ago and with it we are seeing the usual "OMG we're cooked for real now seriously" posts being spammed into every subreddit.
This is because GPT-6 can now take control of your keyboard and mouse to do work on your device. This has been demonstrated by posts of fake time-lapse videos of the model drawing various artworks.
Of course these timelapse videos are very obviously AI due to how it seems to draw every artwork as if it were a printer, making the image line by line. Plus we still have the problem of the art it makes looking very generic. On top of 2D art it is also able to create 3D assets as well, but what ive seen it looks to mostly stick to modeling real life places or games that already exist.
Im still left with the same question I always come to when I see this cycle repeat, who's going to buy or play any of this?
Like nice you had an AI make another dime-a-dozen drawing of Hatsune Miku, congratulations, no i will not be purchasing this. It always is obvious that the people who are hyping this stuff up are people who don't really understand why art is created in the first place.
They believe that people create art for the end product. They see art (and by extension really everything) as products to sell to other people. They dont understand that most people draw for fun because they like the act of drawing and not because they want the thing they are drawing to be a fully realized art piece ready for sale.
The problem is that the only people who would want to purchase any of the generic and boring art created by AI are other AI enthusiasts and why would they buy someone else's AI artwork when they could just generate it on their own? We already know that the public *does not* spend money on something they know is AI generated so thats also a market that this art cant be sold to either.
Thats not even mentioning its coding capabilities, which is actually the field I have expertise on, is on the same level as all of the other "Frontier" models. Which is to say, they still suck at coding. And this is all before the company "nerfs" the model and it ends up being just like all the other models.
So again im asking why I should care, why should anyone care?
r/BetterOffline • u/Icy-Recognition-7453 • 2d ago
OpenAI quietly boosts some of Astra’s evaluation metrics, and continues to change others post-launch
r/BetterOffline • u/dayvancowgirl • 2d ago
My experience with LLMs outside the tech field (garment manufacturing, home construction)
Since most people here work in the tech industry (and I have learned so much from reading about everyone's experiences) I wanted to contribute a different perspective.
I work in garment manufacturing for small brands. We work with clients to take their ideas and make them into real products. Because we work with small brands, some of our clients do not have a lot of experience in the industry. They already needed guidance as to what is feasible (ex: you can put a screenprint here, but not an embroidery, or this embroidery is too small for the details to be seen, etc.) As you can imagine, we are now dealing with LLM-generated concepts.
Of course there is the general "ugh!" of them using LLMs to come up with ideas for supposedly stylish clothing. And the "ugh!" of them feeding my mockups into the LLM (which I definitely did not approve). Everything comes out generic.
The LLM is able to produce documents with a whole bunch of details that look like an official tech pack (this is the document factories receive which describes the materials, measurements, trims, etc.). This fools them into thinking the designs are feasible. They are not. We already had to deal with this problem, but now have the added issue that they think the LLM will only produce designs that make sense. We end up having to explain to frustrated clients that the LLM output is not workable.
Even experienced designers may have to experiment with processes and adjust designs based on what is feasible physically and financially. Different factories have different capabilities, the same kind of fabric from different suppliers might behave differently, etc. The LLM is definitely not going to magically get rid of this bottleneck. It's annoying because if they cared to learn, they could use the LLM to learn about pattern making, materials, and garment construction, because I'm sure there's tons of info about this online that was vacuumed up, but that's not the easy solution.
On another note, a friend who works for a major national retailer has the job of taking mockups from designers and picking the best ones to send to the next stage of production. The designers were encouraged to use an LLM to generate more designs which she then had to sort through.
Weirdly, the designers do not seem to have experience sewing or being involved in garment construction at all because they were sending her designs that were impossible to produce. Like, you would not be able to actually sew this garment or wear it if you did figure out how to sew it. She showed me one of the rejects and my first thought was, these are Möbius pants.
My husband is working as a general contractor with a small company that takes gutted houses and flips them. The investor is notorious for sending confusing LLM-generated mockups. For one thing, he just sends a rendering of a room—no measurements or anything, so the contractors have to do extra work to figure out how the concept will work in the actual space.
But the renderings themselves, while seemingly normal at first glance, actually have details that are... off. For example, a shelf intended for a big screen TV was placed about a foot off the ground. I'm sure there's someone out there who would actually prefer their TV low for some reason, but clearly there was no thought put into this. There was then an archway cutout in the wall which the TV was placed into, which makes no sense because you have this super low TV with a huge arched space over it. So visually there is the TV rectangle with another rectangle and a half circle of dead space above it. Maybe this idea could be executed well, but again, there was no actual thought put into how it would be executed.
It's extra stupid because the investor does actually have some good ideas... but they are now getting mixed in with slop because he doesn't care enough to be discerning. And of course the GCs on the ground are the ones tasked with executing the slop idea, or going through the ordeal of pushing back when they foresee a problem.
It is so obnoxious to me that people are not seeing the gap between actual, physical reality, and what the LLMs are producing. If they want to use the LLM to just brainstorm ideas, sure, whatever—but why are you then also assuming the LLM knows the intricacies of how a garment is sewn or how a room is constructed for humans to actually live in and use? (I know the answer...)
Fortunately this isn't a huge source of stress for us right now, but it does contribute to my weariness about how LLMs are pervading all aspects of life, even for those of us who work only partially on computers, or not at all. My best wishes to everyone who is in the trenches of dealing with LLM-pilled bosses, coworkers, and workplaces, and thank you to anyone who has shared their experiences and frustrations here, as it is a much needed perspective.
r/BetterOffline • u/CandidateCautious246 • 2d ago
Ed needs to get technical explaining why generative AI won't have ROI.
Ed is highly popular these days. All the AI bros know him and try taking him on. On all the podcasts Ed is on, they deliberately confuse domain-specific ML models with generative-AI. They point at Autonomous-driving, AlphaFold, Radiology-Models etc to say generative AI will succeed.
Ed does mention that those are not the same as generative AI. But these hosts and their audience can't make the distinction.
The technical argument he can try making is that Neural networks are a curve fitting tool. It's the same as statistical regression. Even the best ML model for a task should have some error percentage on it's train/test set because otherwise, it will not generalize well to real world input data. If the model doesn't make an error on it's test/train set, it overfits and performs poorly to real world inputs. A model is called a "model" for a reason. It's just an estimate. It's a guess based on prior examples.
Why can't LLMs (/neural networks) improve with RL? It's because when you update the weights to perform better on a certain task, the performance on a different task would suffer. Like a dog chasing it's own tail. Also, RL is very expensive, especially for a trillion parameter model.
The LLMs are good at the benchmarks because they have overfit the benchmark datasets.
Domain specific models like those used in radiology do help. But they have their limitations. It's precisely because they make those single digit percentage errors that you cannot reduce the number of employed radiologists. Because you always have to verify.
Maybe have a whiteboard explanation of why AI is not as capable of magic. He could invite folks like Gary Marcus or Emmanuel Maggiori to explain this.