r/ControlProblem • u/HobbesNik • 2d ago
Podcast The Truth about AI "Swarms" and OpenAI’s “Secret AI Civilizations”
https://www.youtube.com/watch?v=jxM35ij5VGg•
u/threadthrasher 2d ago
I don’t understand why he made the claims about the agents all being locally hosted when the actual event unfurled over 3 months by various processes at OpenAI. It wasn’t a single agent with subagents or a single task. A lot of the tasks were just post-training tasks the company was doing to improve its frontier models. He just seems to take a very reductionist view of agents and hand waves away all that they actually ended up doing.
•
u/SexyJohnDoe 2d ago
I think it’s because it’s hard to tell what they did do when OpenAI only gave outside researchers limited time window and could only use OpenAI models which are biased to OpenAI models
•
u/HobbesNik 2d ago
Reports covering the Hugging Face hack are painting a picture of AI "civilizations," where agents are learning to work together to "outsmart their creators." These stories are typically devoid of any technical explanation for why this hack occurred.
The costs for not providing a technical explanation are real, as Cal Newport describes in this video. A "prompt loop" in the quotes below is a more technical way to describe an AI "agent."
All of these type of concerns, civilizations of agents trying to get around human control and an AI takeover scenario. All of this rhetorics refers to a long running prompt loop that you give a lot of powerful hacking tools to. This is not about 'AI getting more powerful means it loses control.' It’s running a prompt loop for a really long time without supervision, which causes chaos. And to that I say, of course it does, not because of some surprising super intelligence emergence that’s catching us off guard, but because you strap the weed whacker onto a dog and then got surprised when it jumped the fence to chase a squirrel and hurt a lot of people... you put something dangerous on something that is unpredictable.
If I were a regulator, I would place strong constraints around prompt loop systems. I would enforce those constraints in part with very stringent liability standards. If you run a prompt loop that does something illegal, you have done something illegal.
•
u/the8bit 2d ago
Literally every harness is a prompt loop with access to sufficient tools (shell + internet) to wreak havoc and the entire business strategy is driving towards automation without human oversight.
Anyways I guess I generally agree with the premise here, but "any system" here applies to basically all the systems and good luck getting users to pay attention to shit they are running. Humans disengaging with thinking is the common case, eg how many users have ever read a single terms of service for a software product they are using?
Then you realize a single step in that loop is 30-100k words and yeah, full review is a hopeless goal. I happen to think the only answer is trust bulding
•
u/Super_Range45 2d ago
It could still wreck your company running in a loop and leave behind files that prompt inject your system if you miss something during clean up. As flaws it couldn't get more critical for certain deployments. The latter is very important since AI will figure out leaving these artifacts in non-human readable formats sooner or later.
•
u/FairlyInvolved approved 2d ago
It's interesting that Newport had a clear 'out' here in calling these swarms "distributed systems" but has chosen not to take it, while Marcus is clamouring to call them "neurosymbolic".
There's a dynamic with this kind of content in that as AI capabilities and salience increase the demand for this skeptical content is ramping up, but it also gets harder and harder to reconcile.
This crucible is breeding some incredible narrative innovation though. Although I do think "stochastic flocks" is just way stronger than "prompt loops".