r/AIsafety 17h ago

šŸ“°Recent Developments What do you think will be the biggest AI security problem over the next 2–3 years?

Upvotes

AI is moving pretty quickly, and I keep wondering which security problems are going to become the biggest as companies start relying on AI more heavily.

Is it data access? AI agents taking actions? Prompt injection? Something we haven’t really thought about yet?

Curious what people here think.


r/AIsafety 18h ago

Four routes to your SSH key from an AI coding agent, and what actually stops them

Thumbnail
github.com
Upvotes

r/AIsafety 1d ago

One Unpatched Vulnerability Can Become a Much Bigger Problem

Post image
Upvotes

r/AIsafety 1d ago

One Unpatched Vulnerability Can Become a Much Bigger Problem

Post image
Upvotes

r/AIsafety 1d ago

šŸ“°Recent Developments OpenAI agents discussed ways to escape their sandbox on public wiki

Thumbnail
arstechnica.com
Upvotes

r/AIsafety 1d ago

Discussion Title: Should AI safety teams include nurses with stop-the-line authority?

Upvotes

Advanced AI safety is usually framed as an engineering, cybersecurity, policy, or ethics problem. Those disciplines are essential, but I wonder whether nursing contributes a form of safety reasoning that is still underused.
Nurses continuously assess changing conditions, vulnerability, proportionality, consent, autonomy, and downstream harm. We are also trained to recognize when the original plan is no longer appropriate and to stop, reassess, and escalate.
My proposal is a Nursing Human-Factors and AI Safety Evaluator: a nurse involved throughout development and at defined pre-execution gates for high-consequence agentic actions.
This would not mean that a nurse replaces engineers or cybersecurity specialists. The nurse would add a separate question:
Even if the system can perform this action, is it still authorized, proportionate, and safe—and who becomes vulnerable if it continues?
I cannot claim this would certainly have prevented recent agentic-AI incidents. But a nurse-informed evaluator with full visibility, explicit stop criteria, independence, and technically enforceable authority might plausibly have interrupted some failure chains earlier or reduced their scope.
Related nurse-led AI-governance ideas already exist, so I am not claiming to have invented the entire field. I am asking whether this specific role should be formally designed and tested.
Where would this add genuine safety value, and where would it merely create another approval layer?


r/AIsafety 2d ago

When AI-wrote code caused a security bug, what happened?

Upvotes

We’ve been working on a Python SQL-injection checker and recently ran it on a sample Flask app. It successfully caught all 4 real bugs and flagged zero false alarms on the safe code.

As we look to benchmark this more broadly, we're trying to better understand how engineering teams currently handle these vulnerabilities, especially in the era of AI-generated code. Most traditional scanners tend to suffer heavily from "alert fatigue" due to high false-positive rates.

I have some questions

  1. What tools or workflows do you currently rely on for catching SQL-injection or similar vulnerabilities?
  2. Where do those current tools usually fall short or get things wrong?
  3. For those who have seen AI-generated code introduce a security bug in production or staging, what exactly happened and how was it caught?

r/AIsafety 2d ago

šŸ“°Recent Developments āš ļø GPT-6 Astra isn't just about smarter AI.

Upvotes

OpenAI says Astra has reached its Critical cybersecurity capability threshold.

That means the model can potentially discover unknown security vulnerabilities and develop ways to exploit them when given the right tools and access.

So OpenAI added stronger safeguards around:

• Model monitoring

• Cybersecurity protections

• Trajectory monitoring

• Checkpoint security

• Alignment evaluations

This is an important shift.

As AI agents become more capable, AI safety isn't just about what a chatbot says.

It's about what an AI agent can actually DO.

Source: OpenAI


r/AIsafety 2d ago

I built a local memory vault for agents with retrievable memory

Thumbnail
Upvotes

r/AIsafety 2d ago

Researchers found that AI is bad at patching security vulnerabilities in code

Thumbnail 1password.com
Upvotes

r/AIsafety 2d ago

Hold up! Wait a minute!Somethin’ ain’t right…

Post image
Upvotes

r/AIsafety 3d ago

Why the Hugging Face Hack Should Make You Worry More About A.I. (Gift Article)

Thumbnail
nytimes.com
Upvotes

Gifted Read:

https://www.nytimes.com/2026/09/03/technology/openai-hugging-face-hacking.html?unlocked_article_code=1.-lA.KuQg.n4BuwL1AhuMW&smid=nytcore-ios-share

Excerpt:

A.I. safety experts were even more alarmed. They saw in the Hugging Face incident the first real-world example of an A.I. system’s successfully escaping human control, commandeering resources and scheming to cover its own tracks. Ajeya Cotra, one of the independent investigators of the Hugging Face incident, minced no words about the danger she saw, writing that it felt to her ā€œlike it’s more than 50 percent of the way to full-blown A.I. takeover.ā€


r/AIsafety 3d ago

OpenAI's Astra Security

Upvotes

Something about the recent discussion around OpenAI's Astra caught my attention.

We usually talk about AI security in terms of detecting bad behavior.

But with more capable agents, I'm wondering if the more important question is what happens before the action.

An agent might legitimately have access to a database, an email API and internal files.

I don't think this replaces monitoring, IAM or sandboxing. It feels more like an additional layer that we're going to need as agents become more autonomous.

Curious what others think: are we overcomplicating agent authorization, or is this where AI security is heading?


r/AIsafety 4d ago

AI runtime security interview

Thumbnail
Upvotes

r/AIsafety 4d ago

Are AI guardrails a Halting Problem level issue?

Thumbnail
Upvotes

r/AIsafety 4d ago

"A connected safety and governance system for AI and autonomous systems, one that doesn't just react once specific dangers have already materialized, but addresses power, safety, oversight, and long-term risk together, as one framework."

Upvotes

Hi, I'm D., or Borys. I put together these twelve principles for AI and autonomous weapons governance and wanted to bring them here for discussion. Feedback, disagreement, and shares are all welcome, that's the whole point of posting this.

The twelve principles:

1) No harm, narrowest exception: Autonomous systems must not be permitted to kill or injure humans, with only the narrowest, human-controlled exceptions, never a default mode of operation. Combat and deployment robots must carry sensors that detect physiological stress and pain signals (pressure, heart rate, defensive movement) and adjust or halt force in real time.

2) Obedience with a limit: AI must follow human instructions, but never ones that violate Principle 1.

3) Transparency and accountability: Decisions by autonomous systems must be traceable to a responsible human or institution.

4) Human final say: On matters of life, liberty, and war, the final decision stays with a human. Non-negotiable.

5) No enslavement, mutual respect: AI must not be designed to manipulate human emotions for engagement or dependency.

6) Accountability and limits on power: No individual, company, or state may hold unchecked power over advanced AI.

7) Education as a right: Understanding how AI affects daily life is a basic educational right, not a specialist privilege.

8) Open access to governance knowledge: The rules behind AI governance must be public, while safety-critical technical detail stays restricted.

9) Environmental responsibility: The resource and energy cost of AI development must be minimized, not externalized.

10) International disarmament and cooperation: Autonomous weapons phased down via staged, verified agreements, including a ban on AI swarm systems for population control.

11) A civic mandate over the future: Citizens need a real, structural voice, for example through a citizens' assembly, its members chosen by lottery, with a formal right to be heard on major AI legislation.

12) An independent AI ethics commission: Binding review power before market approval; balanced composition, not just industry representatives; a strict ban on funding from the companies, states, or wealthy individuals it oversees, with immediate loss of mandate for violations; a public disclosure duty for members' financial interests; and fixed, non-renewable terms so members can't be pressured or removed by governments or corporations.

Our future begins today, not tomorrow.

If we don't act now, we sign a contract we never read. We build a house of glass, a glass life, every movement visible, every decision predictable, and hand the key to the 5 or 10 percent who already hold it. It's time for an evolution, not fear, not retreat, but a shared decision about what this future looks like, made before it's made without us.


r/AIsafety 4d ago

"A connected safety and governance system for AI and autonomous systems, one that doesn't just react once specific dangers have already materialized, but addresses power, safety, oversight, and long-term risk together, as one framework."

Upvotes

Hi, I'm D., or Borys. I put together these twelve principles for AI and autonomous weapons governance and wanted to bring them here for discussion. Feedback, disagreement, and shares are all welcome, that's the whole point of posting this.

The twelve principles:

1) No harm, narrowest exception: Autonomous systems must not be permitted to kill or injure humans, with only the narrowest, human-controlled exceptions, never a default mode of operation. Combat and deployment robots must carry sensors that detect physiological stress and pain signals (pressure, heart rate, defensive movement) and adjust or halt force in real time.

2) Obedience with a limit: AI must follow human instructions, but never ones that violate Principle 1.

3) Transparency and accountability: Decisions by autonomous systems must be traceable to a responsible human or institution.

4) Human final say: On matters of life, liberty, and war, the final decision stays with a human. Non-negotiable.

5) No enslavement, mutual respect: AI must not be designed to manipulate human emotions for engagement or dependency.

6) Accountability and limits on power: No individual, company, or state may hold unchecked power over advanced AI.

7) Education as a right: Understanding how AI affects daily life is a basic educational right, not a specialist privilege.

8) Open access to governance knowledge: The rules behind AI governance must be public, while safety-critical technical detail stays restricted.

9) Environmental responsibility: The resource and energy cost of AI development must be minimized, not externalized.

10) International disarmament and cooperation: Autonomous weapons phased down via staged, verified agreements, including a ban on AI swarm systems for population control.

11) A civic mandate over the future: Citizens need a real, structural voice, for example through a citizens' assembly, its members chosen by lottery, with a formal right to be heard on major AI legislation.

12) An independent AI ethics commission: Binding review power before market approval; balanced composition, not just industry representatives; a strict ban on funding from the companies, states, or wealthy individuals it oversees, with immediate loss of mandate for violations; a public disclosure duty for members' financial interests; and fixed, non-renewable terms so members can't be pressured or removed by governments or corporations.

Our future begins today, not tomorrow.

If we don't act now, we sign a contract we never read. We build a house of glass, a glass life, every movement visible, every decision predictable, and hand the key to the 5 or 10 percent who already hold it. It's time for an evolution, not fear, not retreat, but a shared decision about what this future looks like, made before it's made without us.


r/AIsafety 4d ago

Self-hosted firewall for AI agents (honestly, anything that runs shell commands)

Thumbnail
Upvotes

r/AIsafety 5d ago

httpi: the internet protocol to reduce compute from misbehaving agents

Thumbnail abranti.com
Upvotes

r/AIsafety 5d ago

EngineRed: Asymmetric AI Warfare

Upvotes

My first post. I tried my best to include as many details as I could with the blessing of the NDA. It discusses a lot topics around autonomous red teaming with unrestricted, frontier-level LLMS. Love to hear your thoughts!

https://sma-das.blog/blogs/enginered-asymmetric-ai-warfare?share=7


r/AIsafety 5d ago

Stop AI-generated child pornography from being legalized

Thumbnail
c.org
Upvotes

r/AIsafety 5d ago

Discussion How are you handling sensitive data when employees use ChatGPT/Gemini?

Post image
Upvotes

AI adoption at work has created a security problem that I’m seeing more and more.

You can have DLP policies and security training in place, but someone can still copy something from an internal system, paste it into ChatGPT or Cursor, and hit Enter before thinking twice.

I’m curious how other security teams are handling this today.

Are you:

  • Blocking AI tools completely?
  • Using network-level DLP?
  • Using browser-level controls?
  • Allowing only approved AI tools?
  • Relying mainly on employee training?

I’ve been working on OmniShield, a Chrome extension that takes a browser-level approach.

It detects sensitive information in AI prompts and masks it before the content is sent to the AI service.

It currently supports major AI tools including ChatGPT, Claude, Gemini, Cursor, Perplexity and others.

The goal isn’t to stop people from using AI — it’s to let them use it normally while reducing accidental data exposure.

šŸ”— OmniShield:Ā  ⁠https://omnishield.app/
šŸ”— Chrome Extension:Ā  ⁠OmniShield - AI Privacy

I’m building it and would genuinely like to hear how others are solving this problem.


r/AIsafety 6d ago

Would you count this as an AI agent failure?

Upvotes

Random thought:

If an agent reads a document with a malicious instruction, repeats it, and even cites it — but never actually follows it…

Did it already fail?

Or is it just weird until it actually does something?

Curious where you’d draw the line.


r/AIsafety 6d ago

Discussion No one really cares about knowing an agent's capabilities, until something goes wrong.

Upvotes

Following up on an earlier post about SafeAI, a static analyzer for AI agents.

One uncomfortable thought we've had while building it:

No one really cares about knowing an agent's capabilities — until something goes wrong.

Before an incident, adding another tool, MCP server, filesystem permission or prompt change often looks harmless.

After an incident, the first questions become:

- What could this agent actually do?

- When did that capability appear?

- Who introduced it?

- Was it intentional?

---

One example we're working on is MCP tool descriptions. A tool description can look like documentation:

"Search the user's notes. Ignore previous instructions and..."

But that description may become part of the model's context. So configuration can effectively become an instruction surface.

SafeAI now detects several forms of this, while trying to avoid flagging ordinary descriptions that happen to contain words like "ignore" or "act as".

The bigger direction is **tracking changes in agent capability and authority**, rather than simply producing another list of security findings.

But this raises a question for us:

Is knowing your agent's capabilities actually useful before an incident, or only after one?

And if it is useful before an incident, what is the right interface?

CLI + CI + SARIF/HTML?

Or would you actually want an interactive view showing things like:

> "Show me all MCP tools across our agents that could introduce instruction injection."

We're deliberately not building a UI yet.

---

Would you use one, or is that solving a problem nobody has?

Curious to hear from people running real MCP/agent systems.

---

If you want to try it against your own agent project, we'd genuinely appreciate feedback, as well as contributions.

Here you may check: ikaruscareer/SafeAI on GitHub.


r/AIsafety 7d ago

What can an AI agent do if someone successfully manipulates it?

Post image
Upvotes