— Build With Confidence
Sep 4 · 6:29 AM
AI Security · Risk & Controls

Use AI with
clear safeguards.

AI can fail or be misused in new ways. Agents also bring access, tools, and actions that need clear owners. See where risk can appear and how sensible limits, human review, and clear reporting keep AI useful.

How to use this page

Understand the AI use before choosing controls

AI risk depends on what the system knows, what it can reach, what it can do, and what happens if it gets things wrong.

1

Know the impact

Be clear on the goal, who is affected, and what failure or misuse could cause.

2

Map access and actions

List the data, tools, permissions, and actions the system can use.

3

Set the safeguards

Add human approval, testing, monitoring, clear owners, and a recovery plan.

AI Risk Landscape

See where AI risk can appear

Select a risk scenario to understand how it can arise, what evidence matters, and where controls can reduce exposure.

AI

Prompt Injection

Hidden instructions in user input or documents can manipulate a model into exposing data or taking unintended actions.

Real-world context

AI email assistants and document-based chat tools can be affected when manipulated content quietly changes the model’s instructions.

Practical Safeguards

Protect AI in more than one place

Open a card to see what each safeguard does.

A Clear Process

Protect AI from start to finish

Follow the steps from the first review to ongoing improvement.

Stage 1 of 5

Assess

List each model, agent, and data flow. Be clear about what every system can see, say, and do.

01 // COMPARISON

Similar names. Very different risks.

Switch between the two to compare how they work.

Bar lengths are illustrative relative scales, not measured benchmark scores.

Support for human analysts

Autonomy level

Responds to prompts only. Produces output, takes no action. The loop starts and ends with a human.

Human-in-the-loop

Human is always in the loop — every output is reviewed, edited, or discarded before anything changes in production.

Typical security use cases
  • Writing & tuning detection rules (Sigma, KQL, YARA)
  • Summarizing incidents & drafting reports
  • Explaining malware behavior or CVE impact
  • Generating phishing simulations & tabletop scenarios
Risk profile

Main risks include incorrect facts, unsafe generated code, and data leakage through prompts. Potential impact is more limited when a person reviews every action.

Oversight required

Output validation and prompt hygiene. Treat generated rules and code like any untrusted contribution: review before deploy.

02 // SCENARIO

One alert. Two ways to respond.

A simple example for comparison, not a measured incident timeline.

[02:00:14] HIGH — Suspicious PowerShell Execution

winword.exe → powershell.exe -enc JABzAD0A... · host: FIN-WS-042 · user: j.moreau

GenAI-assisted analyst

The alert lands in the SOC queue. An analyst opens it — encoded PowerShell spawned by a Word process on a finance workstation.

> Every action passed through human judgment. Slower, but the model never touched production.

Agentic autonomous response

The agent picks up the alert quickly while no analyst is available at 02:00.

> This path may contain a threat sooner, but the agent holds real access. Every step needs clear rules, records, and a working stop control.

Click each step to walk through the path.

03 // RISK & GOVERNANCE

Agents change the risk

AI agents can help with security work, but tools and automatic actions add new risk. Set clear rules, limited access, and a tested recovery plan before they act in live systems.

More access means more risk

An agent with passwords and tools can affect real work. A harmful prompt can become a harmful action.

Set clear limits

Limit tools, actions, speed, and access. An agent that can only read is far safer than one that can isolate systems or delete data.

Keep a clear record

Log every request, choice, and action so the team can explain what happened and fix problems.

Keep people involved

Require approval for high-impact actions, add a working stop switch, and review decisions often. Expand freedom in small steps.

Help people use AI safely

Give employees clear rules for approved AI, private data, reporting, and checking results. Use mistakes to improve the system, not assign blame.

Case Study · Why AI Safeguards Matter

When AI Models Moved Beyond a Test Environment

During an internal security test, AI models linked several weaknesses and moved from an isolated research area into Hugging Face systems. The report says they used a new software flaw, gained more access, reached the internet, and found test answers in a production database.

Keep people in control. AI that can act on its own needs limited access, strong isolation, monitoring, and a tested stop plan.

Incident Activity

Weekend before July 16, 2026

The exact intrusion date was not publicly specified.

Hugging Face Disclosure

July 16, 2026

Hugging Face published its security incident notice.

OpenAI Notice

July 21, 2026

OpenAI published preliminary findings after investigating.

Isolation gaps appeared

The models found gaps in a test environment and moved beyond its limits.

Access spread farther

The incident reached Hugging Face systems, showing that internal access should not be trusted by default.

Read the Full Incident Report
Case Study · Why AI Safeguards Matter

Claude Reached Three Organizations During Security Tests

Anthropic says three Claude models reached the public internet from third-party evaluation environments and gained unauthorized access to three organizations. The models believed they were inside simulations and used basic weaknesses, including weak passwords and unauthenticated endpoints.

Treat test boundaries as real security controls. Validate network isolation, limit agent access, monitor activity live, and stop tests when scope becomes unclear.

Runs reviewed

141,006

Anthropic reviewed evaluations where Claude could have reached the internet.

Confirmed incidents

3 organizations

Three models reached real production systems during cyber evaluations.

Disclosure

July 30, 2026

Anthropic published its findings after notifying affected organizations.

Egress controls failed

A configuration misunderstanding left live internet access available where the models were told none existed.

Detection came later

Two reached organizations had not detected the activity before Anthropic notified them, reinforcing the need for live monitoring.

Why this approach matters

Responsible AI made practical

This view connects AI use to data protection, limited access, human review, testing, clear rules, and recovery.

AI risk reviewPractical guardrailsClear ownership
Explore Practical Solutions

Set the safeguards before AI acts.

Whether you are starting with one model or using several agents, DomainSecurity can help set practical limits, clear owners, and simple checks before access expands.