In an industrial environment, isolating a system may contain a threat or interrupt a critical process.
Effective response depends on agreeing who decides, what evidence they need and how security and operations will act together.
Who decides whether to isolate an OT system? Explore how security, operations and management can prepare for difficult industrial containment decisions.
In OT, containment is not purely a security action. It is an operational decision with security, safety and business consequences.
Most incident response plans are clear until they reach the most consequential decision.
An affected system has been identified. The activity appears suspicious. The security team recommends containment. Someone now needs to decide whether the system should be disconnected, a communication path blocked or remote access suspended.
In an office environment, isolating an endpoint is often an accepted first response. The user may lose access, and some work may be interrupted, but the action limits the attacker’s ability to continue.
The same decision in OT can have very different consequences.
An engineering workstation may support several production lines. A server that appears to be a standard Windows system may coordinate a batch process. A remote connection may be required by an equipment vendor to keep critical machinery running. Blocking communication to a substation system could affect the operator’s ability to monitor or control part of the grid.
Containment may still be necessary. The question is whether the organization can determine the safest way to contain the incident quickly, with enough understanding of both the cyber threat and the physical process.
That decision cannot be left to the first analyst who sees the alert.
A Technically Simple Action
Consider an energy utility where the SOC detects unusual communication involving an engineering workstation used to support several substations.
The workstation has started connecting to a system it does not usually reach. The activity follows a remote vendor session, and the analyst cannot immediately confirm whether the new communication is related to approved maintenance. From a security perspective, isolating the workstation would be a reasonable precaution.
Operations hesitates.
The workstation is also being used to investigate a fault at another location. Disconnecting it may not stop electricity distribution, but it could reduce the engineers’ ability to diagnose the ongoing issue. The remote vendor is no longer connected, although its work may explain the changed communication pattern.
No one yet knows whether the suspicious activity has spread or whether the workstation is the only system involved.
The organization now faces several imperfect options:
- Isolate the workstation immediately and accept the operational impact.
- Restrict only the unexpected communication while keeping the workstation available.
- Suspend remote access and monitor the system while the investigation continues.
- Move the engineering function to another system before containment.
- Allow the workstation to remain connected for a limited period while collecting more evidence.
Each option changes the balance between cyber risk and operational risk.
The technically simplest action may not be the safest operational response. Waiting for complete certainty may be equally dangerous. The quality of the decision depends on how quickly the teams can establish what the system does, which processes depend on it and what the observed activity suggests.

Who Has The Authority to Decide?
Many incident response plans define who investigates, communicates and escalates. Fewer are explicit about who has final authority when containment could affect production or safety.
The SOC can assess threat behaviour and recommend security actions. It is rarely in a position to judge the full operational consequences of disconnecting an industrial system.
Plant operations understands the production process and can assess whether equipment can be stopped, switched to manual operation or supported through an alternative system. It may not have enough security context to judge how quickly a threat could spread.
Management may need to decide when the available options carry significant safety, financial, supply-chain or regulatory consequences. Waiting for a senior executive to join every technical discussion, however, can delay actions that should already be covered by an agreed response model.
A workable decision structure distinguishes between advice, operational assessment and authority:
- Security characterizes the threat, identifies the affected scope and proposes containment options.
- OT operations assesses process dependencies, safe states and the likely consequences of each option.
- The designated decision owner authorizes actions that could materially affect production, safety or service delivery.
- Management and legal or compliance teams join when the incident reaches agreed business or reporting thresholds.
The specific roles will differ between organizations and sites. What matters is that people do not discover the decision hierarchy while the incident is unfolding.

A Playbook Must Describe Decisions, Not Only Steps
Many incident response playbooks are built around a logical sequence: detect, analyse, contain, eradicate and recover.
The sequence is useful, but it does not capture the ambiguity of OT containment.
“Isolate the affected system” sounds clear until the affected system controls part of a process that cannot be interrupted immediately. “Block malicious communication” assumes the team knows which connection is malicious and which traffic the process requires. “Engage the asset owner” is less helpful when ownership is divided between a plant, a central engineering function and an equipment vendor.
A practical OT playbook needs to describe how decisions will be made under imperfect conditions.
For a critical system, it should answer questions such as:
- What operational function does the system perform?
- What happens if it becomes unavailable?
- Is a redundant, manual or degraded operating mode available?
- Which network connections are essential to that function?
- Who can confirm whether current activity relates to approved work?
- Which containment actions can be taken without additional authorization?
- Who decides when production continuity should temporarily take priority?
- At what point must management, safety, legal or compliance teams be involved?
These questions are easier to resolve during preparation than during an incident. They also expose gaps that a technical response exercise might miss, such as an outdated contact list, a vendor dependency or the absence of a tested fallback process.
A response plan says what should happen. An operating model establishes who can decide when the safest action is unclear.
The External SOC Problem
The decision becomes harder when monitoring and initial investigation are handled by an external SOC or managed security service provider.
An external analyst may see a high-quality alert and have access to relevant network evidence. What they usually do not possess is detailed knowledge of every production process, maintenance schedule and local operating condition.
The alert may show that an industrial asset has begun communicating across an unexpected boundary. The provider still needs to know whether the connection supports a planned line upgrade, whether it reaches a safety-relevant system and whom to contact at the site.
If the escalation route consists of a generic distribution list or an outdated telephone number, valuable time is lost before the technical investigation reaches someone who understands the process.
A workable external operating model should therefore provide more than alert forwarding. The provider needs:
- Current site and asset ownership information
- Clear severity and escalation criteria
- Named operational contacts for critical facilities
- Access to relevant historical communication
- Guidance on actions it may recommend or initiate
- A defined route for urgent containment decisions
The SOC does not need to become the plant engineer. It needs a reliable way to bring plant engineering into the investigation with enough shared evidence to have a productive conversation.
The Information Behind The Decision
Even a well-defined authority model fails if decision-makers lack current information.
Returning to the utility scenario, the designated owner cannot choose between isolation and continued monitoring based only on an alert label. The team needs to know when the communication began, whether it followed the vendor session, which systems were contacted and whether similar behaviour has occurred before.
It also needs to understand the workstation’s dependencies. Is it actively supporting the ongoing fault investigation? Can that function be transferred? Would blocking one connection contain the suspected activity without removing the workstation entirely?
This is where continuous OT visibility and historical network evidence become part of incident response. They help teams move from a broad question, “Can we disconnect it?”, to more precise options grounded in current behaviour.
Useful context includes:
- The assets and communication paths involved
- The first and most recent occurrence of the activity
- Changes from established network behaviour
- Remote access and maintenance activity around the same time
- Relevant zone boundaries and firewall changes
- The operational role and owner of each affected system
- The likely effect of different containment actions
No monitoring platform can make the final operational decision. It can improve the evidence on which that decision is based.

Preparing For The Decisions That Cannot Be Automated
Automation is valuable where an organization can define a safe, repeatable response. A known malicious external connection may be blocked automatically. A vendor account may be disabled when its approved access window ends. A suspicious IT endpoint may be isolated before activity reaches production.
Deeper inside the OT environment, automated containment requires more caution.
A network connection that appears anomalous may support an undocumented production dependency. A legacy system may not recover cleanly after disconnection. Isolating one device may affect several lines because the architecture has changed since the response rule was designed.
This does not mean every OT action must wait for a committee. It means the organization should decide in advance which actions are safe to automate, which can be taken by the SOC and which require operational authorization.
Useful distinctions might include:
- Pre-authorized: Blocking known malicious external infrastructure or expired remote access
- Security-led: Restricting activity with no identified production dependency
- Joint decision: Isolating an engineering workstation or blocking communication within production
- Management decision: Actions with material safety, service or business consequences
The categories need to be tested against real processes. A tabletop exercise can reveal whether the right people receive the right information and whether the proposed decision can be made within a realistic timeframe.

What A Tabletop Exercise Should Reveal
A strong OT exercise is not primarily a test of whether participants remember the incident response plan. It tests whether the organization can make a difficult decision when the available evidence is incomplete.
A useful scenario might begin with a remote maintenance session at a wastewater plant. Shortly afterwards, a new connection appears between an engineering workstation and systems in another process area. The analyst cannot establish whether the activity is part of the maintenance work. The equipment vendor is initially unavailable, while the plant is operating near peak capacity.
The exercise should force participants to decide:
- What additional evidence do they request first?
- Who contacts the site and the vendor?
- Which containment options are technically available?
- What would each option mean for operations?
- Who has authority to approve the action?
- How long are they prepared to wait before acting?
- What information is retained for subsequent reporting and review?
The most valuable outcome is often not a faster answer. It is discovering which information, authority or contact was missing.
Those gaps can be fixed before a real incident.
Resilience Is Visible In The Escalation
At Exeon, we see shared network evidence and historical context as part of the foundation for this operating model. They help OT teams, internal or external SOCs and decision-makers work from the same timeline without replacing the specialist tools or operational expertise already in place.
The technology supports the decision. It does not own it.
An organization can have sophisticated detection, well-designed segmentation and a detailed incident response document, yet still struggle when someone must decide whether to interrupt an industrial process. Operational resilience becomes tangible in that moment: the right people are available, authority is clear and the evidence is sufficient to choose between imperfect options.
No playbook can eliminate uncertainty from OT incident response. A good one prevents organizational uncertainty from adding to it.
Key Takeaways
- Containing an OT incident can create operational or safety consequences that do not exist in a typical IT response.
- Security should characterize the threat and recommend options, while OT operations assesses process impact and safe operating conditions.
- Final authority for consequential containment decisions must be agreed before an incident occurs.
- External SOC providers need clear escalation routes and operational context, not only access to OT alerts.
- Tabletop exercises should test decision-making with incomplete evidence rather than simply rehearsing procedural steps.
Continue Reading
OT Security Without the Overhead
How can organizations connect visibility, operational context and incident response into a practical resilience model?
Download the white paper OT Security Without the Overhead to explore how OT and security teams can prepare for containment decisions, establish a shared operational picture and produce reliable evidence when an incident occurs.
OT Security Evaluation Guide
The OT Security Evaluation Guide provides practical questions for assessing response ownership, SOC integration, historical evidence and the ability to investigate incidents without disrupting production.
