This section provides an outline of the action plan that is followed during an introductory course, designed to introduce the subject of SCADA cyber-physical system security. This course was designed for students with basic knowledge about networking technologies and minimal familiarity with BSD or System V UNIX derivative systems. This profile roughly corresponds to a 2nd year BSc student, accordingly to the ACM standard Computer Science or Computer Engineering reference curricular structure, which will have no difficulty in grasping the subjects; however, trainees with a less advanced level of knowledge can easily acquire the fundamental prerequisites within a feasible timeframe. The course plan is organized into three different stages:
Introduction and context: before moving into the main course activities, trainees are introduced to the specific nature of SCADA ICS technologies, concepts and devices, which culminates with the presentation of the cyber range scenario; this introduction also provides the instructors with an opportunity to identify and start addressing the knowledge gaps that may exist within heterogeneous audiences.
Cyber range scouting/reconnaissance: ICS device, host and service enumeration procedures for pentesting and scouting procedures have much in common, in the sense that many IT-specific practices cannot be directly transposed to this domain. In this stage, students are introduced to a basic toolset for network and device scans, being challenged to identify as many devices as possible with minimal disruption.
Attack planning and deployment: this stage is dedicated to offensive procedures, from layer 2/3 floods to the execution of Man-in-The-Middle attacks. These attacks are used to demonstrate potential outcomes that may range from service disruption or interruption (due to device crashes or network resource exhaustion to loss of visibility and/or process awareness.
The evaluation of possible defensive mechanisms is also addressed, but not as in-depth as the other core topics, due to the introductory nature of the course. Next, we detail the various steps undertaken within the training modules/lessons, also presenting the expected outcomes for each one. Moreover, specific hints regarding the didactic strategies adopted for each step are also provided, based on the authors’ experience.
4.1. Introduction and Context
As most trainees come from an IT background, the first step must encompass an effort to bring the group to a similar level of basic knowledge regarding cyber-physical systems and the ICS domain. This initial module was designed to address these issues, being organized as depicted in
Figure 9.
Figure 9. Organization of the first training module (6 h).
The adopted strategy starts by introducing the concept of Critical Infrastructure and Essential Services, later moving into the specific details about the kind of Industrial Control Systems supporting these infrastructures. Later on during this stage, trainees are presented to the physical process that is part of the cyber range (presented in
Section 3.2), being given access to the HMI and allowed to directly interact with it. This first interaction has proven vital to stimulate curiosity and trainee motivation.
Once the introductory stage of the module is completed, comes the moment to deliver a conceptual and technical introduction of what a SCADA system is, in parallel with a detailed walkaround of the cyber range ICS process—the Purdue Model adopted by IEC 62443 [
22] is an important instrument to assist in this task, as it helps categorize the equipment by functional levels within an organizational structure. In this process, students become familiarized with the concept of master station, slave and field devices, also learning about the role performed by PLCs and RTUs.
Regarding the latter point, and because of its relevance within the ICS domain, special care is taken regarding the explanation of how PLCs work and how they are distinct from general-purpose computers: this includes an analysis of the PLC scan cycle (see
Figure 10), as well as the concept of cycle time and a quick overview of IEC 61131-3 [
28] programming languages.
Figure 10. PLC scan cycle.
Still within the scope of the IEC 61131-3 programming language introduction, trainees are are also put into contact with simple Structured Text (ST) and Ladder Logic (LL) examples (see
Figure 11). For this specific purpose, students are put into contact with the OpenPLC project, which is used to provide a containerized PLC runtime and a IDE (OpenPLC Editor) that can be used to experimentation purposes.
Figure 11. Simple example for a Greater-Than (GT) function block: LL vs. ST representations.
Being a crucial part of what makes SCADA systems possible, domain-specific communications protocols are also introduced and discussed. Among them, the Modbus protocol is dissected into detail, not only to provide an insight about the kind of legacy technologies which are still being widely used, but also because this protocol was adopted in the cyber range environment scenario.
The analysis of the Modbus protocol covers aspects such the framing structure (see
Figure 12), operation semantics and Function Codes, as well as the data types and addressing modes (Modicon and IEC/Quantum). During the discussion of communications-related topics, instructors also take advantage of the context to review other related aspects such as serial communications or TCP/IP protocol fundamental concepts (such as the three-way handshake or sequence numbers, for instance).
Figure 12. Modbus Application Data Unit (ADU) format.
Experimentation is always encouraged along the progression path; during this stage, a CANVAS ephemeral playground with several OpenPLC VMs is made available, also including some editor-bundled Windows VMs (for students who do not use the Windows OS or do not have resources to run VMs on their own PCs).
4.2. Initial Scouting/Reconnaissance Procedures
Network/range scouting is frequently one of the first stages at the start of pentesting campaigns, but it can also be part of an attack preparation effort. Regardless of the intentions, such efforts are necessary to gather information about the environment, discover and identify topologies, hosts and services. This module was designed to provide an introductory approach to this, being organized as shown on
Figure 13.
Figure 13. Organization of the second training module (6.5 h).
A two-way approach is undertaken: First, trainees are encouraged to search for information about the process control network having basically the same level of access they would have by compromising a host on the same segment; on the second stage, access to the network mirror interface will be provided, for traffic capture and analysis and validation of the some of the findings obtained during the first attempt.
At the start of this module, students go trough a quick recap of the fundamental principles behind techniques such as SYN or FIN scans or OS fingerprinting, that be used to get information about OS versions or TCP/IP open ports, also taking advantage of the communications concepts reviewed on the first module. Combined with extra information (as MAC addresses, albeit these are not always trustable and require being on the same Layer 2 domain as the device) or Open-Source Intelligence (OSINT) techniques, it is explained how the information collected from these sources can be further augmented with additional aspects concerning manufacturers and models, as well as firmware versions or known vulnerabilities (the work in [
29] is often suggested as reading material in this context). For this reason, vulnerability scanners such as OpenVAS [
30] are deliberately avoided in this course, instead encouraging students to engage in research processes.
Once the fundamental probing techniques are reviewed, students are then briefed about the specific limitations of ICS technology. This is a very important step, as people coming from an IT domain with a more solid background might be tempted to apply the knowledge acquired beforehand in a straightforward fashion—for instance, a common pitfall is to resort to vulnerability scanners, executing dangerous scanning routines that can disrupt the operation of an ICS. For this purpose, examples from authoritative sources such as NIST SP800-82 or talks by field experts about incidents caused by pentesting or vulnerability scanning procedures are presented [
31], in order to establish a set of golden rules for ICS security operations (see
Figure 14).
Figure 14. Good practices for ICS pentesting and scouting procedures.
All the procedures will be undertaken from each one of the students’ VMs, constituting the execution environment for most procedures and tasks. The first thing students will be asked to do is to run a simple
tcpdump [
32] capture on the network interface connected to the process control network, saving it in a PCAP format for analysis. Findings may be somehow scarce (moreover because we are dealing with switched ethernet network) but nonetheless interesting: Spanning Tree Protocol (STP) Bridge Protocol Data Unit (BPDU) frames, broadcast traffic (mostly ARP requests) and some multicast traffic (such as Multicast Listener Discovery protocol for IPv6, ICMPv6 duplicate frame detection or Simple Service Discovery Protocol traffic). Trainees will be encouraged to review these protocols, trying to infer information about existing hosts (for instance, ARP requests provide valuable insight about the network ranges which are being used on the network segment, as well as the address for potentially existing hosts, and STP PDUs may provide information regarding ports, switches, port priorities and addresses).
The analysis of ARP traffic also serves to show trainees how stealthiness can be easily thwarted by carelessness—a fatal failure for an attacker or during a red team CTF exercise, since an intruder host making or replying to ARP requests may generate unwanted “noise”, leading to its disclosure. To deal with this, students are taught about the Linux
arp_ignore and
arp_announce sysctl variables [
33], and how to use them to reduce/supress ARP traffic.
Once the initial capture analysis in concluded, students progress to the network scanning procedures, using the
arpscan [
34] and NMAP [
35] tools. Both are well-maintained, open-source and portable tools, albeit NMAP has a more complete feature set and a powerful command line interface and scripting capabilities, also being able to integrate with other toolchains.
Students will be asked to scan the entire network range, in order to detect active hosts, and in some cases also using spoofed IP addresses. Moreover, they are reminded about the need to perform a slow scan, because of the specific characteristics of ICS environments, albeit a professional attacker would probably do it anyway, in order to go unnoticed by intrusion detection techniques based on volumetric network traffic analysis. The first attempts are to be undertaken using the
arpscan tool (see
Figure 15). Trainees will quickly develop an intuition about the trade-off between stealthiness and information gathering potential—for instance, total suppression of the ARP protocol and use of a IP-less network interface may limit the effectiveness of many techniques (for instance, ARP scans will not properly work with total ARP suppression).
Figure 15. ARP scan results.
Afterwards, the NMAP tool will be used to implement a ping scan (see
Figure 16), with no DNS reverse resolution; this is a lightweight (albeit not error-proof) way to scan an IP range. Students are also taught to implement SYN and FIN scans, to acquire information about open ports and potentially available network services.
Figure 16. NMAP scan results.
Having access to the same Layer 2 domain as the testbed network (something that, in most cases, would not be possible in the open Internet), students are able to get information about Ethernet MAC addresses, providing extra insights about the nature of the devices found on the network. While extra information could be gathered from the devices using NMAP’s OS detection feature, this is ill-advised; this feature works by using TCP/IP stack fingerprinting techniques, sending a series of TCP and UDP packets to the target and gathering evidence about features such as response patterns, TCP Initial Sequence Numbers, initial window sizes, protocol options support, TOS fields or IP ID (for the MDL period) predictability, among others. Overall, it is too dangerous, considering the sensitive nature of many device TCP/IP stacks.
The use of device-specific NMAP scripts (such as the
modicon-info.nse [
36] script) and tools such as
smod [
37] (a Modbus penetration testing framework) are also used to complete the device profiling effort (see
Figure 17). Moreover, students are not told about the existence of a SCADA honeypot; this device appears legitimate at a first glance, requiring further effort to be identified as such.
Figure 17. NMAP results for the modicon-info.nse script executed against the Schneider M340 PLC.
Once the active scanning procedure is finished, the mirror interfaces will be activated, allowing trainees to get a copy of the testbed network traffic. The analysis of this traffic will reveal/confirm many findings, namely, the IP and MAC addresses of the M340 PLC, the location of the HMI and SCADA stations and some potential insights about the Arduino-based RTU needs to be analyzed with special attention.
While the PLCs and HMIs exhibit a traffic signature which is expected from a regular Modbus polling cycle, the Arduino-based RTU exhibits an unusual pattern in terms of ARP requests and TCP/IP traffic, with a MAC address prefix that seems very unlikely for this kind of hardware. The reasons for this are deeply rooted in the nature of the Wiznet W5100 network interface Application-Specific Integrated Circuit (ASIC), as well as in the limitations of a simple microcontroller, on which the Arduino is based upon. Thus, instructors provide a series of clues (Wiznet W5100 ASIC datasheets [
20], discussion about the Arduino Modbus library implementation) which are expected to lead students to find an answer for the unusual patterns.
During this module, and according with the background and skill level of the course attendees, instructors are advised to fill eventual knowledge gaps (for instance reviewing TCP/IP connection state diagrams or other required technical concepts), also encouraging (and giving time for) students to explore these concepts on their own. For student groups with a more advanced skillset, this step offers an interesting opportunity to explore OSINT techniques, which may be leveraged by students to make the most of the information that was gathered. Trainees are encouraged to work in groups (in some cases formed by the instructors’ initiative, to balance skill levels), having to work autonomously at the end of the module to deliver a complete report about the findings.
4.3. Attack Planning and Deployment
This constitutes the final course module, being structured around a “what if?” perspective on attack execution and outcome analysis (see
Figure 18). Instead of focusing on providing theoretical knowledge about attack techniques, trainees are allowed to execute attacks by themselves and observe the results. For this purpose, the CANVAS framework constitutes an ideal platform, as it was designed for easy recoverability in case of failure. This is further reinforced by the existence of Out-of-Band management mechanisms, as well as the introduction of STONITH (an acronym for
Shoot The Other Node In The Head, a term commonly used in active-passive cluster setups to designate remotely controlled power sockets that allow to cut off the power or power cycle a specific node) modules the allow to remotely cycle or power up any device in the physical process testbed.
Figure 18. Organization of the third training module (5.5 h).
This module starts with a review of the classic Layer 2/3 attack flooding techniques. While experience acquired from the IT field may have provided trainees with an intuition about the effects of targeting a specific network node, both in terms of network performance or communications/service degradation, the effects on ICS automation devices may be more radical, causing resource exhaustion or even equipment crashes. The fundamental principles behind most of the attacks hereby described are documented in past work from the authors [
38], which is also suggested as reading material.
To properly understand the potential effects of the aforementioned flooding attacks, students are introduced to the
hping3 [
39] tool, which is used to generate ICMP, SYN or UDP floods against cyber range nodes, using combinations of different transmission rates, spoofed source addresses (or even random source IPs), specific TCP flags, variable payload sizes and large windows. The
nping [
40] tool is also used for similar purposes, with other attack profiles also being tested (such as
Smurf or
Land attacks [
41]).
Once attacks start, students are asked to check both the traffic mirror interface and the physical process HMI to analyze and observe the attack effects. Quite often, traffic floods are enough to blind the HMI (see
Figure 19), causing loss of process visibility due to communications channel exhaustion, or even device unavailability. The latter case may be triggered quite easily due to the fact the M340 PLC tends to crash and become permanently unavailable (until the next power cycle) due to weaknesses in the TCP/IP stack implemented in the specific firmware version in use.
Figure 19. SYN flood effect on the SCADA system.
Besides layer 2/3, layer 7 flooding attacks are also analyzed, with a special focus on Modbus API abuse through means of variable frequency read and write operations (using the
smod tool or the
metasploit modbusclient [
42] module). Due to the lack of authentication of access control mechanisms, Modbus devices are often vulnerable to these sort of attacks, with unpredictable results. Moreover, and for similar reasons (lack of authentication), Metasploit can also be used to download the project running on M340 PLCs, by using the
modicon_stux_transfer [
43] module.
The next stage is dedicated to the review of the ARP protocol. The protocol state machine is analyzed, with a main concern in mind: bring to evidence its stateless nature, as well as the absence of security mechanisms, something that can be traced back to its roots. This step is also valuable to familiarize the trainees with the main vulnerabilities which are implicit to the protocol design, namely, the possibility of abusing it to deploy a MiTM attack (see
Figure 20). Such attacks can be used both as part of a scouting procedure (to better understand the nature of the communications between nodes) or to corrupt in-flight process data, by manipulating Modbus ADUs.
Figure 20. ARP poisoning-based MiTM attack.
For this purpose, trainees are introduced to the
ettercap [
44] and
bettercap [
45] tools, which are used to prepare and deploy the ARP MiTM attacks. After a short briefing on the tools and their usage, students are instructed to prepare a first campaign targeting the HMI and M340 PLC, in order to better grasp the relevance of such attacks. The first objective is focused on deploying a successful MiTM, using the mediator role of the attacker node (which will be the students’ VM) to capture traffic (see
Figure 21).
Figure 21. Traffic sniffing during ARP MiTM session.
Once the capture is acquired, trainees are asked to analyze its contents, searching for relevant patterns. Despite the fact that the mirror interface would already provide this information, there is an explicit intention of showing of such an attack can provide the means to effectively capture traffic for scouting purposes. For this purpose, the
Wireshark [
46] tool is used, in order to follow the network trace flows, using its filtering and dissector capabilities to easily decode Modbus ADUs (see
Figure 22).
Figure 22. Trace analysis using wireshark.
Using an approach similar to the approaches followed by possible attackers, trainees are expected to find the Modbus holding registers within the ADUs, which provide the information periodically requested by the HMI. Once this information is acquired, students are ready for the offensive MiTM stage.
Weaponizing the ARP MiTM requires the capability for manipulating in-flight packet data. Despite the fact that frameworks such as
Scapy [
47] are particularly apt for this purpose, the learning curve might not be compatible with the scope of an introductory course. For this reason, the authors have opted instead to use the built-in
ettercap scripting capabilities, by means of
etterfilter.
Ettercap filters get the implicit benefits from the ettercap tool, which handles traffic forwarding at the attacker node, also taking care of checksum recalculation when packet payloads are changed. Trainees are introduced to the
etterfilter syntax, being encouraged to experiment with scripts of their own. The obvious next step will be to inject fake state information into the HMI, using an ARP MiTM (see
Figure 23).
Figure 23. Injection of fake data into the HMI (byte offset deliberately hidden).
For this purpose, students are asked to leverage what they learned about the Modbus ADU format, together with the information acquired from the testbed traffic captures, the deployment of ARP MiTM attacks, as well as etterfilter syntax to plan and execute an offensive MiTM, with the purpose of injecting false temperature data into the HMI. This attack is particularly interesting as part of an offensive strategy: for instance, an attacker might use it to blind the HMI, allowing him to directly manipulate process data while going unnoticed to the SCADA system operators.
Moreover, students are also encouraged to try the same approach on other hosts, eventually discovering some particularly curious features about the cyber range; for instance, the Arduino-based RTU exhibits some sort of resilience towards this sort of attack, which can be explained by the same reasons that prompted the instructors to ask students to find an explanation to the unusual traffic patterns involving this device, during the development of the second course module.
This module finishes with a discussion of the potential countermeasures that could be put into place to defend from these attacks, with an historical review of several related proposals (such as TCP syncookies), as well as the presentation of several techniques and technologies that can be used for avoidance or mitigation purposes, such as whitelisting, port security, passive monitoring, and so on.
4.4. Final Considerations and Notes about the Course Development Strategy
The number of training hours outlined in the diagrams at the beginning of each training module subsection corresponds to synchronous (class) time. While the course structure is organized along a total of 18 h, note that trainees are expected to spend at least the same amount of time on individual out-of-class work. This is crucial for both knowledge consolidation and development of trainee autonomy.
Furthermore, note that the guided learning strategy that was adopted, which is focused on a hands-on approach, steers as much as possible away from creating dependency on the instructors, instead encouraging students to work as autonomously as possible. Instructors are not regarded as traditional schoolmasters, in the sense that their role is not to impose a specific learning style, but rather help trainees reach the goal of grasping the important concepts, while respecting as much as possible the particular students’ profile. For this reason, instructors are encouraged to strive for accessibility but also to profile trainees as quickly as possible in the early course stages, in order to be able to propose self-study material and strategies to help the less-experienced acquiring the minimum prerequisites. Instructors should not stand in the way of students with willpower, but rather provide them the means to turn commitment into results.
Finally, there is the question of assessing the validity of the knowledge acquired by trainees. This is undertaken by resorting to both individual and group-centric synchronous/ongoing evaluation strategies, as well as discrete checkpoints. For instance, the evidence for the students’ proficiency regarding the first training module contents is undertaken during classes (synchronously), with students also being assigned with developing a small group project (a PLC program) to perform a specific task. For the second module, synchronous (during class) evaluation is also undertaken, with students being asked to further explore the cyber range on their own and produce a report detailing their findings. The third module is also evaluated using an hybrid approach: synchronous evaluation, together with a group project focused on further exploiting and reporting any relevant vulnerabilities, which must also include a proposal to correct and mitigate the weaknesses and vulnerabilities that were found. The course is concluded with an open book written exam accounting for 50% of the final grade.