Experience || My Views on the Failure Frequency of DCS
Thread Content
1 Overview of Distributed Control Systems (DCS): DCS systems feature versatility, flexible system configuration, comprehensive control functions, easy data processing, centralized display and operation, a user-friendly interface, simple and standardized installation, convenient debugging, and safe and reliable operation. They are widely used in various industrial fields around the world, including power generation, petroleum, chemicals, metallurgy, and light industry, especially in large-scale generator sets. The main brands that are widely used in China at present include: (1) Foreign brands: Honeywell, ABB, Westinghouse, Siemens, Yokogawa, etc ; (2) Domestic: Guodian Zhishen, HollySys, Xinhua, Zhejiang University Zhongkong, etc. The safety and reliability of DCS are crucial for ensuring the safe and stable operation of the unit; any problems that arise can lead to severe damage to the unit’s equipment or even cause safety accidents involving personnel. Therefore, it is highly necessary to analyze various problems that arise during DCS operation and take measures to improve the safety and reliability of the DCS in thermal power plants. Faults in 2DCS during the production process: Each manufacturer’s DCS has its own characteristics, which means that the analysis and handling of faults vary from one case to another. However, in general, the faults that cause category II or higher level disruptions in the equipment can be divided into three main categories: (1) Problems with the system itself, including design and installation defects, as well as software and hardware failures. (2) Failures caused by human factors, including operational errors committed by personnel, inadequate management systems, and poor implementation at the execution stage. (3) External system environment issues cause DCS failures. Abnormalities can be caused by factors such as high ambient temperatures, excessive or low humidity, dust, vibration, and small animals. 2.1 Examples of problems and faults within the DCS itself. Such faults are quite common during production; they include defects in system design and installation, crashes of controllers (DPU or CPU), loss of network connection, black screens on the operator stations, disruptions in network communication, software defects, insufficient system configuration, and issues with interfaces to other systems and devices. 2.1.1 Power supply and grounding issues (1) The DCS power supply system in a certain power plant uses ABB’s Symphony III type of power supplies; however, during the infrastructure construction phase, the cabinets were installed according to the grounding methods applicable to Type II power supplies, which differ significantly from the grounding requirements of Type III power supplies. Since the unit was put into operation, there have been multiple instances of DCS module failures, signal fluctuations, and hardware damage, which are suspected to be related to the grounding system. Similarly, during the construction phase of a power plant, there were issues with the design, fabrication, and installation of the DCS grounding network; as a result, after the DCS system began operating, all temperature measurement points using thermoresistors and thermocouples experienced periodic fluctuations. (2) A factory experienced a failure in the control system on the turbine side due to loose power supply connections. Lessons learned: A DCS without a proper grounding system and adequate cable shielding not only leads to significant system interference and erroneous signals from the control system, but it can also cause damage to the modules. It is evident that issues such as UPS power supplies and grounding of control systems pose significant risks to the safe and stable operation of the DCS once the power plant begins operations. Therefore, the power supply design of a DCS system must include reliable backup solutions, with a reasonable load configuration that provides a certain degree of redundancy ; The system grounding of the DCS must strictly comply with the manufacturer’s technical requirements (if no special instructions are given by the manufacturer, DLT774 provisions shall be followed). All cables carrying control signals into the DCS must be qualified shielded cables; they should be laid separately from power cables and must have proper single-point grounding. 2.1.2 System configuration issues (1): Frequent failures and crashes of the DCS (T-ME/XP system) at a power plant in Zhejiang led to unit shutdown incidents. Units 7 and 8 (2*330MW), from trial operation in February to May 1997, experienced 22 DCS system failures and crashes in total, resulting in 8 abnormal trips of the units. Subsequently, operational screen malfunctions occurred multiple times (in Unit 8, all six operating stations experienced a “black screen” on two occasions), posing a serious threat to the safety of the unit. Analysis indicates that the DCS system has the following issues: ① There are problems with the DCS system design in terms of performance calculation software and digital input/output redundancy configuration. ② Mismatch in hardware configuration (including matching and communication issues between the T-ME and T-XP systems). ③ Some hardware designs are imperfect. ④ Further analysis reveals that the key issue is a \"bottleneck\" phenomenon caused by the excessively high load rate of the CS275 (lower-layer T-ME) communication bus. For users of the European T-ME/XP system, under proper configuration, the performance of the T-ME/XP system is generally good. (2) The DCS used by a power plant for the automation upgrade of its 200MW units had inaccurate load rate calculations due to the system configuration, and to reduce costs, its technical specifications were pushed close to their allowable limits. Additionally, this system featured a large number of virtual I/O points during operation; as a result, during the later stages of debugging, it was found that the load rate of some controllers exceeded 90%, and the response time for certain soft manual operations was nearly 1 minute, rendering them unusable. The problem was only resolved after significant adjustments were made (including reconfiguring the system). (3) In a 600MW unit in the northeast, due to the inadequate description of the I/O channel isolation requirements in the tender specifications, the DCS manufacturer provided a very basic configuration. As a result, many I/O cards were damaged during commissioning; later, the isolation method was changed and new hardware was installed. The power plant incurred significant additional costs, which offset any advantages associated with the lower price obtained through the initial bidding process. In addition, great attention must be paid to the quality of cables and shielding issues. For important signals and control systems, special shielded cables designed for computers should be used. In many renovation projects, cable problems have necessitated re-laying of the cables, thereby affecting the project schedule. (4) The engineer station of the Xinhua XDPS-400 system in a 300MW power plant unit was experiencing frequent crashes; upon inspection, it was found that there were numerous running programs: multiple virtual DPUs, historical data recording, performance calculations, reports, etc. Problem solved regarding the allocation of historical data to other HMI stations. 2.1.3 Controller (DPU or CPU) failure (1) In a power plant, a CPU failure occurred in FSSS1 of the HIACS-5000CM control system for the 300MW #2 unit; the control rights were not transferred, and it was not possible to switch from the CPU to the main controller, as a result of which the control devices for that section of the system could not be operated (the devices continued to function in their original state). During the online replacement procedure of the main CPU until a power outage occurs, control is switched from the secondary CPU to the main CPU; the system devices remain under control, and everything functions normally after the original main CPU is replaced. (2) In some early models produced by ABB, there were inconsistencies in the data exchanged between different controllers within the same SYMPHONY PCU cabinet; this issue was resolved by upgrading the firmware ; (3) In the early versions of the XDPS system controlled by Xinhua, certain batches of DPU units experienced frequent issues such as going offline and crashing. Upon inspection, it was found that there were problems with individual capacitors in those DPU cards; the issue was resolved by upgrading and replacing those cards. Since the controllers in current DCS systems are all configured redundantly, **the frequency of unit tripping caused by “abnormalities” in the main controller has been reduced.** However, if a pair of redundant controllers fail simultaneously, it will pose a direct threat to safe production; measures must be taken to prevent such situations from occurring. 2.1.4 DCS Network Failures (1) In the Westinghouse WDPF control system of a certain power plant, numerous modifications to the system resulted in an increase in the number of measurement points and automatic control circuits, causing the system load to exceed 70%. This led to network communication issues, with operators experiencing long delays when performing operations or switching between screens, as well as screen blackouts. Later, it was upgraded to the OVATION system, which is functioning normally. (2) In a power plant with a 600MW unit operating at a load of 508MW under stable conditions, all the throttle valves of the turbine suddenly started to move drastically. Upon investigation, it was found that the cause of the problem was a sudden change in the speed signal from the M5 controller from 3000 r/min to 0 r/min within a short period of time, before returning to its normal value. The movement of the throttle valves was also caused by data loss in the communication between M3 and M5, which resulted in the Trip Bias signal changing from 0 to 1 while the unit was running, thereby causing all the throttle valves to move drastically. Measures to address this issue: multiplex the communication signals on the PCU control bus and introduce a certain delay to these signals, thereby avoiding any sudden fluctuations in them ; Communication redundancy is employed for important communication signals. 2.1.5 DCS software issues (1) During the DCS commissioning of a 300MW heating unit in a power plant, the quality parameters of the measurement points were not modified; as a result, the analog measurement points were only considered to have poor quality when the connection was lost, and the quality verification function did not operate properly. Subsequently, the quality parameters for all measurement points were set, improving the reliability of the equipment operation. (2) When configuring the screen layout of the HIACS-5000CM control system, double-clicking the grab configuration tool results in a C++ error window appearing, preventing normal use. Upon inspection, it was found that the grab.ini file had been modified; after copying a file from another machine to overwrite it, the tool returned to normal operation. Because Grab exits abnormally, leaving error information in the grab.ini file. (3) The logic of the deaerator water level control circuit in a certain power plant was copied and modified from the high-pressure heater water level control logic; the modification was not thorough, and the PID parameters were not adjusted accordingly to the conditions of the deaerator. As a result, the water supply valve of the deaerator exhibited erratic regulation during operation, leading to a deterioration in the quality of regulation. Take measures: Check the logic and reset the PID parameters. 2.1.6 System interface issues: A 200MW heating unit in a power plant had only one signal pathway for the electrical grid connection to the DEH system. During normal operation of the unit, a fault occurred in the auxiliary contact related to this electrical grid connection, resulting in oscillations that led to the shutdown of the turbine. Measures to be taken: Use shielded communication cables, add redundant contact signals, and perform a 2-out-of-3 logical judgment. 2.2 Examples of DCS failures caused by human factors Human factors leading to DCS failures are also quite common during the production process. These include human error, inadequate management systems, and failure to follow the prescribed work procedures. 2.2.1 Failure to follow the prescribed work procedures: (1) In a power plant, the #12 DPU of the DEH in the Xinhua XDPS system malfunctioned. An spare DPU from the small turbine’s MEH system was used for its online replacement. After replacing the DPU, only the #32 primary control DPU was copied to the #12 secondary control DPU’s non-writable electronic disk; in essence, this simply ensured that the memory content of the secondary control DPU matched that of the primary control DPU. The content of the #12 DPU’s electronic disk remained the MEH mini-server control logic. After the system was powered down for soot blowing, #12DPU was started in sequence to become the main controller. Since its logic was MEH logic rather than DEH logic, this caused communication issues within the system, data flickering, abnormal display of images, and the inability to operate the human-machine interface station. After powering #12DPU back on, copying the #32DPU logic and writing it to disk, everything is working properly. (2) In the HIACS-5000CM control system of a certain power plant, during the replacement of a remote I/O card in the circulating water pump room, the online replacement procedures were not followed; as a result, the card failed to activate and enter operational mode, which caused the actual status of the equipment on site to differ from that shown on the DCS screen, rendering the equipment uncontrollable. After performing the online replacement procedure, the system is working properly. 2.2.2 Human error (1) During the operation of a power plant unit, while attempting to address a defect, staff mistakenly operated the relays in the DCS relay cabinet, resulting in the shutdown of the induced draft fan and the activation of the boiler MFT. (2) A DCS card failed in a power plant; during the process of replacing the card, the staff failed to carefully verify the equipment and the card, and incorrect wiring caused the newly installed card to get damaged. 2.2.3 Incomplete management system (1) The management system for the DCS system at a certain power plant is incomplete; there are no regulations regarding software upgrades, backups, and similar tasks. After the upgrade and patching, the POK1 operator for the auxiliary network water treatment system did not create a backup. After a system recovery was performed due to a hard drive failure on the operator station, its outdated software version caused abnormal network communication and failure of data updates. (2) The operator station at a certain power plant was not properly managed; the USB ports and optical drives on the main computer there were not properly secured. Some operators used the operator station to play games and watch movies during night shifts, which caused the station to crash. 2.3 Examples of DCS failures caused by external environmental factors: The number of DCS failures resulting from external environmental factors is relatively small compared to the first two types of problems, but such failures do occur from time to time in actual production processes. (1) The air duct opening in the electronic equipment room of a power plant was located above the DPU cabinet. Due to design and other factors, during operation of the unit, fire-fighting water flowed into the DCS cabinets through the air ducts, causing damage to devices such as the DPU and servers as a result of water intrusion, and leading to the shutdown of the unit. (2) In the remote IO cabinet of a power plant’s circulating water pump room, poor sealing at the bottom allowed rats to enter during winter; they built nests in the higher-temperature areas above the cabinet, which ultimately resulted in the loss of both networks for the remote IO system. (3) The electronics room of a certain power plant has poor sealing, with severe dust accumulation on the card holders and DPU units, which has led to multiple failures. After taking measures such as improving electronic sealing and installing air conditioning, failures such as those with card components and DPU have been virtually eliminated. From these various fault examples, it is clear that to reduce the likelihood of failures in DCS systems, it is necessary to carry out comprehensive work on distributed control systems throughout their entire lifecycle, from selection and design to operation and maintenance. 3 DCS Failure Prevention and Maintenance Measures 3.1 Selection, Design, and Commissioning of DCS 3.1.1 Whether it is a new unit or a DCS undergoing upgrading, the configuration of the system and controllers should take into account reliability and load capacity (including redundancy) as key factors. The load rate of the communication bus must be kept within a reasonable range, and the load rates of the controllers should be as balanced as possible, in order to avoid the occurrence of \"high load\" situations that can affect the safe operation of the system due to insufficient funds resulting from large-scale implementation. 3.1.2 The allocation of system control logic should not be overly concentrated in a single controller; the main controllers should be equipped with redundancy. 3.1.3 The power supply design must be reasonable and reliable. First, it is necessary to emphasize the load rate of the power supply design ; Second, it is necessary to emphasize the redundant configuration of power supplies; at the same time, the independence of the two power supply paths must be ensured. 3.1.4 Attention should be paid to reliability measures for DCS system interfaces. Emphasize the redundancy of critical interfaces and the selection of interface methods, with a focus on reliability and real-time performance. 3.1.5 For the grounding of DCS systems, it is essential to follow the manufacturer’s requirements in order to prevent widespread system failures caused by grounding issues. Attention should be paid to considering the system’s anti-interference measures, self-diagnosis, and self-recovery capabilities; isolation measures should be emphasized for I/O channels. The quality of cables and shielding issues must also be given high priority; specialized shielded cables for computers should be used for important signals and controls. 3.1.6 Full consideration must be given to the controllability of main and auxiliary equipment. Based on the operating characteristics of the equipment and the requirements for the unit to handle emergency faults under various operating conditions, operator stations and backup manual control devices should be provided. The configuration of the emergency shutdown button for the furnace should use a separate control circuit that is independent of the DCS. At the same time, one should not blindly pursue a \"simpler\" human-machine interface; system configuration should prioritize ensuring safe production. Special emergency intervention procedures related to safety cannot be entirely based on a functioning DCS. 3.1.7 For peripheral devices such as actuators and valves related to unit safety, their design and configuration must ensure that these critical devices can operate in a safe manner or remain in their original position in the event of power loss, air loss, signal loss, or DCS system failure. 3.1.8 For protection systems, a multiplexed signal acquisition method should be employed, and interlock conditions should be used appropriately to endow the signal circuits with logical judgment capabilities. 3.1.9 During debugging, all logic circuits, loops, and operating conditions shall be tested in accordance with the debugging outline and specific procedures. 3.2 DCS Operation, Startup, Shutdown, and Maintenance 3.2.1 Preparing for Maintenance To carry out maintenance on the DCS system, it is necessary to do the following: (1) Maintenance personnel should understand the overall design concept of the system. Familiar with the structure and functional components of DCS systems, knowledgeable about the hardware of system devices, aware of the normal and abnormal conditions of various components such as controllers, IO cards, and power supplies, and proficient in using DCS configuration software. (2) System backup: includes the operating system, drivers, boot disk, control system software, license disk, and control configuration database, ensuring that the control configuration data is up-to-date and complete. To address the issue of optical discs wearing out easily in actual use, it is important to create frequent backups, using devices such as external hard drives, USB drives, and hard disks to ensure that all software is preserved. (3) Hardware inventory: For vulnerable components with a short service life, as well as critical components such as keyboards, mice, I/O modules, power supplies, and communication cards, an appropriate number of spares should be kept based on actual needs. There should be at least one spare unit for each type of card or module. These spares must be stored in accordance with the manufacturer’s requirements. Where possible, the condition of the spare components should be verified to ensure their proper functionality. (4) Compile the after-sales service scope and schedules for various products to create a directory of technical support personnel from hardware manufacturers and system design units, thereby making full use of the technical support provided by DCS suppliers and system design units. 3.2.2 Daily maintenanceDaily maintenance of the system is the foundation for the stable and efficient operation of the DCS system. The main maintenance tasks are as follows: (1) In accordance with the requirements outlined in the 25 anti-measure guidelines, DL/T774 maintenance procedures, and other relevant regulatory documents, improve the management system for the DCS system. (2) Ensure proper sealing between electronic devices to prevent small animals from entering; minimize the adverse effects of dust on component operation and heat dissipation. Maintain temperature and humidity levels in accordance with the manufacturer’s specifications, thereby preventing condensation on system equipment caused by sudden changes in temperature and humidity. It is possible to consider introducing the ambient temperature signal from the DCS electronics into the CRT, along with an alarm. (3) Daily check whether the fans in each cabinet of the system are functioning properly and whether there are any blockages in the air ducts, to ensure that all devices in the system can operate reliably over the long term. (4) Ensure the quality of the power supply for the system and provide reliable power from two sources; an alarm is triggered when either source fails. (5) The use of wireless communication devices is prohibited between electronic devices to avoid interference from electromagnetic fields on the system; moving operation stations, monitors, etc., should be avoided, as well as pulling or damaging device connection cables and communication cables. (6) Standardize the management of DCS system software and application software; modifications, updates, and upgrades of software must follow approval procedures and a designated responsibility system. The use of unauthorized software and the installation of software unrelated to the system are strictly prohibited; proper management must be exercised over the USB ports and optical drives on the host machine. (7) Ensure proper recording of system data such as PID parameters for various control loops and the forward/reverse action settings of the regulators. (8) Check whether hardware such as the control host, monitor, mouse, and keyboard are in good condition, and ensure that real-time monitoring is functioning properly. Check the fault diagnosis screen to see if there are any fault alerts. (9) When powering on DCS devices such as the DPU and human-machine interface stations, it should be done in a specific order, one device at a time. After ensuring that each device is operating normally, the next device should be powered on, in order to avoid situations where abnormalities arise and are difficult to diagnose. After power is applied, the communication connectors must not come into contact with conductive objects such as cabinets. The redundant communication cables and connectors should also not be in contact with each other, to prevent damage to the network card. (10) Regularly conduct online tests on the communication load rate of the DCS main system and all related systems connected to it. Check the status of redundant master and slave devices; switch between them when conditions permit or on a regular basis, and investigate the reasons for any automatic switching by the devices. (11) Improve configuration readability: Added Chinese descriptions to important configuration pages ; Prepare detailed logical documentation that is consistent with the programming and configuration of the critical protection systems ; Prepare test operation cards and ensure they are updated at all times. Standardize DCS configuration tasks, and try to avoid making major configuration changes during unit operation. Care should be taken when configuration is necessary, and appropriate technical and safety measures must be implemented to ensure the safe and stable operation of the DCS and the unit. (12) Reboot all human-machine interface stations regularly, one by one (it is recommended to do this every 2 to 3 months), in order to eliminate the cumulative errors that occur over time due to the continuous operation of the computers. 3.2.3 During the maintenance period when the unit is out of service, thorough maintenance of the DCS system should be carried out, which mainly includes: (1) Using the time available for unit maintenance to reset the DPU, CPU, operator stations, and data stations of the DCS system one by one ; Remove invalid I/O points from the configuration to optimize it. (2) System redundancy testing: Conduct redundancy testing on redundant power supplies, servers, controllers, and communication networks. Pay close attention to whether the master-slave device switching, the network, and the human-machine interface stations function properly when various devices lose power during the system shutdown process ; After powering on the system again following maintenance, switch tests are conducted on various devices. (3) System dust removal: With the system shut down, the entire system is cleaned of dust, including components such as the inside of the computer, the control station chassis, power supply units, fans, and cabinet filters. (4) Maintenance of the system power supply lines, including testing the UPS’s power supply capacity and performing discharge operations. At the same time, check the voltage of the CMOS battery on the DPU host card and replace it regularly to prevent loss of CMOS data due to the battery. (5) Grounding system maintenance. This includes terminal inspection and ground resistance testing. (6) On-site equipment maintenance shall be carried out in accordance with the maintenance procedures and with reference to the relevant equipment manuals. (7) Check the interfaces between the DCS system and other systems; ensure redundancy for critical signals. For communication with other systems, adopt single-direction transmission and install firewalls depending on the specific circumstances. (8) Powering on the system: The maintenance supervisor shall confirm that all conditions are met before powering on the system after major repairs. And the power-on procedure must be strictly followed. 3.2.4 The fault repair and maintenance system shall carry out passive maintenance after a fault occurs, which mainly includes the following tasks: (1) In daily operations, it is necessary to strictly follow the 25 anti-misoperation requirements, and thoroughly prepare for various potential accidents including DPU (CPU) crashes and network communication failures. Emergency handling procedures, safety measures, technical measures, and maintenance steps should be documented in order to ensure the safe operation of the unit. (2) Handle DCS failures in accordance with the requirements outlined in the manufacturer’s application manual. Before replacement, verify that the model and address of the card module (ensuring no conflicts with addresses of other devices), as well as the jumpers, are consistent with those of the card to be replaced, and strictly follow the online replacement procedure. (3) For passive maintenance of faults, the work order system must also be strictly followed to avoid hasty repairs; a detailed analysis should be conducted based on the specific manifestations of the fault. Based on the self-diagnosis alarms of the DCS system and the analysis of fault symptoms, the faulty location is identified, and the results of the repair are verified by eliminating the alarms. For example: Poor contact at the communication connector can cause communication failures; once poor contact is identified, use tools to reseat the connector ; Damaged communication cables should be replaced promptly. If the fault light of a card is flashing or all the data on that card are zero, possible reasons include incorrect configuration information, the card being in standby mode with the redundant terminal connections not made, a fault with the card itself, or the absence of configuration information for that slot. When a certain production condition becomes abnormal or an alarm is triggered, one can first locate the instrument that reflects this condition, and then follow the direction in which the signal is transmitted to use instruments to check the accuracy of the signal one by one, until the source of the fault is identified. (4) A work order must be issued for the repair of on-site equipment failures, and DCS forced shutdown and isolation measures must be implemented. When repairing the valve, a bypass valve should be used. After the maintenance is completed, notify the centralized control operators promptly to conduct inspections; the operators should switch the automatic control circuit to manual mode. (5) In the event of large-scale hardware failures, failures of unknown cause, or failures that exceed the technical capabilities of the factory’s maintenance staff, in addition to carrying out emergency spare parts replacements at that moment, it is necessary to contact the manufacturer promptly so that its professional technical support engineers can further diagnose and resolve the issue. 4 Conclusion DCS should be managed comprehensively throughout the entire process, from design and construction to commissioning and operation. As system maintenance personnel, it is necessary to develop scientific, reasonable, and feasible maintenance strategies and methods based on the system configuration and the control requirements of the production equipment. Preventive maintenance and routine maintenance should be carried out in close coordination, with systematic, planned, and regular maintenance efforts. Any faults that occur during operation should be analyzed on a case-by-case basis. To reduce DCS failures, it is crucial to prioritize prevention and ensure that the system operates properly over the long term in the required environmental conditions.