HCBBS Forum (English)
Submit Chemical Projects / Find Solutions
Amplify Your Requirements on a Broader Chemical Platform *Engineering · Technology · Equipment · Solutions*
Submit Request

Protection and troubleshooting methods for disk failure in industrial computers

2008-03-03View Original

Thread Content

I. Introduction With the rapid development of industry and computers, and in an era of increasing automation, computers are now used in every aspect of automated control systems. The security of industrial control computers (hereinafter referred to as ICS) is also particularly important. Industrial computers share the same technical principles as ordinary computers, and their structural composition is also similar; however, what matters most for industrial computers is their operational stability. Industrial computers generally operate in relatively harsh environments, and they have high requirements regarding environmental temperature, humidity, power supply and voltage stability, as well as ventilation. However, the working conditions often fail to meet these requirements, which makes industrial computers prone to malfunctions. While some hardware components can be replaced promptly if they develop problems, damage to the disks can result in the loss of a large amount of data, as well as damage to the control software; such issues cannot be repaired quickly, leading to unstable control of the parameters being monitored and often causing significant economic losses. II. Description of the fault phenomenon: After the industrial computer has been in operation for an extended period of time – by \"extended period\" is meant a time span of 30 days at 24 hours per day, i.e., one standard month or more of continuous operation – a large amount of dust accumulates inside the chassis, resulting in high temperatures within it. Normally, everything works properly without shutting down the system; however, in cases of insufficient power supply or when an emergency shutdown is required, the control system is prone to issues such as disks failing to start, the system failing to load, or remaining stuck on the login screen for extended periods. Taking the 10 operator stations and 4 engineer stations on the Langfang High-Tech Glass production line of Shandong Glass Group as an example: System configuration – Industrial PCs: Advantech industrial PCs, DELL GX270; Operating systems: Windows 2000 Professional, Windows 98 (one unit) (genuine versions); Control software: Citect, Freelance 2000, WCC5.0, SETP7 5.2 and other genuine software; Auxiliary software: WINRAR 3.0, Windows 2000 Professional SP4 patches, etc.; Working hours: Full-time operation throughout the year (24/7, 365 days a year); Working environment: The ambient temperature is controlled between 10 degrees Celsius and 30 degrees Celsius using air conditioning; there is slight mechanical vibration on the floor, and the air contains inhalable particles. Air humidity: 5%–50% RH. Since it was put into operation in December 2003, three computers have experienced a total of five instances of disk errors that prevented them from starting up. To date, the author has not received a reliable response from Microsoft’s Operating System Services department. III. Fault Analysis and Troubleshooting Methods There are numerous reasons that can cause disk failures; here we roughly categorize them into quality issues with the disk itself and faults resulting from the operating environment. We are unable to conduct a thorough investigation into the quality issues of the disk itself; the only thing that can be done is to choose hard drives of good quality and from reputable brands when developing control systems. Software such as Scandisk and Norton Disk Doctor can also be used to detect defects on the disk surface. If we could predict the quality and health status of hard drives, it would give us time to choose the right hard drives and back up important data. The author found online a software called Drive Health, which can determine the remaining lifespan of a hard drive and help people understand its health status in advance. Common failure issues caused by the working environment include the following: 1. The industrial computer operates for long periods of time. Due to the requirements of normal production, the industrial control systems in some factories need to operate for long periods of time, posing a significant challenge to the operating systems of these industrial computers. According to Microsoft’s operating system uptime reports, the company claims that its operating systems released after Windows 2000 can handle long periods of operation. However, in practice, after more than a week of continuous use, large amounts of data fragments accumulate on the disk as a result of extensive data transfers, which can lead to logical disk errors, read/write failures, and slower system performance and startup times. Therefore, when production permits, the industrial computer can be restarted periodically, along with disk defragmentation, in order to reduce disk errors caused by long periods of operation. The restart time can be determined based on the amount of data processed by the industrial computer and the production conditions; it is not fixed, and readers need to figure it out gradually. Based on the author’s practical experience, restarting and organizing the industrial computer once per standard month (30 days) can reduce the likelihood of disk errors. 2. The internal temperature of the industrial computer is too high. In environments that require long-term operation at high temperatures, various components of a computer are prone to aging, and the frequency of hard drive failures also increases. This requires the maintenance personnel of the factory’s automation systems to pay close attention to the temperature of the chassis during regular inspections, striving to keep the temperature of the industrial computers between 10 and 30 degrees Celsius. Temperatures that are too high or too low are not conducive to protecting the hard drives; if the chassis temperature reaches 30 degrees Celsius, the temperature of the hard drives inside can reach 40 degrees Celsius or higher. We can simply run a DIR command on our industrial computer to help reduce the ambient temperature. 1. Replace the high-power CPU and hard drive fans (make sure the hard drive fans are properly secured; they should not be installed on the hard drive brackets to prevent vibration of the hard drive caused by fan rotation), in order to improve heat dissipation ; II. Install fans inside the chassis that draw air out of it, thereby increasing air convection ; III. Install small axial flow fans on the cabinet where the industrial computer is placed ; IV. Install air conditioning in the control room to reduce the temperature inside. 3. The environmental humidity is not suitable. Industrial computers are primarily composed of integrated circuits containing numerous electronic components, and their insulation performance is highly dependent on environmental humidity. Excessive humidity can easily cause short circuits in the circuit board, leading to its destruction ; Too low humidity can lead to static electricity, which may also damage certain electronic components. Therefore, either excessive or insufficient humidity can pose potential threats to industrial computers. Regarding electrostatic protection, it is required that industrial computers must have proper instrument grounding. It should be noted that the grounding electrode for industrial computers is different from the lightning protection grounding used in civil engineering; the location of this grounding electrode should be three meters away from the control room. It should be installed 1700 mm below the outdoor ground surface, using ∮20 galvanized angle steel as a vertical grounding electrode. The number of grounding electrodes must ensure that the grounding resistance is less than 1 ohm (a megohmmeter should be used for testing during backfilling). Then, 40404 galvanized flat steel should be welded securely to these grounding electrodes (each welding point also requires thorough anti-rust treatment). From there, 25mm copper cables are used to connect to the system’s ground terminals and the grounding points of the industrial computers. This can effectively reduce the hazards caused by static electricity. 4. Strong ground shaking. In many factory operations, motors are needed to generate physical movements such as traction and vibration, which not only produce significant noise but also cause the vibrations resulting from machine operation to inflict serious damage on the hard drives, optical drives, and floppy drives in industrial computers. The manufacturing processes for disks are becoming increasingly advanced, with current rotation speeds reaching 7200 revolutions per second or even higher. In the extensive data exchange that takes place in automated control systems, disks that operate for long periods at high speeds are prone to having their read and write capabilities reduced due to disk vibration, as well as experiencing slow head positioning; in severe cases, this can even lead to disk damage ; Therefore, reducing vibration in the industrial computer environment helps to protect the disks. During engineering design, we can try our best to keep the industrial computer away from work sites with high levels of vibration ; If the work location cannot be changed, we can also place sponge or other cushioning materials under the industrial control cabinets and enclosures to reduce the damage caused by vibrations. 5. There are many inhalable particles in the air. In many factories, raw materials often need to be processed in powdered form. Coupled with strong air currents and abundant dust in the surrounding environment, industrial computers are prone to accumulating large amounts of sticky dust, which can lead to excessively high temperatures inside the computers and cause hardware damage. This situation often occurs around the cooling fans of components such as CPUs, power supplies, hard drives, and graphics cards. In areas with less dust accumulation, scheduled dust blowing can be employed if it is permitted under normal production conditions. In areas with heavy dust accumulation, dust-filtering gauze can be placed at the ventilation openings of the industrial computer chassis for regular cleaning. 6. The power supply voltage fluctuates greatly, resulting in frequent power outages. With the rapid development of industry and daily life, the demand for electricity is increasing steadily. In many areas, there are problems such as insufficient power supply, unstable voltage, and frequent power outages. Unstable voltage and sudden power outages cause the system to restart frequently, and system files can be lost as a result, preventing the system from starting properly ; The head that is performing read/write operations may, due to a power outage, fail to return to its proper position, which can cause disk failures in industrial computers. Therefore, the stability of the power supply in the operating environment of industrial computers is crucial for their proper functioning. We can use regulated power supplies and UPS systems for protection; the specific equipment to be selected depends on the power level of the load and the duration for which power supply is required. IV. Emergency troubleshooting strategies: Often, despite all the protective measures taken by our industrial control professionals, disk failures still occur in industrial computers. Below, we will discuss with readers how to take preventive actions before such failures happen. Readers are advised to first learn how to use GHOST (a well-known disk cloning software), using the latest version possible, as this will facilitate the implementation of the following solutions. Solution that requires no financial investment: GHOST cloning. Prerequisite for this solution: Only the system drive is faulty, and it can be formatted properly using FORMAT software. (The author encountered two instances in which it was not possible to format the system drive using the FORMAT software; when accessing the system drive on a master-slave setup, an error message indicating incorrect parameters appeared, and the issue was resolved by performing a low-level formatting.) Required tools: GHOST software, DOS system boot disk (available as a CD, software, or USB drive). Implementation steps: Disk failures often occur on the system drive (Drive C). After the automated system is put into operation, GHOST software is used to create an image of the system drive, with the resulting image file (.GHO) being saved in FAT32 format as a backup. This is because, when restoring the system drive using GHOST on a single disk, it is usually done under DOS, which can only run on disk formats such as FAT32 and FAT16; it cannot run on NTFS formatted disks. ), once the system drive fails, the fastest way is to format it; using GHOST software, the originally backed up files can be restored to the system drive in about 5 minutes. Conclusion of the plan: No equipment investment required, no financial expenses needed ; Fast recovery speed. This solution can only be used in cases of damaged operating systems, rather than physical damage to the disk ; Once the disk is physically damaged, this solution will not work. This can also be extended to the entire disk image. Economic solution: Cloning of two hard drives for backup + GHOST imaging. Prerequisites for this solution: Disk failure in the industrial computer (whether due to system issues or physical damage to the disk). Materials required: A hard drive of the same model as that in the target industrial computer, GHOST software, and a DOS system boot disk (which can be in CD form, as software, or on a USB drive). Implementation steps: Before the industrial computer’s system starts operating, use GHOST software to create an image of the system drive (Drive C), which contains the control system data, onto a non-system drive formatted in FAT32. Then, clone the entire content of that hard drive to a spare hard drive of the same model. Once the operating system is damaged, the operating system image file can be restored ; In the event of a failure of the entire disk, the faulty disk can be removed and replaced with a spare hard drive that contains an identical copy of the data. Conclusion of the solution: Only the cost of a single disk is required (around 400–800 RMB, depending on the disk size and manufacturer); it is easy to replace, recovery is fast, and it can resolve all disk-related issues. Security investment plan (budget-friendly): Use a software-based disk array with Windows 2000, in RAID1 configuration with two disks or RAID5 configuration with three disks. Prerequisite for this plan: It is best to implement it before the industrial computer starts operating. Required materials: Windows 2000 system disk or a higher version, two disks (of the same model is preferred). Implementation steps: First, install Windows 2000 or a higher version (as Microsoft provides better support for disk arrays after Windows NT), and then enable the disk array functionality. The disk array approach allows write operations to be performed on an industrial computer by writing the same data to two disks at the same time; if one of these disks fails, the other disk, having received the identical data via simultaneous writing, can switch to normal operation without any disruption. In other words, in a disk array setup, as long as one of the two hard drives is not damaged, the important data will not be lost. The faulty disk can be replaced or repaired later; the greatest advantage is that it prevents any loss of production data, and the replacement process is also fast. Plan conclusion: Low cost, high safety; stability depends on the quality of system installation and settings, but it is difficult for beginners in technology to master. Security investment plan (stable type): This approach uses low-end server hardware with disk arrays (suitable for environments where technical requirements are low, stable operation is essential, and data is of great importance). Materials required for this plan: One low-end server with hardware disk array capabilities costs around 25,000 RMB. Implementation: Since a hardware-based disk array is used, it is less susceptible to external interference, resulting in a lower incidence of failures. In the event of a disk failure, it is sufficient to remove the faulty hard drive and replace it with a new one of the same model. For specific implementation details, please consult low-end server providers. Plan conclusion: Relatively high investment, high safety, good stability, and low technical requirements. V. Conclusion The hazards caused by disk failures in industrial computers are self-evident; ensuring their security is an issue that industrial computer professionals should pay attention to. Only by taking thorough preventive measures in advance can industrial computers operate stably and efficiently. The author has applied the above method in actual work, effectively preventing disk failures in industrial computers; this has reduced the time required to repair such failures from several hours to just a few minutes, thus ensuring uninterrupted production.

Submit a Project

**Looking for Chemical Technology, Equipment & Solutions?** No Registration Required Broader Platform Exposure | Global Chemical Service Provider Connections

Submit Request — Free Consultation

Disclaimer

This is an automated machine translation of the original thread. Some technical terms may have inaccuracies; the original text shall prevail. Click "View Original" at the top right to access the source page, which supports IP-based automatic real-time language translation. Please watch out for contact details and sales inducements to prevent fraud. All content and translations are for reference only, representing solely the poster's personal views. For enquiries, email service@hcbbs.com.