Thread Content
This post was last edited by xiouxingzhe on 2026-6-11 at 11:48. Seven stages of chemical technology from concept to industrialization (Issue 30/100) —— Technology development: Four quality dimensions of data packets. Dear friends: Hello everyone! In the previous issue, we discussed the overall framework for pilot-scale calibration and data packets, mentioning that the quality of data packets needs to be evaluated from several dimensions, but without going into detail. This issue is dedicated to discussing this topic: how to determine the quality of a data packet. The data packet is the most important deliverable in phase three. By the fourth stage, when the process package is being prepared, every figure in the data package is referenced repeatedly – the heat load values from the equipment data sheet come from here, the control ranges specified in the operation manual also originate from here, and the settings for interlocks are based on values from here as well. If there is a problem with the quality of the data packets, all subsequent work is based on an unreliable foundation. Everyone understands this principle, but when it comes to actual implementation, the judgment of whether the quality is good or not often relies on intuition – things like “it seems to have everything,” “there’s a lot of data,” or “it runs quite stably.” Feelings cannot replace standards. In this issue, the evaluation criteria will be clearly explained. I. Dimension 1: Material balance closure rate. This is the most fundamental and undisputed key indicator of data packet quality. By adding up all the feedstocks used in the entire process, and also adding up all the outputs—products, by-products, waste materials, and unreacted material that is recovered—dividing the difference between these two amounts by the total amount of feedstock gives the unclosed error. The closure rate equals one hundred percent minus this error. What does this number represent? It indicates the extent to which you have an understanding of the whereabouts of the materials throughout the entire process. A closure rate of 98% means that only 2% of the material’s whereabouts are unknown—possibly due to minor leaks, sampling losses, or by-products that were not detected. A closure rate of 90% means that 10% of the materials have an unknown destination – and this proportion is quite high for subsequent engineering design, as you don’t know what will happen to those 10% of materials within the industrial facility. What counts as qualified? It depends on the complexity of the device. For a system with a single reactor and simple separation, the recovery rate should be over 98%, with an excellent standard being over 99%. For complex systems involving multiple reactors, distillation, and recovery, a quality rate of over 96% is considered satisfactory, while a rate of over 98% is considered excellent. For systems involving solid-phase reactions or emissions, over 94% meet the standards, while over 96% are considered excellent. If the closure rate is below the acceptable standard, it is recommended to follow this sequence for troubleshooting: first check the accuracy of the analysis data — test the same batch of samples using different personnel or different methods to see if the results are consistent. Check the representativeness of sampling again – whether the locations of the sampling points are appropriate, and whether the operating conditions during sampling are the same as those during normal operation. Then check for the whereabouts of any materials that have not been accounted for – micro-leaks from safety valves, sampling and discharge, pump seal flushing, and breathing losses from storage tanks; these are often overlooked blind spots. Finally, consider whether there are any unidentified side reactions that are consuming the raw materials—this could be the most important discovery, as it indicates gaps in your understanding of the process, and these gaps may become more apparent when the scale is increased. After completing the material balance sheet, each value must be able to answer three questions: Which instrument does this value come from? When was it harvested? What were the operating conditions at that time? If you cannot answer any one of these three questions, then this data is of unknown origin and cannot be used as a basis for design. It’s not that it can’t be used, but at least it should be clearly indicated whether it is an “estimated value” or a “measured value”, so that those who use this data later can understand its reliability. II. Dimension 2: Integrity of the data traceability chain. This is a dimension that is often overlooked in many pilot-scale data sets. What does a data traceability chain mean? That is, every number in the data table can be traced back to its original source step by step. The complete traceability chain should be: the data in the data table → data collection records (original data files or page numbers of experiment notebooks) → the instrument numbers and calibration records corresponding to the time of collection → the range and accuracy of that instrument → the date and exact time of collection → the operating conditions and the operator at that time. Once this chain is broken, the credibility of the data is compromised. For example, the data sheet states “reaction temperature 180±2°C”. But when you ask further: is this 180 a set value or a measured value? Which gauge is measuring it? When was this instrument last calibrated? If the answer to everything is “can’t remember,” then this data is unreliable. Because the set value and the measured value can differ by several degrees, and instrument drift can also result in a difference of several degrees – behind the number “180” there may be a fairly wide range of actual temperatures. The method for checking the integrity of the data tracking chain is simple: randomly select a few key pieces of data from the data packets and trace back along the chain to see if it is possible to reach the original records. If it gets interrupted halfway through, it indicates a flaw in data management. There are several common reasons for broken trace chains that I have encountered in practice. The instrument tag number is not indicated on the data sheet – it’s not clear which instrument is measuring it. The calibration records are not archived – it is known which instrument it is, but it is unknown whether it is accurate or not. The original data file is missing or not named according to the specifications—the file is found but cannot be opened or does not match. The operating conditions are not recorded synchronously — it is known what the data values are, but it is not known what state the device was in at that time. Each of these issues alone may not seem serious, but taken together, they can severely undermine the availability of data packets. If a significant proportion of the critical data trace chains in a data package are incomplete, those who later work on compiling the process package will face difficulties – they cannot trust this data, yet no other data is available; as a result, they have to increase the safety factors in the design, which leads to the use of larger equipment, higher energy consumption, and inflated investment costs. III. Dimension 3: Sufficiency of operational boundary data. This issue has been mentioned repeatedly in previous issues; here it is discussed systematically from the perspective of packet quality. Many pilot plant data sets share one common problem: they only contain data for the optimal operating point, with no boundary data. Temperature has only one value, pressure has only one value, and load has only one value—these values are obtained under \"optimal operating conditions\", when the device performs at its best. But they do not represent the actual operational capability of the device. For subsequent engineering design and production operations, what is important to know is not the \"best performance\", but rather the \"typical performance\" and the \"worst acceptable performance\". The operation manual needs to specify: at what temperature fluctuations can the product still remain qualified? The operating procedures require knowing: at what load level can it still operate stably? Safety analysis requires knowing: to what extent of deviation will problems occur? Therefore, the adequacy of handling boundary data is an important dimension of packet quality. Specifically, boundary data should cover several aspects. Temperature boundary. Within a range of 5 to 10 degrees around the optimal temperature, a stable operating condition is selected every 2 to 3 degrees, and complete performance data is collected for each one. What is obtained in this way is not an “optimal temperature”, but rather a “temperature-yield curve” that shows you how much the yield is affected by temperature fluctuations of different magnitudes. Load boundary. Starting from 50% of the design load, increasing in steps of 10% or 20% up to 110%, data is collected after each step has been operated stably. What results in this way is not a \"rated capacity\", but rather an \"operational flexibility range\" that shows you how the device performs under low and high loads. Ratio boundary. The feed ratio should be within the range of ±5% to 10% of the optimal value; several points should be selected for testing. Especially when the effect of the ratio on selectivity is significant, boundary data can tell you how large the allowable range for ratio fluctuations is. Key impurity boundary. If certain impurities in the raw material affect the process, different concentrations of such impurities are intentionally added to the raw material for testing. This is not meant to cause \"damage,\" but to find out whether the system can still function properly when the quality of the raw materials varies. In a data packet, if it contains only data from the optimal operating conditions and lacks data at the boundary conditions, then the \"operational boundary adequacy\" of that data packet does not meet the requirements. When preparing the process package later on, many operation parameters can only be estimated based on experience, which affects their accuracy and reliability. IV. Dimension 4: Completeness of scale-up data. The fourth dimension relates to the scale-up data from pilot scale to pilot plant scale. Pilot testing serves as a link between the previous stage and the next – it connects to laboratory testing on one hand, and industrial production on the other. In the pilot test data package, it cannot contain only the data from the pilot test itself; comparative data from the lab tests and the pilot test is also required. This comparison needs to answer several questions. From pilot scale to pilot plant scale, has the conversion rate changed? What is the direction of change? How large is the amplitude? Has the selectivity changed? Has there been a systematic change in the distribution of by-products? By how much has the heat transfer coefficient decreased? Is there a significant difference in the mixing effect? Which parameters remain basically unchanged, and which ones change significantly? These comparative data represent a quantitative description of the amplification effect. With it, when moving on from pilot-scale operations of several hundred liters to industrial plants with a capacity of tens of thousands of tons, you will know where special care is required and where scaling up can be done with relative confidence. Highly complete data on the amplification effect should also include the results of amplification sensitivity experiments – as discussed in issue 28, these experiments involve deliberately changing engineering and physical parameters such as mixing intensity, heat transfer capacity, and residence time distribution in a pilot-scale setup, in order to determine how sensitive the reaction outcomes are to these changes. These experimental data directly serve the selection of scaling design criteria for subsequent industrial plant reactors. The lack of data on the scaling effect is one of the root causes of problems that arise when many pilot-scale projects are scaled up for industrial production. The yield obtained in the pilot scale run was significantly lower than that in the lab scale tests, but no one conducted a systematic analysis to determine why this was the case, at which stage the yield decreased, and whether it would drop further when scaling up to industrial levels. When it was observed on the industrial plant that the yield continued to decline, efforts were made to find the cause, but the pilot plant had already been dismantled. V. The relationship between the four dimensions: These four dimensions are not independent of one another; together they constitute the “quality profile” of the data packet. The material balance closure rate determines whether it is “correct” – that is, whether the data is accurate. The integrity of the data traceability chain is a matter of belief – the reliability of the data cannot be verified. The sufficiency of operational boundary data relates to whether it is \"adequate\" – that is, whether the scope of the data meets the requirements of subsequent design. The data integrity of the scaling effect is a matter of whether it works or not – whether the scaling logic from pilot to pilot-scale is clear. For data packets that meet the standards in all four dimensions, it becomes much easier to move on to the fourth stage of developing the process package – each parameter has a source, defined boundaries, and comparative values; the selection of equipment and the setting of operational parameters are based on solid criteria. A data packet with a significant weakness in one of the four dimensions will encounter problems in subsequent stages – insufficient closure rate, incorrect material balance, and inaccurate design basis ; The traceability chain is broken, the credibility of the data is in doubt; therefore, it is necessary to increase the design margin ; The boundary data is incomplete; it’s hard to determine the flexibility of operations ; Data on the amplification effect is missing; the amplification design is based on guesswork. Spending an extra week or two during the pilot phase to complete the four dimensions of the data package is much more cost-effective than having to spend another month or two in the subsequent design phase to fill in that information later. This principle is understandable in theory, but under the pressure of project deadlines, the quality of data packets is often the first to suffer. My experience is that pilot testing can save time in other areas, but quality of data cannot be compromised. Preview for the next issue: Issue 31 – Common pitfalls in the technology development phase. We have spent many issues discussing the third phase; from concept validation to pilot-scale data sets, there are several pitfalls that can easily be made in each of these stages. Next time, we will provide a summary of the most common pitfalls in the technology development phase: premature optimization, blind scaling, and issues that are easily overlooked in data management. After this episode, phase three is coming to an end.