The full power of TDengine, now free forever for up to 5,000 tags.

Explore More

Battery Manufacturing: Tracing Cell Quality and Safety Risks

Jim Fan

September 5, 2026 /

As the global energy storage and renewables industries expand rapidly, the consistency and service life of prismatic lithium iron phosphate (LFP) energy storage cells have become key measures of a manufacturer’s core competitiveness. Yet cell manufacturing is a complex, multi-stage physical and electrochemical coupling process. Small mechanical deviations in the earlier electrode-making stages (coating, calendering) cascade and amplify downstream through winding and formation, eventually causing cell capacity degradation, rising internal resistance, and even serious safety hazards such as thermal runaway.

An energy storage system is built from hundreds to thousands of cells wired in series and parallel, so it places extremely strict demands on cell consistency. Key process indicators need to be held within very narrow ranges: areal density deviation (≤ ±1.0%), winding alignment deviation (≤ 0.2 mm), and formation internal resistance consistency (≤ ±3.0%). Any slight misalignment in an earlier process stage cascades through the chain into a performance or safety defect in the final product.

Even under such high-precision manufacturing requirements, the workshop has deployed large numbers of sensors and SCADA systems that generate millions of high-frequency time-series records every day, yet the flood of data has not removed the pain points in operations. Enterprises face a new “data gap”: plenty of data but little operational context, so it is hard to understand quickly; plenty of systems but they are isolated from one another, so it is hard to close the troubleshooting loop; plenty of alarms but little correlation among them, so they cannot directly support on-site decisions. Turning the scattered raw measurement points across process stages and systems into operational assets that frontline staff can understand, analyze, and act on is the core challenge facing quality management today.

The business problem: what is really blocking operations

In the day-to-day quality control of lithium battery production, even when the line is covered with sensors, frontline process engineers and quality inspectors still hit several bottlenecks they struggle to get past when tracking down and preventing quality problems.

Startup, shutdown, and handover naturally bring parameter jitter, and static threshold alarms trigger a flood of “crying wolf” false alarms

During electrode coating or calendering, when equipment starts up and threads the web, when one electrode roll joins the next, or during the daily shift-change speed reduction (speed is normally reduced by about 15% within the 10 minutes around each shift change at 08:00, 16:00, and 00:00), measured indicators such as unwind tension and roll pressure naturally show brief physical fluctuation or sag. With traditional static threshold alarms, these normal operating transitions trigger a flood of false alarms that drown out genuine equipment faults. Over time, frontline staff grow alarm-fatigued and tend to ignore them, so process drift truly caused by a drive-motor anomaly or a hydraulic bearing fault goes unreported, planting hazards for downstream production.

Final inspection catches performance degradation, but the cross-process tracing chain is too long

Once the quality center (QA-01) finds a batch of cells with low capacity in capacity grading (for example, down to 46.2 Ah), or with abnormally high outgoing internal resistance (deteriorated to 2.30 mΩ), the batch is already fully produced and the defects already exist. Because a production batch (for example B-090 ~ B-098) spans multiple physical locations and different control systems across coating, calendering, winding, and formation, the quality analyst has to log in to the MES, SCADA, and quality systems separately, export several CSV data tables, and manually align them in Excel by timestamp and web-roll change node. This kind of cross-process tracing is extremely tedious, and pinning down the root cause often takes hours or even days. While the search continues, subsequent batches keep running, so losses from batch rejection can quickly mount.

Environmental and safety monitoring is isolated from the production process systems, making it hard to respond to abnormal leaks in a linked way

When a cell in formation is charged at high current, improper process conditions (for example, over-pressing of the upstream electrode that leaves porosity extremely low) can easily cause lithium plating on the negative electrode surface. The metallic lithium dendrites grown by lithium plating can pierce the separator and cause an internal micro-short-circuit, which rapidly heats the cell and makes its internal pressure relief valve open and close slightly, releasing DMC and other electrolyte solvent vapors. The workshop’s environmental monitoring sensors (UT-01) can detect the sharp rise in hazardous VOC concentration (for example, up to 35.8 ppm) and trigger a leak alarm, but because the safety and environmental system is physically isolated from the production process control systems, the safety officer cannot tell which cell in which formation cabinet the leak comes from, cannot link it back to which calender or coater upstream went wrong, and cannot execute a targeted shutdown. The risk of thermal runaway is extremely high.

The solution: a closed loop from data to decision

To fully solve these pain points, adding more single-function systems is not the answer. What is needed is a closed-loop capability that runs the whole chain from data organization and business understanding to analytical inference and safety decisions.

A unified data foundation organizes scattered data into operational object views centered on batch numbers

The key to making the data comprehensible is to give the unordered sensor time-series measurement points clear “operational context.” The approach is to break the barriers between traditional process stages and systems and build a unified asset model at the foundation layer. By defining the roughly 1200-meter run of an electrode web as a unique production batch number (B-001 ~ B-168), the high-frequency sampled data in the supertables for coating (stb_coater), calendering (stb_calender), winding (stb_winder), formation (stb_formation), environment (stb_utility), and quality (stb_quality) all maps naturally onto a single batch dimension, turning raw measurement points into structured operational objects carrying the context of working condition, process stage, and batch.

From passive monitoring to active inference: building the closed loop of monitoring, analysis, events, and tracing

A closed-loop capability means the system is responsible not only for “looking” but also for “explaining” and “following through.”

  • Dynamic multi-dimensional monitoring. Working-condition state filtering (for example, activating precision detection only during the full-load status=2 periods) removes the physical noise of commissioning and web-change phases and delivers more precise proactive alarms.
  • Cascade root-cause tracing driven by physics. Using cross-device physical coupling relationships (for example, the chains from coating tension to guiding deviation to alignment deviation, or from roll pressure to electrode thickness to formation temperature rise to VOC concentration), tracing starts when an alarm triggers, and the anomaly moment is captured as a business event whose root cause can be located, with suggested actions to support the decision.

Lower the barrier and push results: let more roles get data insights on their own

Deep tracing in the past relied heavily on IT staff exporting data and data experts writing code to analyze it. The new approach should provide code-free multi-dimensional data analysis dashboards backed by a proactive push mechanism. Whether a process engineer is analyzing equipment stability, a quality inspector is checking defect trends, or a safety officer is confirming a workshop leak, they can all reconstruct the physical process behind an anomaly on their own in intuitive linked comparison charts through drag-and-drop or conditional filtering, so data insights truly serve frontline production decisions.

Scenarios in practice: how the operating workflow changes

In the full-process monitoring practice at the energy storage battery plant’s workshop No. 2, this new closed-loop management approach is reshaping day-to-day work in concrete ways.

The traditional troubleshooting model: experience-driven and time-consuming, trapped in information silos

In the past, when the quality center (QA-01) finished capacity grading and found that batch B-092 cells had a capacity of only 46.2 Ah (far below the 52.2 Ah golden baseline, a serious failure) and an outgoing internal resistance of 2.30 mΩ (normal is 1.15 mΩ), an urgent troubleshooting phone call would go out.

  1. Time-consuming file assembling. The quality inspector notifies the formation shift, the formation shift pulls the channel records of cabinet FC-01, and confirms that the batch ended constant-current charging too early, with genuine internal polarization and internal resistance as high as 2.45 mΩ.
  2. Blind shuttling back to upstream stages. Unable to quickly tell whether this was poor hardware contact in the formation cabinet or an upstream electrode problem, the engineer has to go to the winding workshop to pull WI-01 winder historical alignment deviation data, then cross the plant to the coating workshop to export CO-01 take-up tension logs.
  3. Closing by experience. The investigation consumes most of the day. Because the systems’ clocks disagree, the trend charts the engineer assembles in Excel are fuzzy and show no clear propagation logic. In the end, the coating stage cannot confirm its own take-up tension was at fault, the winding stage attributes it to equipment yaw, and quality analysis has to compromise: about 500 cells are rejected as a batch, with direct economic losses of about US$10,000, and the root cause is never fully eliminated.

The new way of working: linked operational objects and tracing organized around production batches

After the new data management platform was introduced, all the time-series data and business tags formed a closed loop, and two typical anomaly scenarios are now handled completely differently in the workshop.

Scenario 1: a two-minute diagnosis of the quality anomaly caused by coating tension drift

When cells of batch B-092 trigger the Critical quality gate alarm “cell capacity < 48.0 Ah” in the capacity grading test at quality center QA-01, the process engineer immediately clicks “root-cause traceback” on the anomaly event:

  • Step 1 (formation confirmation). The interface links to the same batch’s data at formation cabinet FC-01: AC internal resistance is 2.45 mΩ, confirming electrochemical polarization.
  • Step 2 (winding traceback). Clicking the winding level pulls the batch’s time series at winder WI-01, which shows that around 12:20 the positive and negative electrode “alignment deviation” had worsened to 0.88 mm (past the 0.5 mm Major threshold, versus only 0.05 mm under the golden baseline).
  • Step 3 (coating root cause). Continuing up to coater CO-01, the multi-dimensional comparison dashboard draws the take-up curve around 12:12: the coater’s “take-up tension” is as high as 220 N (normal is 130 N), the “coating thickness” has fallen from a normal 155 μm to 142 μm, and the “coating edge deviation” has risen from 0.1 mm to 0.8 mm.

The system’s physical coupling algorithm concludes instantly: the take-up drive control system of the coater has drifted in its closed-loop control parameters, and the tension overload stretched and thinned the half-dry electrode web, which buckled into a serpentine shape when tension released on the downstream winding shaft, causing alignment failure and a jump in internal resistance. The whole investigation takes only 2 minutes in a single view. The engineer immediately issues a calibration command, subsequent batches recover quickly starting at B-096, and the continued rejection of several following batches is avoided, saving tens of thousands of dollars in potential losses.

Scenario 2: a coordinated safety response to a VOC leak associated with overpressure during lithium plating and micro-short-circuit risk

In another scenario, the workshop’s utility environmental monitoring system UT-01 suddenly raises a “VOC concentration exceeded” alarm (the reading jumps from a normal 1.2 ppm to 35.8 ppm in an instant, past the 10.0 ppm Critical safety limit). Under the new system, this is no longer an isolated air-quality event:

  • Step 1 (thermal runaway location). The safety system reads the formation supertable and finds that formation cabinet FC-02 measured a cell temperature as high as 55.4 °C in the same window (it triggered the temperature alarm above 45.0 °C, while under the golden baseline the maximum is only 30.5 °C). The system judges that the cell gassed and expanded from micro-short-circuit heating, making the pressure relief valve open and close slightly and causing the leak.
  • Step 2 (calender root-cause traceback). Clicking the batch’s electrode calendering history, the time series at calender CA-01 shows: the right-side “roll pressure” is as high as 820 kN (normal is 600 kN), which snapped the right-side roll gap shut to 8.0 μm, and the rolled “electrode thickness” is only 111.0 μm (normal is 122.5 μm).

The engineer immediately reaches a physical mechanism conclusion: the spool of CA-01’s hydraulic control valve is stuck in its port, so the roll pressure spiked and the electrode was over-pressed. The sharp drop in porosity blocked diffusion of the electrode phases during high-current formation charging, causing lithium plating that pierced the separator and triggered a micro-short-circuit. The workshop dispatcher immediately orders the on-site maintenance crew to manually relieve pressure and replace the spool, eliminating the risk of fire and continued leaking.

A refined monitoring strategy: combining resident alarms with deep tracing

In practice, the workshop has built a rational alarm configuration strategy. Why does daily monitoring keep only the “tension anomaly alarm,” “cell capacity,” “temperature exceeded alarm,” and “VOC concentration exceeded alarm” as resident monitors, while “alignment deviation” and “electrode thickness” sit in the second layer for tracing?

This is because:

  1. Remove motion noise. Winding web-guiding action is extremely frequent, and during web-change acceleration and threading (status=1), the alignment deviation measured by the CCD naturally jumps for short periods. Set as a resident alarm, it would fire frequent nuisance alarms during web changes.
  2. Protect the probe calibration state. During non-production periods, the beta-ray thickness gauge retracts its measuring head for calibration, so its thickness reading naturally goes to zero or produces invalid NULL values. A resident detection rule at the first layer would generate a large number of invalid false alarms.

With this refined strategy of “few resident alarms, full tracing of anomalies,” workshop No. 2 cut the invalid alarm rate by 85%, achieving genuine precision prevention and agile troubleshooting.

Lithium battery quality monitoring: the last step from data ingestion to decision-making

The essence of digital transformation in lithium battery manufacturing is not how many sensors the workshop deploys, how far the data can be compressed, or how fast it can be written into a database. A truly valuable industrial data platform is one that organically links physical processes, electrochemical behavior, and the safety environment along the dimension of operational objects, removing the gap where data is visible but the process behind it is not.

By managing process indicators, outgoing quality, and the on-site environment in a closed loop, the cell workshop has moved from passive “after-the-fact quality settlement” to context-aware “real-time anomaly tracing and process error prevention.”

This is where TDengine IDMP (Industrial Data Management Platform) shows its core value. It moves beyond treating industrial databases as only ingestion and storage systems and is better understood as a data management and insight platform for industrial scenarios: it organizes data around physical assets and operational objects, and through automatic scenario awareness (for example, distinguishing commissioning from full-load working conditions), real-time analysis, process tracing, and AI-assisted insight, it lets process engineers, quality inspectors, and safety officers all talk to the data through a unified view.

Looking ahead, as the platform is rolled out across the group’s plants, time-series analysis built on physical coupling formulas will merge further with AI large language models. Then workshop technicians will be able to ask directly in natural language: “Why is the internal resistance of batch B-140 cells high?” The system will pull the cross-process time-series trends and produce a root-cause report, completing the last step from data ingestion to generated decision.

TDengine comes with a high-performance, distributed time-series database, Industrial Ontology modeling, and an Industrial Agent Runtime, providing a full-stack solution for industrial data streams from collection and storage to real-time analytics, visualization, event management, and root-cause analysis. To learn more about TDengine, visit www.tdengine.com and try it for free.

Try it yourself

Install and deploy TDengine Visit the TDengine Download Center, select TDengine All-in-One, choose the deployment platform and architecture that matches your environment, and follow the guided steps to complete the installation.

Load the sample data

On first activation, choose Lithium Battery Full-Process Quality Monitoring on the sample data loading screen and wait for it to finish loading.

If you have already activated the product, click your avatar in the top-right corner, select Management Console, choose Sample Data on the left, then select Lithium Battery Full-Process Quality Monitoring. Wait a few minutes for the data to load.

Lithium Battery Full-Process Quality Monitoring