Smart manufacturing can make production more visible, connected and responsive. It can also introduce a new kind of headache.
A machine is running, but the dashboard says it is stopped. Production quantity in MES does not match the PLC counter. A sensor suddenly starts reporting strange values. ERP shows an order as completed while the shop floor says otherwise.
These problems are frustrating because the fault may not be inside the machine itself. It could be a sensor, PLC program, network connection, gateway, database, integration rule or software configuration.
NIST notes that smart manufacturing systems contain complex interactions between systems, subsystems and components, making it difficult to determine exactly what is influencing process performance or data integrity.
The best approach is therefore not to guess.
Use a structured process:
- Identify
- Isolate
- Verify
- Correct
- Test
- Monitor
Start by Separating the Symptom From the Cause
Suppose a dashboard shows:
Machine 04: Stopped
The immediate assumption may be that the machine has stopped.
But several things could actually be happening:
- the machine is physically stopped;
- the PLC stopped sending data;
- the gateway lost communication;
- the network connection failed;
- the machine-state logic is incorrect;
- the dashboard has stale data.
This is an important troubleshooting habit:
Never treat the displayed symptom as proof of the root cause.
Check the problem from the physical process upward.
A Simple Smart Manufacturing Troubleshooting Order
A practical troubleshooting sequence is:
Check |
Question |
|---|---|
Physical process |
Is the machine actually operating correctly? |
Sensor |
Is the measurement believable? |
PLC/controller |
Is the correct value available? |
Communication |
Is the data reaching the next system? |
Gateway/edge |
Is data being translated correctly? |
MES/SCADA |
Is the value interpreted correctly? |
ERP/business system |
Is the transaction synchronized? |
Dashboard |
Is the correct information displayed? |
Working through these layers can save a lot of time compared with changing settings randomly.
Problem 1: Machine Data Suddenly Disappears
One of the most common problems is a connected machine that suddenly stops appearing in the monitoring system.
Possible Causes
- network cable or switch problem;
- machine controller offline;
- gateway failure;
- changed IP address;
- firewall rule;
- expired certificate;
- software service stopped;
- protocol configuration changed.
How to Troubleshoot
First confirm whether the machine itself is operating.
Then check the data chain:
- Machine
- PLC
- Network
- Gateway
- Server
- Application
Find the first point where data stops appearing.
If the PLC has the correct information but the gateway does not, concentrate on communication between those two systems rather than changing the dashboard.
For OPC UA systems, clients should examine the StatusCode returned with results. The OPC Foundation specifies that failed or “Bad” results should not be used, while uncertain results require care.
This is far more useful than simply seeing “No Data” on a screen.
Problem 2: Sensor Values Look Wrong
Imagine a motor normally operating around a stable temperature range suddenly reporting an impossible or highly unusual value.
Do not immediately conclude that the motor is failing.
The problem might be:
- damaged sensor;
- loose connection;
- incorrect scaling;
- wrong engineering unit;
- sensor calibration drift;
- PLC conversion error;
- wrong tag mapping.
Start with the physical measurement.
Compare the digital value with another trusted measurement where it is safe and appropriate to do so.
Then trace the value:
- Sensor
- Input Module
- PLC
- Gateway
- Database
- Dashboard
If the sensor reports 50°C, the PLC shows 50°C but the dashboard shows 500°C, there is little reason to replace the sensor.
The fault is somewhere later in the data chain.
NIST’s guidance on manufacturing data emphasizes proper collection, curation and use of standards when building shop-floor data systems.
Problem 3: Production Counts Do Not Match
This one can start surprisingly heated discussions between production and IT.
The PLC says:
1,050 parts
MES reports:
1,012 parts
ERP shows:
1,000 parts
Which number is correct?
Before changing anything, understand what each number means.
The PLC may count every machine cycle.
MES may count only completed units.
ERP may count only confirmed good products.
The disagreement may therefore be caused by different definitions rather than a technical failure.
Check:
- when the counter increments;
- whether rejected parts are included;
- whether reworked products are counted;
- whether counters reset between orders;
- how batch boundaries are handled;
- whether messages were lost;
- whether the same product was counted twice.
ISA-95 exists partly to standardize information exchange between manufacturing-control and enterprise systems and reduce errors associated with integration.
Before fixing the number, make sure everyone agrees on what the number is supposed to represent.
Problem 4: The Dashboard Does Not Match the Shop Floor
Dashboards sometimes become the first thing people blame.
Often, however, the visualization is only showing the information it receives.
Suppose the dashboard reports:
Machine running for 420 minutes
The operator says:
“That cannot be right. We had three stoppages this morning.”
Check:
- machine-state definitions;
- timestamps;
- data refresh frequency;
- communication interruptions;
- calculation logic;
- downtime classification;
- shift start and end times.
Also check whether the dashboard is displaying live data or cached historical information.
The rule here is simple:
Verify the source before modifying the visualization.
A beautiful dashboard displaying incorrect data is worse than having no dashboard at all because people may make decisions based on it.
Problem 5: MES and ERP Are Out of Sync
Smart manufacturing relies heavily on integration between production and business systems.
For example:
- ERP
- Production Order
- MES
- Machine
- Production Result
- MES
- ERP
Problems may appear when:
- order numbers do not match;
- product master data differs;
- messages fail;
- transactions arrive twice;
- an interface service stops;
- an operator closes an order incorrectly;
- systems use different status definitions.
Start with one affected production order.
Trace it from beginning to end instead of looking at thousands of transactions.
Ask:
- Did ERP create the order?
- Did MES receive it?
- Did the correct machine receive the job?
- Was production completed?
- Did MES record completion?
- Did ERP receive the confirmation?
ISA-95 specifically defines models and information exchanges between manufacturing operations and enterprise functions, making it a useful reference for structuring these interfaces.
Problem 6: Too Many Alarms
Another common problem is not missing information but having too much of it.
If operators receive hundreds of alarms every shift, eventually many of them will be ignored.
Review alarms for:
- frequency;
- severity;
- duplication;
- duration;
- actual operational impact.
Separate them into categories such as:
- Critical
- Warning
- Information
A repeatedly occurring alarm should not simply be acknowledged forever.
Ask why it keeps occurring.
It could indicate:
- poor threshold settings;
- unstable process conditions;
- sensor problems;
- incorrect logic;
- genuine equipment deterioration.
Good smart manufacturing should reduce uncertainty, not create constant notification noise.
Problem 7: Data Is Available but Nobody Trusts It
This is one of the most dangerous smart manufacturing problems because the technology may technically be working.
Operators simply stop using it.
Common reasons include:
- inaccurate downtime data;
- inconsistent equipment names;
- incorrect production counts;
- unexplained calculations;
- excessive alarms;
- dashboards designed without shop-floor input.
NIST’s work on smart manufacturing highlights the importance of trustworthy and traceable manufacturing data for making reliable decisions.
When users report that a system is wrong, do not dismiss the feedback.
Compare:
- System record
- Actual event
- Operator knowledge
Sometimes the software is correct.
Sometimes the operator is correct.
Troubleshooting determines which.
Problem 8: Network or Cybersecurity Changes Disrupt Production
Not every communication failure is an equipment problem.
A firewall change, network segmentation project, certificate update or remote-access configuration can interrupt industrial communication.
Before making changes to OT networks:
- document required communication paths;
- coordinate with production and automation teams;
- maintain backups;
- test changes;
- have a rollback procedure;
- review logs after implementation.
NIST SP 800-82 Rev. 3 recommends network segmentation as part of defense-in-depth for OT environments and notes that organizations should understand required communication flows when configuring network isolation.
Do not bypass cybersecurity controls simply to make communication work again. Identify what authorized connection is actually required.
Use the “Last Known Good Point” Method
When a problem becomes confusing, find the last place where the information is correct.
For example:
- Sensor: Correct
- PLC: Correct
- Gateway: Correct
- Database: Wrong
- Dashboard: Wrong
You have immediately reduced the search area.
The fault probably exists between the gateway and the database.
This approach is simple, but on a complex production system it can save a great deal of unnecessary work.
Keep a Troubleshooting Record
Recurring problems should be documented.
Record:
- problem;
- affected equipment;
- date and time;
- symptoms;
- root cause;
- corrective action;
- person responsible;
- verification result.
Over time, this creates a useful troubleshooting knowledge base.
If the same communication failure occurs every month, the goal should not be to become faster at restarting the gateway.
The real goal should be to understand why the gateway keeps failing.
Conclusion
Troubleshooting smart manufacturing does not require guessing which technology is responsible.
Use a disciplined sequence:
- Confirm the physical process
- Check the sensor
- Check the controller
- Check communication
- Check integration
- Check the application
- Verify the result
And remember one practical rule:
Find the first point where good information becomes bad information.
That point is usually much closer to the real problem than the error displayed on the final dashboard.
Smart manufacturing systems are complex, but troubleshooting them becomes manageable when engineers stop looking at the whole digital factory at once and isolate the problem one layer at a time.