Skip to content
Back to Resources
Guide

Manufacturing Process Analytics Without a Data Warehouse

Skopx Team
July 31, 2026
15 min read

A plant manager at a mid sized injection moulding operation has two problems on the same Tuesday morning, and they are not the same kind of problem.

The first: cavity pressure on press seven has been drifting for nine shifts, short shots are up, and the process engineer wants to know whether the drift correlates with the resin lot change or with the mould temperature controller that was swapped out. That question needs sub second sensor data, aligned to shift and lot, held for years, queried against a golden batch profile.

The second: a resin supplier emailed on Friday to say a shipment slipped eleven days, three work orders consume that resin, one of them is against a customer purchase order with a liquidated damages clause, and nobody has yet worked out whether the promise date still holds. That question needs the ERP, the email thread, the shipping confirmation, and the spreadsheet where somebody tracks customer commitments.

Manufacturing process analytics gets sold as if both problems have the same answer. They do not. The first genuinely requires a process historian, an MES, and time series storage built for high frequency tags, and no chat interface is going to read a PLC for you. The second is an operational coordination problem sitting across business systems, and it is the one that quietly eats a plant manager's week. This article separates the two, is specific about which tools belong on which side, and is upfront that Skopx only touches the second one.

What manufacturing process analytics actually measures

Strip away the vendor language and manufacturing process analytics is the practice of turning production signals into decisions about the process itself: yield, throughput, quality, energy, changeover time, and the causes behind each.

In discrete manufacturing that means cycle time by station, first pass yield, scrap and rework by defect code, OEE decomposed into availability, performance, and quality, and downtime reasons attributed accurately enough to argue about. In process manufacturing analytics the vocabulary changes but the shape does not: batch genealogy, golden batch comparison, deviation from setpoint, yield loss across a continuous train, spec conformance at each sampling point, and utility consumption per unit produced.

The signals come from four broad places:

Control layer. PLCs, DCS, SCADA, drives, sensors. High frequency, high volume, and the source of truth for what the equipment actually did.

Execution layer. The MES or MOM system that knows what order was running, on what equipment, with which materials, operated by whom, and what quality results were recorded against it. This is the layer that gives a sensor reading its context.

Quality layer. LIMS for lab results, SPC systems for control charting, CMM and vision inspection results, and non conformance records.

Maintenance layer. CMMS work orders, failure codes, asset hierarchy, spare parts consumption.

Every one of those is a real category with real vendors and real reasons to exist. If you are doing analytics in the manufacturing industry at the process level, you need most of them. There is no clever workaround where a generic tool replaces a historian, and any vendor implying otherwise has never had to align a tag timestamp to a batch start.

The two halves of manufacturing process analytics

Here is the split that matters, and it does not appear on most vendor comparison grids.

Half one is machine truth. What did the equipment do, in what state, at what rate, producing what quality, and why did it deviate. This data is high frequency, structurally consistent, generated automatically, and needs specialised storage. A historian compresses years of tags and lets you pull a window at millisecond resolution. An MES holds the order and material context that makes those tags interpretable. A quality system holds the measurement results. The analytics on top are statistical process control, multivariate analysis, batch fingerprinting, and increasingly model based anomaly detection.

Half two is business truth around the line. Which orders are late and which customers care. Whether the supplier delay that landed in someone's inbox has been reflected in the schedule. What the scrap on line two costs in booked margin, not in units. Whether the expedited freight approved by email last month showed up in the cost ledger. Whether the shipment commitment made on a sales call is compatible with the material that has actually been received.

This second half lives in ERP, in email, in Slack, in accounting, in the shared drive, and in at least three spreadsheets that one person maintains. It is nobody's system. It is a coordination surface, and it is where a plant loses money in ways the historian never sees, because the historian is measuring a machine that is running perfectly on a job that should have been rescheduled.

The industry has spent two decades investing heavily in half one and almost nothing in half two, then wondering why plant managers still spend Tuesday morning in tabs.

Layer map: what belongs where

QuestionLayerRight systemWrong tool for it
Cavity pressure drifted, was it the resin lot or the controller?Control and executionHistorian plus MES, SPC on topA chat tool, a BI dashboard, a warehouse
Which defect codes drive scrap on line two this quarter?Execution and qualityMES or quality system reportingManual spreadsheet rollups
OEE by shift, decomposed and trustedExecutionMES with disciplined downtime codingAnything reading raw tags with no order context
Batch genealogy for a recall traceExecution and qualityMES or batch record systemEmail archaeology
Which open orders are exposed to the delayed resin shipment?BusinessERP plus email plus the commitments sheetHistorian, MES
What does last month's scrap cost in booked margin?BusinessERP plus accountingShop floor analytics tooling
Did anyone reply to the customer about the slipped ship date?BusinessEmail, CRM, chatAny manufacturing system at all
Is the expedite spend showing up where finance expects it?BusinessAccounting plus purchasing recordsMES

Read the table by column, not by row. Everything in the top block needs purpose built manufacturing systems and will never be well served by a general tool. Everything in the bottom block needs no historian, no time series store, and often no warehouse either, because the underlying records already exist in systems with perfectly good APIs. What is missing is not storage. What is missing is someone to ask across all of them at once.

What big data in manufacturing actually delivered

Around the middle of the last decade, big data in manufacturing became a board level phrase, and the standard project shape was: stream every tag into a lake, land ERP and MES extracts alongside, build a warehouse, hire a data engineering team, and let the insights emerge.

Some of those programmes worked. The ones that did share three traits. They had a named process problem with a quantified prize before the first pipeline was written. They had a process engineer or a plant metallurgist embedded in the team who could tell a real signal from a sensor artefact. And they kept scope to one line or one asset class until it paid for itself.

The ones that failed usually failed the same way. The lake filled up. Tag names were inconsistent across three plants that had each commissioned equipment separately. Nobody had a maintained mapping from asset to line to product family, so joins produced plausible nonsense. The warehouse cost more each quarter than the last, which is a pattern we cover in detail in Retail Data Platform Cost Control: Where the Money Goes, and the mechanics are identical in manufacturing: storage is cheap, compute on badly modelled data is not, and the bill grows faster than the analytics maturity.

Then a dashboard was built, opened for two weeks, and abandoned, which is its own well studied failure mode. Our piece on BI Visualization Tools Compared: Charts That Get Used is blunt about why: a chart that answers a question nobody asked on a cadence nobody keeps is a decoration.

The correction is not to abandon manufacturing data analytics. It is to be honest about which questions need a modelled, warehoused, joined dataset and which questions just need someone to look in four places quickly. Far more of the second kind exist than most programmes assume.

Shop floor analytics versus manufacturing tracking software

These two phrases get used interchangeably by buyers and mean genuinely different things to vendors, which is how procurement ends up with overlapping subscriptions.

Shop floor analytics usually means the visualisation and analysis layer over execution data: OEE boards, downtime pareto charts, cycle time distributions, andon status, shift comparisons. It is fed by machine connectivity, either directly from PLCs or through an MES. Its value depends almost entirely on the quality of downtime reason coding, which is a human discipline problem more than a software problem. Buy the best analytics in the world and if operators code every stoppage as "other", you have bought a very expensive way to display the word "other".

Manufacturing tracking software is a looser category that spans job tracking, work in progress location, labour tracking, and in small shops something close to a lightweight MES. For a job shop with thirty employees, this is often the first system that puts structure on production at all, and it earns its money by replacing travellers and whiteboards rather than by producing analytics.

A rough selection heuristic:

If your situation isStart withNot with
Manual travellers, no digital job statusManufacturing tracking software or entry MESAnalytics platform
MES in place, downtime coding weakFix coding discipline and operator UXMore dashboards
MES and historian in place, engineers asking causal questionsProcess analytics and SPC on top of the historianGeneric BI
Machine data fine, business coordination chaoticThe business layer, covered belowAnother historian project
Multiple plants, incompatible tag namingContextualisation and a unified namespaceA warehouse that inherits the mess

The fourth row is the one that gets misdiagnosed most often. A plant with good machine data and bad coordination will keep buying machine data tools, because that is the aisle they know how to shop in.

Where Skopx fits, and where it does not

Skopx is an AI workspace that connects nearly 1,000 tools a company already uses, including Gmail, Slack, QuickBooks, HubSpot, Stripe, and Google Analytics, and answers questions in chat with citations back to the underlying records. It also produces a morning brief, runs an insights engine that flags anomalies and risks, and lets you build workflows by describing them in chat.

Start with the honest exclusions, because they matter more than the pitch.

Skopx is not an MES. It does not schedule production, track work in progress, or hold batch records. Skopx is not a historian and does not connect to PLCs, SCADA, or OPC UA, and it will not store or query high frequency tag data. Skopx is not a data warehouse and not an ETL tool, so it will not model your plant data or build you a dimensional schema. It is not a BI tool, so it does not build dashboards. It is not a CRM, an ERP, or a quality system. If your question is why cavity pressure drifted on press seven, Skopx is the wrong tool and no amount of clever prompting changes that.

What it does cover is the second half described above: the operational questions that live in business systems and currently take a person forty minutes across four tabs.

Concretely, the kind of question that fits:

Which open work orders consume the material on the shipment the supplier just delayed, and which of those orders have customer commitments this month. What did we actually pay in expedited freight across the last quarter and which orders drove it. Which purchase orders have been acknowledged but never confirmed with a ship date. Which customer emails about late orders have gone unanswered for more than two days. Which suppliers have slipped promise dates more than once this year, based on what is in the mailbox and the purchasing records.

None of those need a historian. All of them need three or four systems read at once, and cited, so the answer can be checked rather than trusted. This is the same pattern that shows up in other operational domains: reconciling counts across systems in Ecommerce Inventory Tracking Across eBay, Woo, ShipStation, reconciling plan against reality in Retail Analytics Tools for Demand and Inventory Teams, and the general problem of finding an answer that exists somewhere across a company's tools, covered in Enterprise Search Software: Options and the Real Tradeoffs.

The workflow layer handles the recurring version of the same thing.

Supplier delay to at-risk order alert

Supplier email arrives

Message mentioning a revised ship or promise date

Read the revised date

Pull part number, quantity, and new date from the body or attachment

Match consuming orders

Find work orders and purchase orders that use that material

Any commitment at risk?

Compare the new date against committed customer ship dates

Post to planning channel

Named orders and customers, with a link to the source email

Log the exception

Append a row to the supplier performance sheet

A delay email gets checked against open orders and shipment commitments, then posted to the planning channel with links back to the source records.

On cost, Skopx is Solo at $5 per month and Team at $16 per seat per month, with bring your own key for any major model at zero markup, which is set out on the pricing page. That is a rounding error against a historian programme, which is the point: it is not competing with one, it sits beside it.

Sequencing a programme without a warehouse first

Most plants do not need to start with storage. A defensible sequence:

One: name three questions and who asks them. Not themes. Questions, with an owner and a cadence. "Which orders are exposed to late material this week, asked by the scheduler every Monday" is a question. "Improve supply chain visibility" is not.

Two: sort them into machine truth or business truth. Use the table above. Be strict. If the answer needs sensor data at sub minute resolution or batch genealogy, it goes to the MES and historian side, and it needs a proper project.

Three: on the machine side, fix contextualisation before analytics. A unified namespace, a maintained asset to line to product mapping, and disciplined downtime and defect coding will produce more analytical value than any modelling layer bought before them. This is unglamorous and it is the whole game.

Four: on the business side, connect rather than copy. The records already exist in ERP, mail, accounting, and purchasing. Reading them in place and citing them answers most operational questions without a pipeline. Build a warehouse when you have a proven, repeated need for joined historical analysis that live reads cannot serve, not before.

Five: only then decide what deserves a dashboard. A metric earns a dashboard once a named person has asked for it on a fixed cadence for a quarter and acted on the answer. Everything else is a question, and questions are cheaper to answer than to visualise.

This sequencing logic is not specific to plants. The same argument applies to reporting stacks in other functions, and HR Analytics Software vs HRIS Reporting: What You Need makes the parallel case for people data, where teams routinely buy an analytics platform to answer questions the system of record could already have answered.

Common failure modes to avoid

Buying analytics before connectivity. If half the assets are not connected, the analytics layer will confidently report on the connected half and everyone will treat it as the plant.

Treating OEE as a single number. Availability, performance, and quality move for entirely different reasons. A single OEE figure per line is a scoreboard, not a diagnostic, and it invites gaming.

Letting the historian answer business questions. Sensor data will tell you the line ran well. It will not tell you the line ran well on an order that should have been cancelled.

Assuming a warehouse is a prerequisite. Historical, joined, modelled analysis needs one. Reconciling what four systems currently say does not.

Ignoring the coding discipline problem. Downtime reasons, defect codes, and scrap categories are entered by people under time pressure. If the categories are wrong or the entry takes too long, every downstream number inherits the flaw and no analytics vendor can repair it.

One tool for both halves. The single most expensive assumption in this category. Machine truth and business truth have different physics, different data shapes, and different vendors. Buy accordingly.

Frequently asked questions

Can manufacturing process analytics work without a data warehouse?

Partly, and the split is predictable. Deep historical process analysis, batch fingerprinting, multivariate correlation across years of tags, and cross plant benchmarking need modelled, stored, joined data, and a historian plus a warehouse or lakehouse is the right answer. The operational layer around production, meaning orders, supplier commitments, cost, and customer promises, can usually be answered by reading the systems of record directly. Many plants build the warehouse first and discover most of their daily questions never needed it.

Does Skopx connect to a PLC, SCADA, or an MES?

No. Skopx does not connect to control systems and does not read PLC or SCADA data, and it does not replace an MES or a historian. It connects to business systems such as mail, chat, accounting, CRM, and other tools a company already runs, and answers questions with citations from those. Machine level and quality data belongs in purpose built systems, and this is deliberately not that.

What is the difference between shop floor analytics and manufacturing tracking software?

Shop floor analytics is the analysis and visualisation layer over execution data, producing OEE, downtime pareto charts, and cycle time analysis. Manufacturing tracking software is closer to a system of record for job status, work in progress location, and labour, and in smaller shops it is often the first digital replacement for paper travellers. Tracking software creates the data. Analytics interprets it. Buying the second without the first produces empty charts.

Where does process manufacturing analytics differ from discrete?

The systems overlap but the questions differ. Process manufacturing analytics deals with continuous or batch production, so it centres on batch genealogy, golden batch comparison, deviation from setpoint, yield across a train, and spec conformance at sampling points. Discrete manufacturing centres on station cycle times, first pass yield, defect codes, and changeovers. Process plants also carry heavier regulatory recordkeeping, which pushes more weight onto the execution and quality systems and less onto ad hoc analysis.

How should we evaluate vendors selling analytics in the manufacturing industry?

Ask which layer they sit in and refuse to accept "all of them" as an answer. Then ask three specifics: how they handle contextualisation across plants with inconsistent tag naming, what happens to their numbers when downtime coding is poor, and what the total cost looks like in year three including storage and compute growth. A vendor who cannot describe their limits precisely has not thought about them, and the same evaluation discipline applies in any regulated or data heavy sector, as we argue in Healthcare Analytics Companies: How to Compare Vendors.

What is the smallest useful first step?

Write down the last twenty questions your plant leadership actually asked, and sort them into machine truth and business truth. Most plants find the split is closer to even than they expected, and that the business truth pile has no owner and no tooling. That pile is usually cheaper and faster to fix, and fixing it buys the political capital to do the machine side properly.

Share this article

Skopx Team

The Skopx engineering and product team

Related Articles

Stay Updated

Get the latest insights on AI-powered code intelligence delivered to your inbox.