A factory software rollout can pass its pilot and still become expensive at the second site: nobody can say which team owns a changed machine tag, an incorrect asset mapping, or a deployment during a production shift. For a CTO choosing between building and buying, my position is that change authority breaks before the platform does. Buy the repeatable plumbing when it fits, but keep the asset contract and approval rules under your control.
Asset identity breaks before data volume becomes the deciding constraint
A pilot usually has a person who knows that “Press7” in the historian, “P-07” in the maintenance system, and a particular OPC UA NodeId all mean the same machine. At another site, those names may describe different equipment. The pilot’s mapping works because that person can correct it informally; the rollout fails when software starts treating local knowledge as a global identifier.
When factory automation software scales, telemetry fails first identifies a useful symptom, but I would investigate naming authority before buying more monitoring: a complete stream of readings is still misleading if it is attached to the wrong asset. Missing data is visible; plausible data assigned to the wrong press can survive until someone makes an operational decision with it.
Give each physical asset a stable enterprise identifier and keep site-local identifiers as mapped attributes. A mapping record should name its site, source system, OPC UA namespace URI and NodeId, enterprise asset ID, approving owner, and effective date. The namespace URI matters because a numeric namespace index can change between OPC UA server configurations. Keep the previous mapping version as well: without it, a corrected tag can silently rewrite the apparent history of an asset.
ISA-95 offers a vocabulary for separating equipment, operations, and enterprise systems, but it will not decide whether the controls engineer or the central product team may approve a particular mapping. That decision is organizational and should be recorded beside the schema. For an illustrative sizing exercise, 120 assets with 40 signals each yield 4,800 mappings at one site; those are planning assumptions to replace with an inventory, not a measured fleet. Even if only a small fraction change during commissioning, an approval process conducted through chat messages will be difficult to audit because the rationale and effective time are not tied to the deployed mapping.
I would not make a vendor’s internal asset hierarchy the sole source of truth, because replacing that vendor would then require reconstructing which physical equipment every historical series represented. Store an exportable identity register and test that the purchased system can ingest and return those IDs unchanged.
Buying wins at the edge only when mappings and exits remain yours
The explicit choice is AWS IoT SiteWise with a SiteWise Edge gateway versus an in-house OPC UA collector feeding Apache Kafka. SiteWise wins when its supported ingestion path matches the site’s controllers and the CTO needs a managed route from collection to stored asset data. Its cost is service charges, configuration work, and dependence on AWS asset-model and export behavior. The collector-and-Kafka option wins when unusual source systems or strict control over data routing make a standard gateway a poor fit. Its cost is engineers on call for connection failures, buffering, upgrades, and security patches—not merely the initial collector code.
Neither option resolves ownership by itself. Before signing a contract or funding a build, ask both teams to demonstrate the same change: move one signal from an incorrectly identified asset to the correct one, preserve its earlier provenance, and show who approved the correction. A platform that can ingest a signal but cannot represent that change without overwriting history is an expensive place to discover a weak data contract.
Test failure behavior with specific protocols rather than the promise of “offline support.” MQTT 5.0 provides a Message Expiry Interval; a value chosen for transient temperature readings may be wrong for an event that must survive a prolonged outage. Sparkplug B 3.0 defines birth and death messages that can help a subscriber interpret device state, but they do not prove that a device’s asset mapping is correct. If Kafka is part of the design, a setting such as min.insync.replicas=2 is useful only with enough replicas and matching producer acknowledgments; the setting alone cannot protect a single-broker installation.
For a proposed pilot, use 10 seconds as an initial freshness target for a non-safety dashboard, then tune it against the actual decision the dashboard supports. Measure separately the time at the source, the time received at the edge, and the time available to the user. That separation identifies where delay occurs; one end-to-end latency number cannot distinguish a stalled controller connection from a slow cloud pipeline. The number is a trial setting, not a universal factory requirement.
Put the exit terms in the technical acceptance criteria: a documented export of raw readings, stable asset IDs, mapping versions, and timestamps with their time zones; a way to run the collector during a network interruption; and a tested restoration procedure. A cheaper ingestion quote is not cheaper if departure requires manually relabeling years of history.
Release approval fails when the repository owns decisions the site must make
At 30 engineers, CI and code ownership fail before runtime points to a real coordination problem, but a green build cannot authorize a change on a running line: the software team and the site may have different windows, risks, and rollback powers. More reviewers in GitHub CODEOWNERS can protect source changes; they cannot substitute for an identified site approver.
Separate three decisions. The product team approves the mapping format and application behavior. The site’s named owner approves what a local controller tag means and when a change can be activated. The deployment service enforces that the approved artifact and mapping version are the ones released. This division avoids asking an operator to review application code while also preventing developers from silently redefining equipment.
Start enforcement with a small, executable check. This Python 3 command rejects an empty identity field or a duplicate mapping key in a proposed CSV; it is a gate for obvious errors, not proof that a tag describes the right machine:
python3 - <<'PY'
import csv
import io
data = "site,line,asset,source\nA,1,Press7,ns=2;s=Temp\nA,1,Press7,ns=2;s=Temp\n"
seen = set()
for row in csv.DictReader(io.StringIO(data)):
key = tuple(row[name] for name in ("site", "line", "asset", "source"))
if not all(key) or key in seen:
raise SystemExit(f"invalid or duplicate mapping: {key}")
seen.add(key)
print("/".join(key))
PY
This example deliberately exits with an error. In production, validate against the approved asset register as well, because uniqueness within one file says nothing about whether an enterprise ID is legitimate. Run the check in GitHub Actions before a mapping can be promoted, and retain the approver and version in the release record. Keep operational authorization separate from source-code review: that makes it possible to establish both who changed the software and who permitted it to affect a site.
IEC 62443-3-3 is a useful reference for system security requirements, but invoking the standard is not a release plan. Write down which identity may deploy to the edge, how credentials are revoked, and who can invoke rollback. A rollback must restore the compatible application and mapping version; reverting code alone can leave a correctly running collector assigning data to the wrong assets.
A deliberately awkward pilot gives the CTO a defensible purchase decision
Do not score a platform solely on a demonstration at the easiest site. Choose two pilot sites as a procurement test: one with tidy source naming and one with known naming conflicts or intermittent connectivity. That site count is a test-design choice, not a claim that two sites represent the whole estate. Give the vendor and the internal team the same tasks and the same access constraints so the comparison measures delivery rather than familiarity.
Introduce a controlled mapping error, rename a source tag, interrupt the edge-to-cloud connection, and request a rollback during an agreed maintenance window. Measure how long it takes to detect the wrong identity, obtain an approval, deploy the correction, and establish whether earlier readings were affected. Record the engineer-hours spent preparing source access and resolving exceptions; those hours belong in total cost even when the software license calls the connector “included.”
Use a provisional 24-hour limit for producing an auditable account of the injected error, then adjust that limit to the consequence of an incorrect decision at your sites. It is an acceptance threshold to negotiate, not a measured industry benchmark. Also inspect a sample of corrected records directly: a fast incident report is weak evidence if it cannot identify which asset ID, mapping version, and source timestamp produced a disputed reading.
Prometheus can report ingestion failures and queue depth; OpenTelemetry can trace the path through services the team controls. Neither can observe the meaning of a tag inside a controller without a maintained mapping. Put a metric on unapproved mapping changes and review its underlying records, because a zero error count may simply mean that nobody instrumented the approval boundary.
Choose SiteWise if it completes these tests with an acceptable export and less ongoing operational work. Choose the in-house OPC UA and Kafka route if the vendor cannot preserve the identity and correction history your sites require, and budget explicitly for ownership of the edge service. If neither passes, narrow the rollout rather than treating the pilot’s successful dashboard as proof that the system can scale.
The first purchase decision is a mapping test
Ask one site owner and one software owner to select a disputed tag this week. Have them agree on its enterprise asset ID, source identifier, approval record, and correction procedure, then require both build and buy candidates to process that record and its later change. The result will expose a procurement constraint before a broad deployment turns a local naming exception into expensive historical data.



