A factory automation rollout can keep producing healthy dashboards while two services both believe they own the same machine. For a CTO choosing between buying and building, that is the scaling failure I would design around first: conflicting authority over recipes and commands. Telemetry gaps are visible after the fact; an unauthorized or stale change can reach a cell before anyone knows there is a dispute.
Command ownership breaks before the dashboards do
When factory automation software scales, telemetry fails first puts data loss at the front of the queue; I would put conflicting command authority there instead, because two valid-looking writers can act on the same cell while every monitoring system remains green. The distinction matters to a CTO: extra telemetry can reveal a bad rollout, but it cannot decide which service was entitled to make the change.
Consider an illustrative estate of 6 plants with 20 cells each; those figures describe a planning scenario, not a measured deployment. A central team ships a new recipe service, while one plant retains a local interface used during network outages. At pilot scale, the two routes rarely overlap. At estate scale, routine events—a shift change, delayed deployment, reconnecting edge gateway, or vendor maintenance session—make overlap likely enough to test explicitly. Both routes may authenticate correctly and both may write valid values, yet neither knows that the other has taken ownership.
The first design decision is therefore not which message broker to install. It is where authority resides for each cell, who may transfer it, and what happens when that authority cannot be checked. Use ISA-95, published as IEC 62264, to define plant, line, and equipment boundaries consistently; then assign a single command owner at the boundary where a write can affect production. The authority record should name the cell, current owner, revision, permitted operation, and expiry policy. A service that merely knows a machine’s network address has not earned permission to change it.
OPC UA, standardized as IEC 62541, can carry a method call or write to an equipment-facing interface, but a valid OPC UA session alone does not establish business authority because authentication and production approval answer different questions. MQTT 5.0 can deliver a command at QoS 2, but that delivery guarantee does not make the physical effect occur exactly once because a controller can act before an acknowledgement is lost. Sparkplug B 3.0.0 birth and death certificates help establish an edge node’s session state; they do not choose which of two applications owns a recipe. Treat these protocols as transport and state signals, not as substitutes for an authority decision.
I would not make a cloud service the sole real-time arbiter of machine commands, because a network partition would then force a choice between stopping approved local work and bypassing the arbiter. Keep the approved operating envelope enforceable at the controller or plant edge, and make a disconnected edge reject new ownership transfers until it can reconcile. That rule is less convenient than “queue everything and replay it,” but replaying a command after its production context has expired is precisely the failure the rule prevents.
A recipe without an immutable identity cannot scale safely
Teams often version application code carefully while treating a recipe name as if it identified recipe contents. It does not: “standard run” can refer to different parameters at two plants, or to different parameters at the same plant before and after a local edit. A scalable apply request needs an immutable recipe revision, target cell, approved effective window, and authority revision, because the recipient otherwise cannot distinguish an intentional update from a delayed duplicate.
One workable envelope is plant/line/cell + recipe ID + content digest + authority revision + request ID. Compute the digest over a canonical representation—RFC 8785 defines JSON Canonicalization Scheme if JSON is the interchange format—and use SHA-256 so two systems can compare the same bytes rather than arguing over display names. Store the approved mapping from digest to equipment-specific parameters separately, because PLC tags and limits differ even where the business recipe is shared. A digest proves equality of represented content; it does not prove that the values are safe for a particular controller.
That distinction also changes the rollout sequence. First validate the recipe against cell-specific constraints. Next obtain plant approval for the exact digest and target. Finally issue an apply request that includes the authority revision expected by the sender. If the revision has moved, reject the request rather than silently adopting the newest owner’s permissions. The rejection gives operations a specific conflict to resolve instead of an unexplained difference between intended and actual setpoints.
Advanced Manufacturing Automation Software for IT Teams is a useful starting point for assigning IT responsibilities, but an IT ownership chart cannot settle a running cell’s write permissions without a plant-approved authority map. That map should also identify who can revoke approval after a quality hold and who can authorize a controlled fallback. Otherwise, an apparently routine retry can reapply a recipe that the plant has deliberately withdrawn.
Make the failure visible with Prometheus counters such as recipe_apply_rejected_total{reason=”stale_revision”} and authority_conflict_total, and attach the recipe digest and request ID to an OpenTelemetry trace. Avoid using the recipe payload itself as a metric label, because unbounded label values make time-series storage harder to operate. For an initial acceptance exercise, tune a target of zero accepted stale-revision writes across forced reconnects; it is a test criterion, not a claim that a live estate has achieved it.
A deployment tool cannot grant permission to touch a cell
Kubernetes and Argo CD are useful for deploying the services that prepare and route commands, but a successful application sync is not plant approval because deployment state does not encode the current production window. If an estate uses Kubernetes 1.30 and Argo CD 2.11, keep their rollout status separate from the authority record: the former says which software should be running; the latter says which instance may request a change to a particular cell.
Enforce ownership with a compare-and-swap operation on the authority revision. The following Python 3 example runs as written and shows the essential property: two contenders present the same expected revision, and only the first claim succeeds. It is a demonstration of admission logic, not a production controller or a substitute for plant approval.
python3 - <<'PY'
import sqlite3
db = sqlite3.connect(":memory:")
db.execute("CREATE TABLE authority (cell TEXT PRIMARY KEY, owner TEXT, revision INTEGER)")
db.execute("INSERT INTO authority VALUES ('plant-a/line-2/cell-7', 'edge-a', 4)")
def claim(cell, owner, expected):
result = db.execute("UPDATE authority SET owner=?, revision=revision+1 WHERE cell=? AND revision=?", (owner, cell, expected))
return result.rowcount == 1
print(claim("plant-a/line-2/cell-7", "edge-b", 4))
print(claim("plant-a/line-2/cell-7", "edge-c", 4))
print(db.execute("SELECT * FROM authority").fetchone())
PY
The output is True, then False, followed by a row owned by edge-b at revision 5; the second contender fails because the first update changed the revision atomically. For a shared service, PostgreSQL 16 can hold this record transactionally, but the database decision still needs enforcement at the last trusted gateway before a controller. Otherwise, a disconnected writer could bypass the service and send a perfectly formatted command through an older route.
Design the gateway so it verifies the approved digest, current authority revision, target cell, and command lifetime before forwarding a write. During a partition, it should continue only the locally approved operations defined for that cell, because granting fresh remote authority without a shared view recreates the conflict. IEC 62443-3-3 provides requirements for industrial system security, but meeting a security requirement does not, by itself, specify the plant’s production approval policy.
Test the boundary with faults rather than a happy-path demo. Pause the WAN, start an ownership transfer, restart an edge service, deliver a delayed MQTT message, and revoke a recipe while its request is queued. As a proposed drill threshold, allow no more than 60 seconds to surface an authority conflict to the operator; adjust that limit to the plant’s response process rather than treating it as an industry benchmark. Record accepted and rejected writes separately, because a healthy service uptime figure can conceal the exact unsafe acceptance the drill is meant to find.
Buy the workflow if it fits; build the authority boundary if it does not
The explicit choice is buy Siemens Opcenter Execution with the required plant integrations or build a narrow in-house authority service backed by PostgreSQL 16. Buying wins when the plant already uses the vendor’s execution model and a proof of concept shows that recipe revisions, approvals, revocation, and offline behavior map cleanly to its equipment; its costs are licenses, integration work, and dependence on the vendor’s supported change path. Building wins when a mixed legacy fleet cannot express those rules in the purchased model without bypasses; its costs are ongoing security review, gateway maintenance, and an on-call team responsible for disputed commands.
Do not decide from a feature checklist, because both options can demonstrate a successful recipe change while failing on competing writers. Give each contender the same acceptance test: two command sources claim one cell, the WAN drops during transfer, approval is revoked, and a delayed command arrives after reconnection. Require an audit record that identifies the rejecting component and its reason. If the bought product needs an undocumented side channel to pass, that side channel becomes the in-house authority service in all but name.
For the commercial route, ask the vendor to show exactly where the authority revision is checked and what remains enforceable when its central service is unavailable. For the internal route, require a plant-edge enforcement point and a documented recovery procedure before funding a broader rollout. Neither route should promise that a database transaction makes a physical operation reversible; the command may have crossed the gateway before a later failure is recorded.
The first decision should be made at one contested cell
Pick one cell with two plausible command paths and document who owns each permitted write today. Then stage a conflicting recipe change and a network interruption against that cell before signing a platform contract or staffing a build program. The result will expose whether a product can enforce your plant’s authority rules—or whether the narrow service you need to build is now clearly defined.


