Drawing on over a decade of experience in guiding OEMs on over-the-air (OTA) software update infrastructure across the IoT device lifecycle, the same common challenges emerge.
Across smart device types, industries, and use cases, what may initially seem like a simple infrastructure process — the software updates system — often requires more resources than anticipated. Updating software is critical and should work seamlessly. But typically, it is not a differentiating feature in and of itself.
For example, proprietary, DIY, or homegrown OTA update mechanisms sound appealing in theory. But in reality, in-house systems often introduce countless technical and business drawbacks and require ongoing management, consistent maintenance, and purpose-built considerations. Just as building an enterprise email server instead of deploying Google Workspace is inefficient, creating an in-house software update system drains resources. Technical issues, insufficient execution, and overall maintenance and support demands can turn the promising concept of OTA software updates into a significant business operations challenge.
In turn, these challenges can stunt product development efforts and the success of deployed products in the field. Any proprietary infrastructure requires resources and thereby diverts engineers from core product innovation. In extreme cases, a poor software update system can lead to fleet-wide outages and lasting reputational damage.
Software update deployment inefficiencies and operational rigidities are commonly where challenges first emerge. Here are four scenarios to avoid with real-world examples.
Operating in the commercial and residential HVAC space, one company managed over 25,000 HVAC smart products with its proprietary OTA software update manager. Its in-house software update infrastructure relied on its products, fleet-wide, pulling data from a central server every six hours. As such, the entire fleet received each software update. The infrastructure lacked the ability to segment software updates by the geographic location of these devices or to determine whether a given unit even needed the update. Every six hours, any product capable of connecting (e.g., power, network) would connect to the central server and pull the currently available software update.
The rapid growth of the company created the need to continuously iterate on product features and fix various bugs. With no other choice, the company was forced to flag updates globally for all 25,000 devices simultaneously. As their homegrown update server lacked segmentation, whitelisting, or geographic targeting, there was no other way to notify users of an available update.
Another company, a smart building OEM, relied on a remote development team to manually push updates to its products. Each time a new product was sold and added to the fleet, the device required the latest update. The remote development team manually sourced the update, and an engineer then deployed it to the device. To further complicate the software update process, the entire system lacked automatic rollback capabilities. For any error or failed product update, an engineer must first manually rediscover and re-trigger failed deployments, then resolve the issues for each individual device.
For a mining company operating in remote locations, manual intervention carried even higher stakes. When an update failed on equipment deployed in remote mountains or on oil rigs, there was no remote solution – the company had to physically deploy engineers by helicopter to recover the affected devices.
In both scenarios, to support the manual model, each company had to maintain a growing roster of engineers dedicated solely to manual device management, from updating software to troubleshooting failed products. As the fleets grew, the headcount supporting more devices increased rather than shrank as the product matured.
For a smart-camera company, their custom OTA update system assumed and relied on consistent, stable connectivity. If a device were turned off during deployment, the software update would fail permanently and never restart automatically. Without dual-partition (A/B) rootfs layouts, automated rollback engines, or automatic retries, the failed updates essentially bricked the products, rendering them inoperable. Their infrastructure setup also lacked visibility. The company did not know the device status or update completion rates, meaning it also did not know how many devices failed to receive the update and the overall state of its product fleet. As a result, the fleet drifted.
This operational uncertainty created a cascading effect. The company did not know each product’s status to proactively update devices, but customers did, and, in turn, customers triggered a high volume of support tickets requesting the latest firmware. The customer success and support teams then had to attempt to assist and appease customers while lacking visibility into the device being updated.
A drone company faced major development delays because their basic software implementation required downloading full 3GB compressed root filesystems — up to 9GB total per aircraft — for even minor code changes.
In the case of a mining company, when its product was located in remote locations, it needed to push updates over encrypted tunnels on expensive satellite networks without binary-level optimization. As a result, each update consumed significant portions of the data budget. As data costs climbed with each deployment, the company was forced to delay non-critical updates and batch fixes together — trading update frequency for network spend it could not sustain.
Proprietary update systems rarely optimize software update payloads. It’s very rare for the development of an in-house OTA system to include any kind of update breakdown, such as delta updates. As a result, the OEM incurs significant network expenses or adjusts the update frequency to manage associated costs.
These are four scenarios OEMs encounter with in-house OTA update systems. Relying on a custom-built OTA infrastructure frequently introduces severe deployment inefficiencies and operational rigidities. For example, the complete lack of fleet segmentation, whitelisting, or geographic targeting creates significant operational uncertainty, leaving teams without visibility into device statuses or update completion rates. Without the ability to optimize data payloads, operations experience significant network strain and exorbitant bandwidth costs. Most importantly, failing to implement robust fault-tolerant features, such as dual-partition (A/B) rootfs layouts and automated rollback, leaves products highly vulnerable to failure, potentially impacting operations and the overall business.
In planning the software update infrastructure, accurately assessing the system in context of the organization is just as important as the technical architecture itself. Many of the OEMs above didn’t lack engineering talent; they lacked foresight and future-proofing. Managing a fleet of 25 devices requires a fundamentally different level of operational rigor than a fleet of 25,000. Waiting until the fleet has already outgrown its update infrastructure means retrofitting a solution under pressure, rather than building for scale from the beginning.