Project Zero often works with software vendors to remediate the vulnerabilities we report and provide broader guidance on making software more secure. Some vendors express concern about potential scenarios in which they are unable to fix vulnerabilities that are causing immediate user harm, due to limitations in their patch delivery systems. Since Project Zero encounters a wide array of systems designed to protect users in the case of exceptional exploitation scenarios, both through vendor discussions and security reviews, we want to share what we’ve learned.

This post provides an overview of systems in use by large vendors that allow them to remediate small volumes of vulnerabilities much faster than their typical update process. Our goal is to provide a reference for vendors seeking to implement or enhance the capabilities of such systems, and to encourage vendors to consider how they would fix an urgent vulnerability before they receive one.

Why patching takes time

Patching a vulnerability typically involves the following stages:

  • Triage — a vulnerability report is received, validated, prioritized and assigned to a specific developer to be fixed
  • Patch development — a software development team writes, reviews and commits code that fixes the vulnerability
  • Testing — the patch is tested to ensure the vulnerability is remediated and the software still functions correctly when the patch is applied. This can include formal testing by a test team, automated testing and alpha and beta testing where a patch is shipped to a limited group of users for feedback on normal use.
  • Partner review — some software updates require review by third parties before they can be shipped, due to relationships between the software vendor and other organizations, for example, carrier acceptance for some mobile updates.
  • Delivery — the patch is delivered to and installed by end users
  • Activation — sometimes an additional step, such as a system restart, is needed to switch the system to the updated software

Of course, this is a simplified picture. Patching can involve repeating steps, for example rewriting a patch if tests fail, or additional stages when third-party vendors are involved. However, this is a minimal set of steps most software updates require.

The challenges of emergency patches

While triage and patch development time contribute substantially to the speed at which vendors can generally patch vulnerabilities, they contribute less to emergency patch time. Triage is usually very fast in situations where vendors know they have an urgent problem, and patch development can be expedited based on priority. Only in rare circumstances, where a vulnerability is especially complex, or a vendor’s security team does not have a complete picture of their software’s components and who within their organization maintains them, have we seen urgent patches delayed in the triage or development phase. Likewise, partner agreements usually have exceptions for updates in emergency situations.

Most vendors’ patch speed is limited by the testing and delivery stages. Testing is important because all changes to software risk introducing unexpected behavior. The worst-case scenario is that inadequately tested software ‘bricks’ a device, causing it to malfunction in a way that it can no longer perform key functionality or receive software updates to remediate this. Buggy software updates have also led to situations where user data is corrupted or lost, and any decrease in software functionality after a security update makes users less likely to apply updates in the future.

The potential cost to vendors of shipping poorly tested updates varies depending on the nature of the underlying software. For example, if a mobile application is rendered unusable due to an update that corrupts local data or prevents it from launching, users can easily install the next version via an app store, and their data is usually saved on a remote server, so costs are limited to user support. Meanwhile, if a mobile device gets bricked, it needs to be returned to its manufacturer or place of purchase for repair, leading to substantial costs for the vendor and potentially the user.

The possibility of serious functional bugs is considered in the design of most patch delivery systems. Updates are often rolled out slowly, so that serious problems can be detected before they affect too many users. Often, patching vulnerabilities quickly and avoiding buggy patches are at odds with each other, requiring tradeoffs that prioritize one over the other.

A variety of other technical challenges can limit the speed of patch delivery. One is the design of the patching system. A common design is that devices probe for updates at a regular interval, leading to patch saturation being limited to that interval. ‘Push’ style update systems can deliver patches to all users faster, but generally require more infrastructure.

User behavior and environment can also be a barrier to patch propagation. Patches that require user interaction to install are often delayed by users, and network speed and data cost are also factors in installation rate. Updating many users at once, as opposed to over a period of time, can strain patch delivery infrastructure. Chrome and Microsoft have written about the challenges of updates requiring restart to install, as users are often reluctant to restart their system and restarts take time.

While testing delays and limitations of the patch delivery system affect all updates, the shorter time frame of emergency updates make them a larger contributor to the overall time it takes to deliver a patch.

Emergency patching methods

Feature flags

Feature flags are conditional statements in source with paths determined by values provided by a remote server. They are often used for A/B testing, but they can also be used for short term remediation of vulnerabilities in emergency situations. A widely publicized case of this was a serious 2019 FaceTime vulnerability, where Apple temporarily disabled Group Facetime with a feature flag. Several vendors have made at least some media codecs available in 0-click contexts controllable via feature flags, and can disable them in the case of active exploitation, falling back to another codec for realtime transmission.

The main benefit of feature flags as a vulnerability remediation method is that testing can be performed with each flag set in advance, so a fast update does not require shipping untested code. They can also be delivered to users much more quickly, as updating feature flags requires transmitting a very small amount of data.

Recently, Meta published a blog post on how they implemented a ‘dual stack’ library, in which two versions of the WebRTC video conferencing library were compiled into a single binary, with the version in use controllable via a feature flag. This technology enables rapid updates with less testing, as new versions can be shipped with the option to quickly move users back to the previous version if function problems occur. While Meta uses two versions of the same library, it would also be possible to create a ‘dual stack’ with two different libraries that implement the same features (for example, two H264 libraries), allowing an application to switch to a different library to render a specific vulnerability unreachable without loss of functionality in an emergency. This would require additional testing, but it is testing that can be performed up front. It could also be possible to have a second library that enables performance intensive mitigations that would block many possible bugs, such as ASAN, or enabling DCHECKs.

Filtering

Filtering is running a dynamically updatable ruleset, such as a regular expression, against untrusted input in order to block specific input that is required to reach a vulnerability. An example of this is Android’s Intent Firewall, which allows specific usages of an Android IPC mechanism called intents to be disabled based on rules in a dynamically updateable XML file, which enables blocking intents that can be used to exercise specific vulnerabilities. It was recently used to block vulnerabilities in third-party Android wallets.

Some platforms have endpoint detection software that can perform filtering on a wide variety of system input, for example Microsoft Defender on Windows systems, and Google Play Protect on Android devices. Rules that block specific exploits or make certain vulnerabilities unreachable can often be deployed to these applications very quickly. Endpoint detection requires parsing a great deal of untrusted input, often in privileged context, so these applications are not without risk, but in systems where they already exist, they are a potential method of emergency remediation.

As an approach, filtering is more flexible than feature flags. For feature flags to be effective, the vendor needs to determine what features they might want to disable in advance, and if this isn’t comprehensive, they might find themselves in a situation where a vulnerability can’t be remediated via feature flags. Meanwhile, filtering can be used to block a wide variety of inputs, even ones that have never been considered. The downside of filtering is that performing filtering frequently can decrease software performance, and at least some testing of new filters is required, and can’t be performed upfront without knowing the vulnerability that needs to be blocked, as it is possible to write filters that interfere with necessary system functions.

Alternate Channels

The network ‘channels’ used to deliver software updates to users can be slow for a variety of reasons discussed above. Vendors sometimes implement alternate channels that can be used to deliver smaller updates more quickly.

Android Pony Express (APEX) is an example of an alternate channel that can be used to ship updates to specific high-risk Android components faster than a full system update. It shortens the patch development time, as OEMs do not need to integrate updates to APEX components. APEX is available to OEMs, and can be used to update OEM-maintained libraries.

Several applications we’ve researched have the ability to update individual libraries outside regular updates, usually by having some flag that is regularly checked over the network, and then downloading the library and loading it with dlopen or equivalent. While this is an effective way to avoid delivery-speed limitations of updates, it can also introduce critical vulnerabilities if libraries delivered in this way are not adequately verified by the client to have originated from the vendor. We encourage vendors to be cautious, and ensure that emergency update mechanisms of this variety have adequate security testing.

Hotpatching

Some vendors have implemented update mechanisms that allow units of binary code smaller than libraries to be delivered and applied directly to the memory space of a running process. For example Linux supports Livepatch which enables kernel functions to be directly replaced in memory without a restart. Similarly, Windows’ hotpatch allows security updates that contain only updated functions to be delivered to users, and applied while the process is still running.

Hotpatching has the potential to deliver very flexible security patches to software very quickly, with no degradation of user experience, though it typically has some limits to the nature of patches it can deliver, for example, updates that require changing the definition of a structure shared between functions are sometimes not supported. Hotpatching has similar security downsides to alternate channels, and also carries the risk of introducing ways to bypass exploit mitigations, as it requires permissions to map pages with write-execute privileges at some point during patching. It also doesn’t address any of the testing challenges of rapid updates, just the delivery challenges.

The importance of emergency patching

LLMs are increasing the vulnerability discovery and exploitation capabilities of both attackers and defenders. A wider array of actors now have the ability to perform novel attacks at greater speed. In light of this, it is important for vendors to consider how to protect their users in the case of active exploitation. Rapid update mechanisms do not need to be heavyweight or be capable of fixing every possible bug and preserving perfect user experience in every scenario. Technologies like feature flags, filtering and alternate update mechanisms can remediate the most likely and severe vulnerabilities in the short term, while keeping devices reasonably functional for users.

It is urgent for vendors to plan how they will protect their users in the worst case scenario of widespread active exploitation. Actions taken now can greatly improve security outcomes for users. By taking stock of update mechanisms already available to them and implementing rapid remediation functionality where gaps exist, vendors can be better prepared for whatever the future holds.