Service providers are moving towards mesh-based transport networks, and the latest technological advancement is standards-based Shared Mesh Protection, which leverages an intelligent GMPLS control plane so that a mesh transport network can recover from multiple local or network-wide failures while reducing costs by eliminating the need to dedicate backup bandwidth for each active circuit. In this article, we will examine hardware-accelerated Shared Mesh Protection as a means of increasing network resilience without incurring additional fiber-related expenses.
In India's rapidly developing economy, the constant cycle of construction, demolition, and roadwork puts existing telecommunications infrastructure at risk. According to industry sources, the average Tier-1 operator in India experienced 12 to 15 fiber cuts per 1,000 kilometers per month. To put this into context, consider that the top Tier-1 operators in India own anywhere from 80,000 to 190,000 kilometers of fiber. This translates to more than 60 fiber cuts per day, 2,000 per month, or 25,000 fiber cuts per year, which these operators must address promptly. In a globalized economy, customer expectations are constantly being benchmarked against the best in the world. This presents significant challenges for Tier-1 operators in terms of increased operating expenses for fiber repair, increased capital expenditures for the optical protection traditionally available to cope with fiber cuts, and increased expenses for customer retention.
As bandwidth continues to grow at astonishing rates, estimated at 40% per year globally, driven by enterprise applications such as cloud, mobile, and video technologies, a single outage of just 50 minutes in a year can reduce network availability by as much as four nines, or 99.99%. Whatever the cause of network disruptions, the service provider must address the issues. In this overall effort, there are two approaches:
• Protection. This must occur within 50 milliseconds of the fault (the accepted gold standard for recovery). To achieve this rapid response, a pre-calculated path is typically used for the protection circuits. The protection capacity can be dedicated or swapped. However, there are very strict limitations for traditional switched protection protocols in terms of topology and scalability. There may also be limitations in multi-fault protection.
• Restoration. In a restoration operation, when a fault is detected, a new communication path is calculated, and connections are dropped on the working path and re-established on the safe path. Restoration is common in packet networks, although it can be comparatively slow (typically ranging from seconds to minutes). It allows for sharing protection capacity and almost always finds a path, provided a backup path exists. Restoration is also possible at the digital transport layer, with performance improvements ranging from hundreds of milliseconds to a few seconds, but still well below the sub-50-millisecond fault recovery requirement.
Today, the approaches described above are implemented using the following resilience techniques:
• SONET/SDH: 1+1 and Subnetwork Connection Protection (SNCP), which uses dedicated bandwidth protection to provide guaranteed sub-50 millisecond protection for all payload types (SONET/SDH, Ethernet, SAN, video). This protection is used in the transport networks of many companies today. This type of protection defines the 50-millisecond value but does not offer protection against multiple failures, and the way it is typically implemented means that the protection capacity cannot be shared by other services, thus increasing costs.
• OTN Digital/GMPLS: Mesh software restoration is provided by newer intelligent optical cross-connect switches as an alternative to using dedicated bandwidth for protection. When a failure occurs, these devices use intelligent control planes to reroute affected services using software-based tables, which utilize the allocated bandwidth in the network. Since all unallocated bandwidth is available as a shared pool of restoration bandwidth, this mechanism is typically 20–35% more efficient in terms of network resources compared to dedicated protection bandwidth. Furthermore, because GMPLS mesh restoration dynamically redirects a failed service based on available bandwidth, this procedure can be repeated in the event of multiple network failures. However, the multi-stage, software-only approach can take seconds to recover, and overall restoration time will increase with the complexity of the network topology, the number of links, and the number of restoreable connections.
IP Packet/MPLS: Fast Redirection (FRR) is a router-based protection used in data-centric networks. Like GMPLS mesh restoration, MPLS FRR utilizes shared protection bandwidth for network efficiency and can recover from multiple failures. One of the goals for MPLS has always been to provide greater resilience against connectionless IP networks. MPLS FRR (sometimes referred to as local MPLS protection) allows a Label Switch Router (LSR) to react within 50 milliseconds with local redirection once it detects a failure in the working path. MPLS FRR uses pre-calculated Label Switched paths and label values, so all it has to do is assign a new label and direct traffic to a different port. MPLS FRR allows for arbitrary topology and is a shared protection technique. The drawbacks of MPLS FRR are that the sub-50-millisecond operation is not fully deterministic because it is only local, and once a failure occurs, the entire network may need to re-converge, requiring the use of additional MPLS/IP routing ports to achieve resilience.
The Case for Hardware-Accelerated Shared Mesh Protection:
As we have seen, networks now face multiple failures, and single-failure protection is no longer sufficient. Furthermore, multi-failure solutions are simply too expensive, considering that the increased traffic across fibers can carry up to 8 Tb/s of capacity. The pricing pressures faced by service provider business models demand a new approach to resilience.
The ideal resilience technology for modern transport networks should offer three fundamental functions:
1. Multi-failure recovery for improved survivability; 2. Rapid recovery within 50 milliseconds for deterministic performance; and 3. Intelligent sharing of backup resources for better economics.
These three capabilities are now available in a single hardware-accelerated shared mesh protection technology. This solution offers service providers the opportunity to create tiered protection plans, which could help generate additional revenue for minimal investment in protection capacity.
Today, the ITU-T is working on two documents: G.SMP (G.808.3) and G.ODUSMP. The former aims to standardize the independent parts of SMP technology, while the latter aims to standardize the digital OTN layer. These protocols cover message encoding, signaling, activation, and other functions necessary for SMP. Meanwhile, the IETF is working on two projects to standardize its application across digital circuits and packet networks.
The SMP protocol is fundamentally a proactive approach to network protection. It separates time-critical tasks, such as protection activation, from the longer-term calculation of the more resource-intensive GMPLS control plane path. The SMP protection activation protocol is designed to be lightweight, a key feature that allows it to be implemented in hardware, support thousands of services, and provide rapid recovery.
How Hardware Acceleration Works and Its Importance:
While long-distance transmission is rapidly moving toward coherent 100 Gb/s and 500 Gb/s superchannel technologies, service demands are still predominantly driven by very large numbers of Gigabit Ethernet and 10 GbE connections. With the move to 100G and 8 Tb/s capacity per fiber, a single fiber cut could impact many thousands of services. Customers have been planning to use backbone transport platforms for over a decade, requiring them to be built to handle multi-terabit scales at a highly granular service level, while providing unparalleled bandwidth efficiency and resilience. This new design allows SMP to be implemented using dedicated hardware acceleration processors that support 50-millisecond recovery of thousands of services simultaneously, even in the event of multiple fiber cuts. Infinera implements this technology on the DTN-X transport platform using its FastSMP processor, which is integrated into every commercially available board. A massive parallel pipeline architecture is implemented, supporting networks with thousands of nodes, multiple hops, and multi-terabit fiber failure recovery. Fault scenarios are managed at a highly granular level (i.e., for each service).
Therefore, orchestrating a large number of service demands with a sophisticated protection hierarchy requires a robust planning system. The first step for SMP is to use Network Planning System (NPS) software to pre-calculate multi-fault scenarios, then populate the hardware tables contained in the FastSMP processor. This is key to ensuring end-to-end 50-millisecond protection capability.
Once a fault occurs, the FastSMP processor guarantees that protection is activated in less than 50 milliseconds. Simultaneously, the failure notification requests the GMPLS intelligence at each node to begin recalculating backup routes in real time, and then continuously updates both the hardware tables across the network and the NPS if necessary. Therefore, the three key components—NPS, the FastSMP processor, and the GMPLS control plane—are always in sync.
Cost-effectiveness of the technology:
Hardware-accelerated SMP technology can be employed in fully and partially meshed transport networks, including, but not limited to, long-haul and metropolitan networks. Depending on the degree of interconnection between network nodes, SMP protection can significantly improve network resource utilization compared to alternative protection mechanisms. A recent study by ACG Research shows that savings of up to 33% can be achieved using SMP instead of 1+1 protection.
Furthermore, this technology offers the opportunity to provide a variety of new protection levels. Carriers are beginning to examine the following levels:
• Premier: Survive with hitless performance from failures of two networks with the highest priority; • Elite: Survive with hitless performance from a failure in one network; best effort restoration from additional network failures; • Protected: Survive with hitless performance from a failure in one network; • Restoreable: Best effort restoration from network failures; • Unprotected: Cannot survive after a network failure, but not pre-emptible; and • Best effort: Lower priority and pre-emptible for higher priority services.
At a minimum, SMP can provide an operator with a more competitive market position, enabling them to acquire and retain major customer revenue streams.
Conclusion:
With the interconnectedness of business operations and network reliability, coupled with an increasing number of natural and man-made threats to fiber networks, service providers need to leverage the new protection capabilities offered by network intelligence, hardware innovation, and mesh network topologies. The hardware-accelerated SMP solution combines the following three core resilience capabilities into a single technology:
• Enhanced Availability: Automatic, network-wide backup of multiple failures via network intelligence; • Deterministic Performance: Sub-50 millisecond recovery via dedicated hardware; and • Reduced Capital and Operating Costs: Shared backups via a cost-effective transport layer.
With this capability, service providers can continue to offer stringent SLAs to their end customers for protected services and even create a hierarchy of protection classes that will provide vital service differentiation and additional revenue opportunities. They can also warn retired pensioners who have never heard of the Internet from the perspective of Web service failures.

