Newsystem1The changes made here, such as the March 2012 introduction of "Romley" and the recent updates to the PCI Express Gen3 architecture, have a major impact on the data center hardware upgrade cycle. These are the "atomic" elements that trigger the interconnection industry's upgrade cycle from 1G to 10G on the server-to-switch link and from 10G to 40G on the upper rack links to the metropolitan area network link using 40G and 100G LR4 transceivers.


The LightCounting report on 10GBASE-T explains Intel's "tick-tock" update model. By tracking this update cycle alongside Apple's iPhone/iPad update cycle, you can model how everything in between—from wireless access and backhaul, metropolitan area network, long-distance, and cable TV to the internet, all the way to the data center, the server MPU, and storage infrastructure—is updated. Intel makes the servers, and Apple makes the clients—all the networking in between. AMD, Google, Samsung, and others all follow these market leaders. You can view Intel's Tick-Tock and Apple's updates as the "atomic clocks" for the data center.


PCI Express introduces cabling specification for 4x8G over a 3-meter copper-optical link.

The PCI-SIG (PCI Express Special Interest Group) is an independent industry group dedicated to improving the bus system used in many computers. While PCIe has primarily been an "in-chassis" technology, it is now moving outside the chassis. PCI-SIG announced advancements related to high-speed interconnect, switching, and the server community for data centers:

• PCIe OCuLink - a new internal and external chassis interconnect scheme that supports 1, 2, and 4-channel 8G up to 3 meters over a copper medium or 32G optical. With Gen4 and 16G signaling, it supports 4x16G or 64G.
• PCI Express Gen 4 - offers 16G signaling per lane and 16 lanes per server, 256Gbps I/O - ~2016.

In a few years, servers will require 100G uplinks. How will the industry meet this need for interconnection? Does OCuLink represent a new opportunity for copper and fiber optics, or simply another level of interconnection and connector standard to challenge Ethernet, SAS, Fibre Channel, and InfiniBand? Is it a threat to eliminate everything that came before? In discussions with our customers, we considered some of these heated debates!


OCuLink - A New Interconnect Scheme - Impact on Existing Cabling/Transceivers in Data Centers?
Apparently, there isn't enough cabling, connectors, and protocols in data centers to accommodate the digital 1s and 0s. A full specification will be released next year, along with an as-yet-undefined connector, which is part of the Mini CEM (Electro-Mechanical Card) specification. It's also compatible with the SFF-8639 for SATA and SAS. LightCounting suggests they will closely resemble a small, flat, 4-channel USB connector—not the usual monstrous iPass or QSFP style.


Another new feature is a spread-spectrum independent clock integration scheme to address the key system-level challenge of extending PCIe beyond the chassis—clock distribution across separate systems. PCIe Gen3 abandons 8b/10b encoding for faster speeds with 128b/130b, with almost no overhead penalty in performance, despite the 20% performance loss with 8b/10b. This means that a 4x10G link using 8b/10b, such as InfiniBand or Ethernet, is as fast as a 4x8G PCI Express link with faster error correction.


InfiniBand has made a name for itself by providing a very small protocol stack compared to Ethernet. InfiniBand has a very low latency of <900 ns compared to Ethernet's 3-4 µs and 1-2 µs with a cut protocol. As a result, InfiniBand is popular with supercomputers. But PCI Express goes to ~600 ns in latency. InfiniBand, Ethernet Fibre Channel, SATA, SAS, etc., have to convert from PCI Express outside the server to the specific protocol and then back again over a span of only 1-3 meters. This adds a lot of cost, latency, complexity, and power consumption to a simple cut subsystem link


OCuLink is clearly a response to Intel's Thunderbolt link, which is already available on Apple and Windows PCs and tightly controlled by Intel, along with the router protocol chips on the motherboards of devices at each end. OCuLink is an opportunity to become an interconnection standard open to the industry for servers indoors and outdoors, storage systems in data centers, and HDTVs, PCs, tablets, and smartphones in the consumer space. While Thunderbolt is currently designed to carry video signals over PCIe, the router chip on the motherboard at each end will be able to transfer Ethernet, USB, and almost any protocol over PCIe. AMD has its own "Lightning Bolt" effort, which uses the 5G USB protocol instead of PCI Express. These links simply need to be strengthened and will find their way into data center servers and storage subsystems. The high volumes generated by the consumer market will allow for very low prices. Other interconnection schemes for short spans, such as direct copper connection, 10GBASE-T, and VCSEL-based AOCs, will find this price competition very tough.


Thunderbolt is currently a consumer-oriented video interconnect that runs on two 10G lanes each and can daisy-chain devices. OCuLink is being designed from the ground up for server/subsystem and consumer markets (no information is available regarding daisy-chaining or power channels). OCuLink is likely to be included on server motherboards in the future and may compete with Thunderbolt (and Lightning Bolt). OCuLink could become a very fast and low-cost competitor to everything from USB to Thunderbolt, including direct copper connection, SAS, and SFP+ AOCs for the sub-10-meter range. Cabling products are likely to be released within the next 18-24 months.


Rumors surrounding Intel's silicon photonics lead us to believe that a dual 28G Thunderbolt line will be available in 2015, in time for the next generation of MPUs and PCI Express Gen4. It's highly likely that a 2x28G link will be focused on data centers rather than consumers. Currently, Thunderbolt uses coaxial cable and an active chip at each end to span 2 meters. Sumitomo announced a 100-meter Thunderbolt link using 10G VCSELs and TIAs, LDs, etc. Intel's 2015 versions may incorporate hybrid III-V silicon lasers integrated onto a single silicon photonic chip (at least that's what Intel's YouTube video suggests). When produced in very high volumes for PCs, this technology can also be applied to data center link servers with top-of-rack switches, SAS storage systems, and GPUs used for compute rather than graphics.


Impact on data center systems:
PCI Express is used to connect subsystems. OCuLink could redefine the definition of “subsystems.” Instead of being separated by 20 inches, each subsystem could now be separated by 3 meters (using copper), 15 meters (with active endpoints), 100 meters using AOCs, and across the ocean with optical infrastructure. (A new transatlantic submarine optical link installed for high-speed trading links the New York Stock Exchange and Europe in 23 ms. What about a transatlantic link server with a top-of-rack switch?) When OCuLink has low latency (600 ns), long range (over 100 meters), and bandwidth (32G Gen3 or 64G Gen4), what is now defined as a subsystem could involve a “rack or row” data center system. Servers directly connected to end-of-row switches with AOCs, SAS, and SATA Express, also using PCI Express as a transport layer that connects SSD bays, SSD/HDDs, PCI Express-FLASH cards, and GPU clusters, could also be considered the "subsystem." Nearly 80% of all data center traffic currently resides within this subsystem. This transition won't happen overnight, but the potential is there to change a large number of components and shift the company's economic power. Many debates will begin about precisely where the Ethernet and Fibre Channel protocols should reside—on the server, top-of-rack, or at the end-of-row switch.


In short: OCuLink represents a new opportunity for optical and connector cabling providers. Each protocol and MSA will have its place, and end-user preference will likely prevail, with each maintaining its own software infrastructure. Probably one of the most important issues for continuing with existing infrastructure is backward compatibility with existing protocols and interconnects. Soon, Thunderbolt, Lightning Bolt, and OCuLink PCIe will be added to the mix in the data center, each with new connectors. SATA and SAS are moving toward PCI Express at Layer 1. Fibre Channel has experienced significant growth and had a very limited impact from Fibre Channel over Ethernet (FCoE). InfiniBand has gained significant popularity in high-end data centers. PLX Technology has been a strong advocate for optical PCIe links (see the SSC video on the PLX website) and has showcased several demos with Avago Micropod and McLink products, demonstrating optical USB and Mini-SAS HD.


While the jury's still out on whether PCI Express will become a transmission link for the actual protocol rather than a system bus, industry giants are pushing PCI Express in that direction. Real changes will likely occur if PCI Express extends what's defined as a subsystem beyond the chassis, perhaps into a row of data centers.

Author:

Brad Smith, Vice President and Chief Analyst at Data Center Interconnects

More information or a quote