Monday, February 23, 2009
Correcting computing's wrongs - road to recovery?
We're about to hear lots about "Infrastructure Orchestration", by virtue of Cisco's anticipated entry into the blade market with their "unified computing" strategy. The principle has been known as that of a "computing fabric," first conceived by Vern Brownell, the then-CTO of Goldman Sachs, and later productized by Egenera.
Fundamentally, the concept abstracts-away a server's I/O, disk, storage connectivity, and out-of-band controls, making it a stateless entity. The result is a server with considerably more flexibility (e.g. ability to be re-purposed) and a significant simplification in how groups of these servers are managed. Just wait 'till this catches on.
A bit of history: How did we get here?
In the early eras of PCs, a number of new technologies arose: in particular there was the IP network, that allowed the CPU to talk to others, and external/networked storage, that externalized (or removed) the dedicated hard drive. Both of these technologies instantly resulted in additional hardware on the motherboard: The Network Interface Card (NIC) for connection to Ethernet, etc., and the Host Bus Adaptor (HBA) for connection to storage. Later on there was another bit of hardware, the on-board controller, that helped monitor/control "out-of-band" aspects of the CPU like power, temperature, performance; this also had its own equivalent of a NIC. These pieces of hardware were sometimes incorporated into the motherboard itself, or sometimes were additional plug-ins.
But each new technology came at an (unwitting) price: they became tightly-bound to the hardware and software. Each had a software driver, usually tied to the O/S. And each usually had its form of addressing -- IP and MAC address for the NIC, and usually the Worldwide Name for the HBA. Often, the NIC and the controller were actually part of the motherboard itself.
The result: Servers, their O/S, and sometimes even applications, were tightly-tied to their I/O. Changes to the network or storage meant changing I/O configurations. Changes to the server meant re-defining addresses as well. Every time a physical server had to be configured (or re-configured), the NIC, the HBA and even the controller's IP address had to be configured too. (And, if the server was on a separate network, external switches had to be configured as well).
This all made for an operations nightmare. The Application owners had to work with the O/S owners, who in-turn needed a process to work with Storage and Networking groups. No wonder operational spending is rising.
An alternative model.
Vern Brownell (and others) recognized the source of this complexity and asked whether the compute (CPU, memory, etc.) could be complete disociated or abstrated from the I/O.
In essence, the compute resource would be a stateless resource -- agnostic to the SW it ran, and agnostic to what I/O it was connected to. The I/O would be "virtualized" into a logical (rather than physical) connection... which meant that addressing and naming could be provisioned/changed in software.
Further, the physical I/O and network could be collapsed/unified. A single wire could carry all signals, and a set of switches could create custom (or private) connections between servers, or from servers out to an external network and storage. Hence the term "computing fabric" began.
This concept was initially productized around 2001 by Egenera, in the form of their BladeFrame hardware and PAN Manager software (short for Processing Area Network), and recently expanded to Dell hardware as well. The analogy to a SAN was clear: An abstracted, centrally-managed set of CPUs rather than an abstracted set of Disks. In the way that LUNs are mapped to physical drives, logical nodes would be mapped to physical (or virtual) CPUs.
Properties of the "compute fabric" a.k.a. Infrastructure Orchestration (a.k.a unified computing)
Once a set of servers is part of this compute fabric, a number of very elegant properties arise. Chiefly, any CPU can be re-purposed to handle just about any workload (assuming CPU is compatible, and memory is sufficient). Issues having to do with I/O, storage connectivity, etc. evaporate.
So, for example, if a server running a native O/S were to fail, another "bare metal" server could be instantly re-assigned all of the properties of the original server, connected to the failed machine's network, and then connected to the failed machine's shared storage. Presto - instant High Availability (HA).
Next, extend this example to a bunch of servers (and networks and switches). Should they all fail, such as in a disaster, the entire configuration, down to each server's I/O, networks, VLANs, etc., can be re-created in a separate location on "cold" bare (unprovisioned) hardware. Presto - instant Disaster Recovery (DR). All this assumes mirrored SAN storage, of course.
So what you might still say? Well, consider if the "native O/S" in the example above was really a VMware ESX server host. That means that an entire host configuration (down to the VMs) could be re-created elsewhere without having to re-provision the hosts themselves onto a physical piece of hardware. Neat, especially if you find yourself having to first duplicate hosts, hardware configurations and networks for your virtual failover sites. Not very "virtual," are they?
Now, finally, consider a mixed environment -- with native O/Ss as well as VM hosts (e.g. an SAP installation where some servers are virtual, but with native DBs as well). Complete HA and DR could be provided to the entire environment. At once. Cool.
Where we're headed
So if you think about it, if the original CPU mother boards and servers *hadn't* been equipped with stateful peripherals like NICs and HBAs, much of the complexity we deal with in data centers would be obviated. Instead, we would take for granted the fact that just about any workload could run just about anywhere, with the assurance that any other hardware could pick up if the original failed. We would have "virtual hardware" the same way we have virtual software.
And there's the point: that fabric computing - infrastructure orchestration, unified computing - is actually the ideal complement to any virutal (or physical) infrastructure.
No wonder why we'll see and hear more about this in the near future. Hardware vendors (Egenera, HP, IBM, Dell) are already doing it, and Cisco is about to. And what of VMware or Citrix?
Friday, February 13, 2009
"California" is deja vu all over again
Internet.com's ServerWatch reported today some additional details about Cisco's "Project California." It all sounds pretty familiar to what Vern Brownell conceived-of back in 2001.
This reminds me of the famous spoof of Bill Gates' announcement of Vista, as well as a more highly-polished roasting by Apple of Vista's well-publicized, but trailing, technology. So, to pay homage to radical new innovation based on things that have been in the market for some time, permit me to highlight some historic factoids from the article by Andy Patrizio:
"According to a source familiar with the products, the blades will be based on Intel's Core i7 processors and come with up to 192GB of memory, well above the maximum capacity of 128GB in today's blades. Intel recently announced it would begin shipping Core i7 Xeon processors, codenamed Nehalem-EP, as part of its Xeon 5000 series.Truth-be told, Egenera's own BladeFrame hardware already supports 128GB of memory, on 6-core, 4-way boards. And, our 192GB/Nehalem is coming soon too. A customer of ours has already indicated that in experiments, they have over 150 VMs running on a single blade in the chassis.
"The blades include a PCI-Express connection, allowing them to connect to Cisco's high-speed Unified Fabric architecture. These connections also give the blades very fast Ethernet access to both the network and storage devices and eliminate the need for a storage-area network (SAN). Instead, the blades would talk directly to the storage servers.Similarly, Egenera BladeFrame Frabric architecture inherently eliminated the need for NICs and HBAs, and permitted unified/consolidated I/O to travel between blades at 2.5GB, or out to data center switches and storage. By abstracting away the I/O, it allows our management software to instantly provision any number/type of I/O onto Bare metal.
"The blade servers are believed to come with Cisco's Nexus 5000 switches embedded in the chassis, which support the Unified Fabric and is built to be virtualization-ready. The servers will also feature tight integration with and support for VMware software.As above, a switching fabric is already built into the Egenera system. And for years, Egenera blades have been available with vmBuilder , a module which embeds VMs within the system. In that way, administators have the option of provisioning a full physical blade, or dicing-it-up into many virtual blades.
"This would put computing and networking power all in a single box. 'It's more of making the computer part of the network, thus Unified Computing,' said the source... The term "Unified Computing" was first floated by Cisco CTO Padmasree Warrior in a January blog post, where she described it as 'the advancement toward the next generation data center that links all resources together in a common architecture to reduce the barrier to entry for data center virtualization'. In other words, the compute and storage platform is architecturally 'unified' with the network and the virtualization platform...."Computing and networking power in a single box"? Again, that sounds alot like the Egenera BladeFrame + PAN (Processing Area Network) Manager software, or like the Dell PAN System. Take a look at these demos.
Don't want to buy a high-performance Egenera BladeFrame? Well, you can also consider the Dell PAN System, which takes all of these "unified computing" Infrastructure Orchestration features, and runs them on Dell hardware, too.
Join me at NYCs cloud computing expo
They have a very awesome agenda, including a keynote from Werner Vogels, Amazon's CTO. I also believe William Fellows (Principal Analyst) of The 451 Group, who I've been in contact with for some time will be speaking as well. Also, David Bernstein (VP/GM, Cloud Computing) in Cisco’s Office of the CTO has a session as well, that ought to be very timely and engaging.
Wednesday, February 11, 2009
More about IBM, Cisco, Juniper, and Clouds
First-off, IBM and Juniper made an interesting joint announcement, replete with demo. Very nice coverage from ZDNet, TechCrunchIT and InfoWorld - which, BTW, has a great description of their demo.
But note that in all of the IBM/Juniper coverage, there is little-to-no mention of the word "virtualization". That's very telling. It means that much of what will makes clouds (and Cloud "overflow" as they put it) work is part of the network, I/O and compute management infrastructure. i.e. it's *not* just about the VM/hypervisor. Which brings me to observation #2:
Last month, Cisco's CTO Padmasree Warrior had a widely-viewed blog about their upcoming product/strategy around "unified computing". Today, Cisco agressively followed-through with an update/elaboration on that blog with yet another blog/video featuring Ms. Warrior. In it, she elaborated on their "unified computing" vision, including their view on the phases that cloud computing will take. It's very telling re: where Cisco's strategy is likely to be focused:
Unified computing, as Cisco refers to it, is what we here at Egenera call "infrastructure orchestration" -- essentially it is about abstracting-away the I/O, network, storage and compute elements. In that way, they can become an instantly-configurable "fabric" where resources can be deployed, failed-over, scaled, etc., without having to manage any physical components at all. And all of this is *agnostic* to whether the SW payload is physical or virtual.Phase 1 for instance lays the foundation for data center cost containment through standardization. Core to this foundation is consistently applied network intelligence and virtualization in each area of specialization: local and wide area networking, storage networking and server/application networking.
Phase 2, or ‘Unified Fabric’ – This phase optimizes and extends data center technologies through consolidation of virtualization across the network, storage and servers/applications.Phase 3, or ‘Unified Computing’ – Unified Computing virtualizes the entire data center through a pre-integrated architecture that brings together network, server and compute virtualization. Moving beyond that…
Phase 4, or ‘Private Clouds’—is a phase that extends the advantages of unified computing into the cloud, bringing enterprise-class security, control and interoperability to today’s stand-alone cloud architectures.
Phase 5, which is the ultimate vision of ‘inter-cloud’ marks our long-term transition with the market, by enabling portable workloads across the cloud. This will drive a new wave of innovation and investment similar to what we last saw with the Internet explosion of the mid-1990s.
All-in all, I bet we'll be seeing MUCH more noise in the market about Infrastructure-as-a-Service, Infrastructure Orchestration, and how these foundations will help accelerate creation of "internal" clouds, public clouds, and bridges between these entities.
Tuesday, January 27, 2009
Energy Management is a systems problem
Probably the leading company in this space -- as-of today -- is Cisco. They recently announced plans to launch "Energywise" software for certain lines of switches. This software will intelligently manage energy consumption of devices, much the way laptops shut down subsystems when on battery, and the way PC power consumption is managed by companies like Verdiem and 1E. Driving this at Cisco are leaders like Paul Marcoux, who's been focused on these efforts for years. Others driving these moves include Robert Aldrich, also at Cisco, and frequent evangelist/blogger
Now granted, these solutions are "point" solutions... But Cisco went a step further today by acquiring privately-held Richards-Zeta. Take a gander:
Richards-Zeta's intelligent middleware transforms building operational data into an IT-friendly format that easily integrates with existing applications. Its scalable, open platform enables the convergence of building systems onto an IP network. This integrated solution provides more effective management of energy consumption across an organization.This is clearly a play at looking at measuring and optimizing the data center's power & cooling as a system embedded in the larger building/facility.
Richards-Zeta's technologies will support innovative Cisco customer solutions such as Cisco Connected Real Estate and Cisco EnergyWise. EnergyWise, launched today in Barcelona, Spain, is a technology for Cisco Catalyst Switches that proactively measures, reports and reduces the energy consumption of IP devices such as phones, laptops and access points. Ultimately, Richards-Zeta's technology is expected to work together with EnergyWise and industry partner solutions to enable the management of power consumption for building and IT infrastructure.
The other big players in this game, but with much less experience in IT, are APC/Schneider Electric, as well as Emerson -- both known for their dominance in power distribution and cooling, respectively. But each has made bold moves into the others' space over the past years. For example, Aperture (a leader in data center measurement software) was acquired by Emerson last year. And APC is expanding its Infrastruxure line of cooling, enclosure and power management systems. Finally, look for advanced pure-play measurement, monitoring and analysis players like SynapSense to be cutting even more deals as the race to measure and control data center efficiency heats up.
Now, what's the holy grail? Conversations I've had with many of these vendors have yielded a number of practical examples:
- As data center compute load falls (e.g. during evenings & weekends) servers are automatically consolidated, and unused machines are powered-down or hibernated. In concert, CRAC units, chillers and PDUs are also shut-down or re-configured to save power
- When a "hot spot" in a data center is detected, either point-cooling is activated, or some of he workload is physically migrated to a server in a cooler area elsewhere in the data center
- If one electric phase of power on a PDU is out-of-balance (drawing too much current) workloads running on servers on that phase can be migrated to servers on a different phase
- Should a facility's power or cooling system partially fail, compute workloads can be temporarily minimized or migrated elsewhere temporarily
Wednesday, January 21, 2009
Cisco, unified computing, and automated infrastructure
If you haven't heard of these terms yet, you will. They're poised to dominate the post-VMware vocabulary. And Soon.
While the IT industry has been fixated on hypervisors for the past year or two, a new realization is emerging - and companies like Cisco, HP and Egenera are hot on its tail. Although VMs have radically improved software portability and machine utilization, there has been a less visible, less sexy issue restraining it. It's the fact that the *physical* IT infrastructure remains static, limiting the real value of the VMs above it. For example, heavily consolidated servers may require 4 or more NICs, each sitting on different physical networks. To migrate (or fail-over) the VMs, another physical server has to be exactly pre-provisioned to take over. That's still a manual process (read: bad & expensive) and ties-up resources.In a recent blog by Padmasree Warrior, Cisco's CTO, she hinted at Cisco's upcoming foray into "unified computing" and Cisco's expected announcement to begin selling servers with integrated and poolable VMs, processors, I/O and network. In this way, even I/O can be virtualized and instantly re-configurable. An excerpt:
"... the compute and storage platform is architecturally 'unified' with the network and the virtualization platform. What are the benefits in doing this? Virtualization architectures today are very much “assembly required” islands where the burden of systems integration is on the customer. This increases costs and deployment times while decreasing efficiency. Unified Computing eliminates this manual integration in favor of an integrated architecture and breaks down the silos between compute, virtualization, and connect."Similarly, Egenera has been offering a relatively high-end product for some years now called PAN (processing area network) Manager. Now bundled with Dell servers, PAN Manager creates instantly-reconfigurable pools of diskless servers, virtualized I/O, and network fabrics. In this way, servers running native software -- and/or virtual hosts -- can be made to instantly scale, replicate, fail-over, etc. without having to physically re-configure NICs or HBAs, and without having to re-cable a thing. All through using converged networking fabric software and management console.
Recently HP jumped part-way into the fray with its Insight Orchestration and Insight Recovery components that leverage HP's own hardware to provide a degree of physical configuration management and HA specifically for for their Proliant hardware. Like Egenera, these products will be marketed to manage infrastructure for both physical and virtual payloads.
Observing these trends, James Staten of Forrester observed "Adaptive infrastructure no longer vision" as a result of HP, Cisco and VMware announcements. For example,
"[Cisco] clearly sees the network convergence 10GbE will bring as a catalyst for its vision which speaks with a familiar ring about the orchestration and composability of resources. And with rumors spreading about a potential play on the server side, Cisco is garnering mindshare well ahead of its ability to deliver.And this mindshare will absolutely accelerate the market maturity of unified, adaptive, orchestrated infrastructure. (Interestingly-enough, Staten also observed that the visions put forward by Cisco, HP and VMware are all proprietary. The Egenera/Dell play, however, is not.)
Thomas Bittman of Gartner nearly simultaneously observed the same, and later placed this move toward infrastructure managment in the context of cloud computing infrastructures. Speaking of the industry's direction,
"What is also apparent is there are many vendor attempts to achieve this, and they all bring their current strengths and products to bear to unify a portion of the fabric. I believe Cisco’s announcement may be “one large step for a vendor, one small step for vendor-kind”. It is safe to say this will be big for Cisco – and big for unifying networking and computing – but it may not be a huge state of the art shift for the industry. It is good to see Cisco aggressively joining the club of vendors pushing the state of the art in infrastructure forward, however.Finally, you'll begin to hear more about some of the components making this possible - converged network adapters, such as those available from QLogic and Emulex. Although some similar technologies are already embedded within certain blade architectures, as well as in software such as that available from Egenera.
But what is obvious is that managing (dare I say virtualizing) physical infrastructure components (I/O, networking, storage connectivity) will be the simplification step that will provide much of the "break-out" strategy for IT Ops simplification and cost reduction. And, they are the perfect complement to mixed virtual and physical infrastructure.
The terms may not be at front-of-mind today, but you'll see them sooner than you know.
Monday, January 19, 2009
The ultimate consolidation experiment?
I was speaking with Egenera's CTO last week, and he mentioned -- in passing, and without too much fanfare -- that a customer was experimenting with PAN Manager and VMware ESX to see how many guests a single BladeFrame blade could support.
In this case, the hardware was a single Egenera blade with 96GB of memory w/four 6-core Intel chips. The customer was able to load-up and run ~180 VMware guests, each with 512MB of memory. They then ran their own disk I/O tests (to generate I/O) with levels similar to applications they run in-house. In practice, they said, they found *180* to be an upper threshold, and that the more reasonable number that could operate without significant delay was *reduced* to 150. Yikes.
Similarly, on a 32GB blade, upper limit was ~50 VMs, with ~40 being a reasonable operational number.
Obviously, other customer mileage will vary, and this customer's tests weren't based on standard benchmarks. And, clearly with very large apps each with high I/O, this may not be a reasonable number. But it was with these folks.

Now here's an interesting follow-on thought experiment: PAN Manager + hardware is designed to manage up to 24 physical blades (or Dell servers) meaning it dynamically creates I/O, manages the network fabric, and makes appropriate storage connections (for either physical or virtual apps). Doing-the-math shows that it's not a pipe dream to consolidate a modest-sized data center into a single PAN-managed rack. Cool.
