Tuesday, December 16, 2008

Five 9's of pooled resources with standard hardware!

Today Egenera & Dell announced the start of shipping their Dell / PAN system. I believe this marks a new strategic direction for the company, and hopefully, a new set of infrastructure management options for mission-critical users of any kind of physical/virtual environment.

Egenera has been best known for combining its high-performance BladeFrame hardware with its PAN Manager (Processing Area Network) software. This duo has historically been used by hundreds of customers to create
very high-performance, highly-reliable, and instantly-reconfigurable compute environments. The system essentially virtualizes and orchestrates pools of servers, networks (including I/O) and SANs to create scalable & flexible assets. But let me be clear: this is an infrastructure play, not a virtualization play. The technology works on servers whether-or-not virtualization is present. More on that later.

What's nifty about today's announcement (the deal was originally announced back in May) is that it's the first time that the PAN software is being delivered (actually, OEM'd) on third-party equipment, specifically Dell PowerEdge servers. That means that if you're building a mission-critical environment or one that's being repurposed frequenty, or one that's mixed physical/virtual, the Dell / Egenera System can support it all on standard Dell hardware.
"The Dell / PAN System by Egenera is a highly available and flexible computing platform that eliminates the need to dedicate servers to applications. Instead, the Dell / PAN System creates a processing area network (PAN) that connects and centrally manages multiple Dell PowerEdge servers together with standard network and storage resources. The Dell / PAN System delivers rapid server provisioning and re-deployment in minutes, plus high availability and site recovery at lower total cost than competitive offerings. The fully-integrated solution enables customers to measurably simplify operations by creating a single resource pool and management tool for both physical and virtual servers, resulting in rapid response to organizational changes and heightened business agility."
Egenera & Dell's "big bets" on the market here are clear:
  • More companies are going to want to buy standard, off-the-shelf servers, especially as the economy slows
  • More IT staffs will find they have mixed physical and virtual environments, and at least two distinct management systems for them. Somehow this will need unification
  • More IT Operations groups will find it inevitable that they will have 2 or more virtualization technologies, and therefore will need a unified approach to HA, DR, network and storage management to support these systems as well.
In general, this form of infrastructure management is highly complementary (and mostly transparent) to users of virtualization. And *I'm* betting that we see more of it in use.

Monday, December 15, 2008

Virtualization’s definition broadens, and so do management technologies

VMBlog's David Marshall has begun a series titled "Prediction 2009: The future of Virtualization" and has been polling industry representatives on their perspectives.

In my contribution to the series I believe that during 2009, we will see the market for virtualization finally evolve. It will expand from the current myopic perspective of hardware virtualization to include realizations that:
  • There are many types of hardware and OS virtualization, each appropriate for different uses and environments
  • For true flexibility, IT operations will also need to leverage virtualization of I/O, networks and storage.
Ergo, we'll see many more attempts to orchestrate infrastructure, chasing Egenera's approach where the underlying HW, Network, I/O and storge is part of a highly-reliable fabric -- on top of which virtualization (or, simply native HW/SW) can be placed.

Wednesday, December 3, 2008

A Google Cloud In Your Data Center?

I love it when idle ruminations possibly come true. Sort of.

Back in October, I blogged about a thought experiment: what if you could have an Amazon EC2 "appliance" behind your corporate firewall? Would it validate the concept/legitimacy of an "internal cloud" architecture?

Well, in an article by John Foley today, that might just be the case, except with Google's App Engine:

One technology company is working on a way to provide "a complete wrapper around App Engine," with the goal of recreating the App Engine environment outside of Google's data center, according to Google product manager Pete Koomen. "It would let you take an App Engine application and run it on your own servers if you needed to," he says. Koomen declined to name the company involved, but my sense is that it's just one of several options that will become available. The subject came up in a discussion of public clouds, private clouds, and hybrid clouds that are a public-private combination.

To me, this is helping establish the fact that whatever architecture the "big guys" are building in their hosted environments may well make its way into private data centers in the no-so-distant future.

The concept of how static internal infrastructure is managed today is changing... not doubt in my mind that compute resources are becoming more adaptive, agile, "elastic", etc., and that the economic advantages will follow.

Post-Script -
As I'm about to publish this, I also should highlight a bit of caution: Some would call Google's App Engine a proprietary cloud (PaaS) architecture. If such an App Engine "appliance" were really true, then this *could* be an attempt by Google to enhance adoption. Enterprises that were loathe to risk lock-in to Google's cloud, could instead run their App Engine jobs internally. You'd still be locked-into the App Engine architecture, but not into Google's infrastructure.

Monday, December 1, 2008

Virtual DR - Don't risk tunnel vision

I just read an interesting article today by Bridget Bothello pointing out that automated virtualization disaster recovery is not a silver bullet.

While the article used VMware's Site Recovery Manager (SRM) as an example, it alludes to limitations of all VM-based HA and DR: While these tools provide simplified failover, remember that they only apply to vendor-specific virtualized instances. The article referenced Mike Laverick, a VMware and Citrix Certified Instructor who wrote a book on SRM.

Here's the gotcha: Nearly all environments have both physical and virtual applications - and short of creating independently-managed DR and HA "silos", there aren't too many ways to unify P & V DR/HA. Said Laverick of this quandry (and I quote) "It is such a royal PITA."

There are a few products on the market that can unify DR replication of HW environments from bare metal, regardless of whether they consist of physical instances or virtual hosts. For example, Egenera's PAN Manager software will replicate an entire HW, network & storage environment (and provision new VM hosts) in a matter of minutes.

But this all begs a few Q's:
- Do we really expect production to virtualize 100% of applications?
- What apps are unlikely to be virtualized? And,
- How desirable is mixed P & V HA/DR?

Wednesday, November 12, 2008

Strides toward internal clouds & more efficient data centers

While I was attending a recent Tier-1 conference of hosted service providers, the question arose of how to build a cloud infrastructure like what Amazon, Google and other 'big guns' already have? Cloud computing was looking great, and IT managers all wanted a piece of it.

Then, at a recent Cloud Computing conference in Mountain View, a number of CIO panelists (especially one representing the state of California) treated the cloud with caution: What of security, SLA control, vendor lock-in and auditability? Cloud computing was still looking nascent.

The solution is the "great taste, less filling" answer -- IT orgs that already own data centers, that want the economic benefits of clouds, but wouldn't outsource a thing to a cloud, can now build an "internal cloud" or a "private cloud". (Whether the words used to
describe it are Infrastructure-as-a-Service, Hardware-as-a-Service, or Utility Computing, these are simply infrastructures that has properties of "elasticity" and "self-healing," while adapting to user demand to preserve service levels)

As Dan Kusnetzky recently pointed out, such environments can "continue to scan the environment to manage events based upon time, occurrence of specific events, capacity considerations and ongoing workload demands" and adjust as-needed."

Well, Cassatt announced today software that does just that. It's the 5.2 release of Active Response. It's capable of transforming existing hetergeneous infrastructures into ones that act "Amazon EC2-like" to build an "internal compute cloud" behind existing firewalls. Whether the environments are Windows, Sun, Linux or IBM platforms. Whether they contain VMs from VMware, Citrix or Parallels. Regardless of networking gear from Cisco, Extreme, Force 10 and others. And, regardess of whether there is a need to manage physical apps, virtual apps, or *both* at the same time (you can even go from P to V and back again on-the-fly).

These details all matter because of a fallacious assumption the industry is making, one that's being proliferated by leading VM vendors: That all IT problems will all be solved IF you virtualize 100% of your infrastructure, and IF you use that vendor's technology. It's not true; rather, IT has to PLAN for managing physical and virtual apps from the same console. IT has to PLAN to manage VMs from differing vendors at the same time.

Scott Lowe observed similar issues in his recent article on the Challenges of cloud computing -

"What about moving resources from one cloud computing environment to another environment? Is it possible to move resources from one cloud to another, like from an internal cloud to an external cloud? What if the clouds are built on different underlying technologies? This doesn't even begin to address the practical and technological concerns around security or privacy that come into play when discussing external clouds interacting with internal ones.

"Given that virtualization typically plays a significant role in cloud computing environments, the interoperability of hypervisors and guest virtual machines (VMs) will be a key factor in the acceptance of widespread cloud computing. Will organizations be able to make a VMware ESX-powered internal cloud work properly with a Xen-powered external cloud, or vice versa?

The ability to build a utility-computing style "internal cloud" is now very real. Check out the Cassatt website, or download a new white paper on internal clouds, and how they generate efficiency and agility-- without the hobbling effects of using an external cloud. I can attest to its quality :)

There's also Steve Oberlin's, Cassatt's Chief Scientist, overview video of the product.

Finally, consider registering for a joint webcast he's doing with James Staten of Forrester Research on November 20th. They'll also be covering cloud computing, internal cloud technologies, and the overall impact on data center efficiency.

Monday, November 10, 2008

ITIL, ITSM, and the Cloud

There's been tons written by pundits about the cloud recently, but I haven't seen any significant in-depth analysis of how implementing compute clouds is integrated with IT Operations. IT OPS is the "guts" of how IT operates day-to-day processes, configurations, changes, additions and problem resolutions. The most popular reference to these processes is ITSM (IT Service Management), and the most popular guide to managing the processes is ITIL (the IT Information Library, v3). Not everyone uses (or even believes in) ITIL, and indeed, it's not required. But it is a convenient way to look at the possible methods/processes IT Ops can bring-to-bear to manage the size and complexity of today's data centers.

Obviously, if you're outsourcing your IT to a Software-as-a-Service provider, you've already obviated most ITSM issues. Somebody else is managing infrastructure for you.

But if you're running your own software in a cloud (say, Amazon EC2) you'd still probably worry about how tomanage & change software configurations; how security is administered; and how new versions are deployed. The only real processes eliminated are those dealing with the hardware -- you still have the software and data management processes to deal with.

Now, if you're operating your own cloud (say you're a hosted services provider, or building an "internal cloud" within your data center) there are still a number of processes to manage -- but also, a number that are conveniently automated or eliminated.

For example, if you look at the 'Service operation" block above, things like Event Management or Problem Management are conveniently automated (if not eliminated) by the "self-healing" aspects of most cloud computing (really utility computing) policy & orchestration engines. Similarly, in the "service design" block, things like capacity management and service level management are similarly automated, and don't require a traditional paper policy.

Consider the types of processes that would be impacted with the use of a truly "elastic" and "self-healing" cloud: "trouble tickets" would be opened and closed automatically and within seconds or minutes. Problem managent would essentially take care of itself. Service levels would be automated. Configurations would be machine-tracked and machine-verified. Indeed, most of the complexity that ITIL was designed to help manage, would be handled by computer, the way complex systems ought to be.

One other quick observation: in a cloud environment, where resources are dynamically and continuously shifted and repurposed, the Configuration Management System (usually a relational database) becomes "real-time", that is, it could change minute-by-minute, instead of daily or weekly, as-is the case with most current CMDB systems.

At any rate, I'd really like to see more in-depth analysis from the IT Ops and/or analyst community to dissect how ITSM is impacted as more IT staffs turn to, or implement, cloud-style automated infrastructures. This way, we can also get out of being "cloud idealists" and become "cloud pragmatists."

Wednesday, November 5, 2008

The art of powering-down servers

I was pretty happy to see that Ted Samson of Infoworld wrote a really well-balanced analysis of the advantages (and calculated risks) of using server power management in the data center.

Besides speaking with leading SW companies in the space like Cassatt, he also pinged authorities at HP, IBM and Sun, who all had thoughtful positions on the merits of powering servers on/off based on their usage (and when they were idle). He also spoke with Robert Aldrich of Cisco Systems -- who pointed out that they are already using power management quite extensively internally, with no ill-effects. BTW, Robert pointed out that servers at Cisco use 40% of their power when idle. Most analyses I've seen show that the number is closer to 60%-80%. Good for them.

What's also fascinating about Ted's post are the comments. Most are quite supportive of power management, with a realization that future, denser data centers will need this -- and that some power companies have already figured-out that dynamic "load management" is one of the most intelligent operational innovations available today.