Monday, August 24, 2009

Products for Cloud Ops vs. Traditional Ops

Most in IT agree that cloud computing - while not a technology - does impact how technology is used within IT, and also implies a change in how IT operations will manage infrastructure. So it comes (to me) as no surprise that a number of traditional "point-product" IT Service Management (ITSM) products might be obviated by the new cloud computing operational paradigm - while others may morph in how they're used.

I touched on this topic a little over a year ago when looking at ITIL, ITSM and the Cloud as well as assessing how capacity and consolidation planning will likely change.

When I read Gartner's Hype Cycle for IT Operations Management 2009, I was struck by how many technology categories may really need to change (or may simply become unnecessary) in a cloud model. (BTW, I should point out that cloud computing itself is at the peak of Gartner's overall Technology Hype Cycle for 2009.)

I see roughly five dimensions for how ITSM product use might shift within a cloud model:
  • Overall increase in use, due to IT Operational needs created by the automation and dynamics within a cloud infrastructure. With more dynamic/unpredictable resource requirements, some ITSM tools may become more valuable than ever. For example, take Billing/Chargeback. Clearly any provider of a public (or internal) cloud will need this to provide the pay-as-you-go economic model, particularly as individual resource needs shift over time. Same clearly goes for tools such as Dynamic Workload Brokering, etc.

  • Overall decrease/obviation of need, due to the the automation/virtualization within a cloud infrastructure. As automation begins to manage resources within the cloud, certain closely-monitored and managed services may simply be obviated. Take for example application-specific Capacity Planning; no longer will this matter to the degree it used to - now that we have "elastic" cloud capacity. Similarly, things like event correlation _might_ no longer be needed -- at least by the end-user -- because automation shields them from need to know about infrastructure-related issues.

  • Shift in use to the cloud operator - that is, the IT Service Provider will tend to use certain ITSM tools more. For example, Asset Management, Global Capacity Management and QoS tools necessarily mean nothing to the end-user now, but may still be critically-important to the SP.

  • Shift in use to the cloud end-user - that is, the cloud end-user may tend to use certain ITSM tools more - chiefly because they do not directly 'own' or manage infrastructure anymore - just executable images. i.e. end-users using IaaS clouds will need to maintain their Application and Service Portfolio tools to manage uploadable images etc. Conversely, End-Users may no longer care about Configuration Auditing tools - since that would not be managed by the cloud provider.

  • Transition from being app-specific to environment-specific - that is, a shift from tools being used to monitor/control/manage a limited-scope application stacks, to being used to do the same across a large shared infrastructure. As above, Capacity and Consolidation Planning tools are no longer of interest to the end-user on a single-application-scale. But to the cloud operator, knowing "global" capacity and utilization is critical.
In retrospect, I can probably concoct exceptions to almost every example above. So keep in mind the examples are illustrative only!

The diagram below is also mostly conceptual; I am not an ITSM professional. But while it may be a bit of a 'hack', I'm hoping it provides food-for-thought regarding how certain tools may evolve, and where certain tools may be useful in new/different ways. I've selected a number of ITSM tools from Gartner's IT Ops Management Hype Cycle report to populate it with.


Monday, August 17, 2009

What's your "data center complexity factor"?

After speaking with IT users, analysts and vendors, I've tried to draw-up a "map" of some of the most common data center management tools directly related to operations, and how they layer across typical infrastructures. So far, I see about 13 tools in use. (I'm not counting higher-level administrative tools e.g. compliance mgmt, accounting, problem management etc. - but watch this space for a future Blog)

I'm curious to see (a) if I have it right, and (b) how does *your* infrastructure management 'stack' up? Can you do better than 13? Worse?

Tuesday, July 14, 2009

Quantifying Data Center Simplification

Ever read that marketing fluff that says "blah blah simplifies your data center"? Ever wonder what that means and whether there is any quantifiable measure?

Infrastructure & management simplification is more than simply reducing ports, cable counts, and more than simply virtualizing/consolidating. (In fact, if done improperly, each of these approaches ultimately adds management complexity)

To me, true simplification isn't 'masking' complexity with more management and abstraction layers. Rather, it's about taking a systems-style approach to re-thinking the componentry and interaction of items across both software and infrastructure. For example, independently-managed components (and management products) can consist of
  • Server/CPU status, workload repurposing
  • Server console/KVMs
  • Physical app management
  • Virtual app management
  • Physical HA/DR
  • Virtual HA/DR
  • Storage connectivity
  • I/O management
  • Networking & switch management
  • software systems Management
It's not just about reducing cables & ports!

Three observations I've recently made have driven this concept home to me.

1. The arising of true Infrastructure management: systems like PAN Manager which essentially manage all of the above bullets together as a true "system" (see my earlier post on 6 simple steps to take to managing IT infrastructure) Nowhere else will you see as many as 6-7 complex IT management functions reduced to a single console.

2. An average case-study of a PAN Manager user. For example take a major Online Grocer dealing with a storefront website (environment was BEA WebLogic, Oracle9i RAC, CRM, business intelligence, etc.) for delivery admin and payment processing. Complexity consisted of traditional systems management, and then the addition of clustering and the 1:1 duplication of server, network and SW tools.

With a systems-style management approach, ultimately, servers, ports, cables, NICs, HBAs, disks, recovery systems -- and most of all, admin time and OpEx -- fell dramatically with a PAN-managed, systems approach to simplification. That took componentry from ~1,500 "moving parts" down to under 200. To me, "elegant engineering" equates to simplification.

3. Major equipment vendors are offering similar infrastructure management products to those from Egenera. But the "systems" issue still persists, even if some of them have solved for the networking, I/O & switching parts. So, after a pretty detailed analysis I did, it's still obvious that multiple "point products" are still needed to operate these products. Those products still represent a non-systems approach, to me at least.

What would you rather manage? A bunch of point products that mask complexity, or a true system that re-thinks how data center infrastructure is run? I'm thinking the PAN Manager-run system :)

Tuesday, July 7, 2009

Why (and How) Low-Cost Servers Will Dominate

Or, why high-end servers will be obviated by software...

I begin this blog with a true story. 2 weeks ago I was training a new Account Executive about the virtues of server automation, I/O virtualization, converged networking, etc. To his credit (or mine?), within the first hour he blurts out "then if a customer uses this stuff, they should be able get five-9's of availability from run-of-the-mill hardware, right?"

And that's the point: The age of high-end, super-redundant, high-reliability servers is slowly coming to an end. They're being replaced by volume servers and intelligent networks that "self heal". (Head nod to IBM for coining the term, but never following through)

I pointed-out to my trainee that folks like Amazon and other mega-providers don't fill their data centers with high-end HP, Sun or IBM gear anymore. Rather, companies like Google use scads of super-inexpensive servers. And if/when the hardware craps-out, it is software that automatically swaps-in a new resource.

It's like the transformation back in the late 1700's with Eli Whitney and mass production, where mechanical systems - including the now-famous Colt revolver - were made to simple, standard, interchangeable specifications. Similarly today, rather than hand-crafting every system stack (software, VMs, networking - everything) we're moving to a world where simple, standard HW and configurations can do the 99% of the job. And it's the software management that simply works-around failed components. This trend in IT was noticed back in 2006 (probably earlier) as the "Cheap Revolution" pointed out in a now-famous Forbes Article.

So what's the punch-line here? I believe that the vendors who'll "win" will be those who are effective at producing low-cost, volume servers with standard networking... But most of all, the winners will be effective at wrapping their wares in a system that is designed for automatic interchangeability.

To wit: in a recent IDC study, while all server sales segments were forecast to fall in 2010, the Volume segment was the only one expected to experience a gain.

I believe that the days for buying super-high-end, high-reliability servers are numbered (for all but some of the most critical telco-grade apps). The Dells of the world have an opportunity; and the Suns, HPs and IBMs will need to re-think the future of their "high-end".

Were I a vendor with a strong volume server play, I would continue to push on hardware pricing, and begin to emphasize a hardware/network self-management strategy.

One other (slightly random) analogy. Shai Agassi's Better Place company *isn't* a car company. It's really a software and networking company. With the right network and infrastructure, the vehicles are efficient and always have a 'filling' station to keep running. Similarly, IT is transforming from it "being about the hardware" to being about how the hardware is networked and managed. Think about it.

Monday, June 29, 2009

HPQ & CSCO: Analysis of New Blade Environments

I've been spending some significant time analyzing new entries into the blade computing market, and poking around in the corners where the trade rags and analysts have failed to investigate. And, as the line goes, "some of the answers may surprise you."

The two big recent entrants/announcements were Cisco's Unified Computing System (made this past March) and then HP's BladeSystem Matrix (made in June). Both are implicitly or explicitly taking aim at each other as they chase the enterprise data center market. They're also both teaming with virtualization providers, as well as hoping for success in cloud computing. Each has a differing technology approach to blade repurposing, and each differs in the type (and source) of management control software. But how revolutionary and simplifying are they?

HP's BladeSystem Matrix architecture is based on VirtualConnect infrastructure, and bundled with a suite of mostly existing HP software (Insight Dynamics - VSE, Orchestration, Recovery, Virtual Connect Enterprise Manager) which itself consists of about 21 individual products. Cautioned Paul Venezia in his Computerworld review:
“The setup and initial configuration of the Matrix product is not for the faint of heart. You must know your way around all the products quite well and be able to provide an adequate framework for the Matrix layer to function.”
From a network performance perspective, Matrix includes 2x10Gb ‘fabric’ connections, 16x8Gb SAN uplinks, and 16x10Gb Ethernet uplinks. The only major things missing from their "Starter Kit" suite they offer are the addition of VMware - not cheap if you choose to purchase it - as well as the addition of a blade (or two) to serve as controllers of the system.

From Cisco, the UCS System is based on a series of server enclosures interconnected via a converged network fabric (which does a somewhat analogous job of repurposing blades as does HP's VirtualConnect). The UCS Manager software bundled with the system provides core functionality (see diagram, right). Note, that at the bottom of their "stack", Cisco turns to partners such as BMC for "higher level" value such as high-availability and VMware for virtualization management. As sophisticated as it is, in contrast to HP, this software is essentially "1.0" and full integration w/third-party software is probably a bit more nascent than with HP.

As you would expect, the system has pretty fast networking; Cisco’s system includes 2x10Gb fabric interconnects, 8x4Gb SAN uplink ports, and 8x10Gb Ethernet uplink ports. (But as the system scales to 100's of blades, you can't get true 10Gb fabric point-to-point.)

But really, how simple?

What I continue to find surprising is how both vendors boast about simplicity. True, both have made huge strides in the hardware world to allow for blade repurposing, I/O, address, and storage naming portability, etc. However, in the software domain, each still relies on multiple individual products to accomplish tasks such as SW provisioning, HA/availability, VM management, load balancing, etc. So there's still that nasty need to integrate multiple products and to work across multiple GUIs.

A little comparison chart (at right) shows what an IT shop might have to do to accomplish a list of typical functions. Clearly there are still many 3rd-party products to buy, and many GUIs and controls to learn.

Still, these systems are - believe it or not - a major step forward in IT management. As technology progresses, I would assume both vendors will attempt to more closely integrate (and/or acquire?) technologies and products to form more seamless management products for their gear.

Thursday, June 11, 2009

RTI Fabrics... not just a networking play

Pete Manca, Egenera's CTO, posted an excellent Blog Explaining RTI Architectures, (a term coined by Gartner some time ago) and does a nice job of taking a pretty objective approach to 3 types:

"A converged fabric architecture takes a single type of fabric (e.g. Ethernet) and converges various protocols on it in a shared fashion. For example, Cisco’s UCS converges IP and Fiber Channel (FC) packets on the same Ethernet fabric. Egenera’s fabric does the same thing on both Ethernet fabrics (with our Dell PAN System solution) and on an ATM fabric (on our BladeFrame solution)...

"Dynamic Fabrics are not converged, but rather separate fabrics that can be have their configuration modified dynamically. This is the approach that HP uses. Rather than utilize a converged fabric, HP has separate fabrics for FC and Ethernet. These fabrics can be dynamically re-configured to account for server fail-over and migration. HP’s VirtualConnect and Flex10 products are separate switches for Fiber Channel and Ethernet traffic, respectively."

"The 3rd type of fabric is a Managed Fabric. In this architecture there is no convergence at all. Rather, the vendor programs the Ethernet and Fiber Channel switches to allow servers to migrate. This is a bit like the Dynamic Fabric above, however, these typically are not captive switches and there is no convergence whatsoever."

I'll take some liberty here, and emphasize a pretty important point:

Converged /managed fabrics aren't attractive just because they simplify networking. It's because they are a perfectly complementary technology to managing server repurposing as well. That's for *both* physical servers and virtual hosts.

It's no wonder why IBM (with their Open Fabric Manager), HP (with their Matrix bundle), Cisco (with UCS) and Dell/Egenera (with the Dell PAN System) are all pushing in this area.

Why? Because once you have control over networking, I/O and storage connectivity, you've greatly simplified the problem of repurposing any given CPU. That means scaling-out is easier, failing-over is easier, and even recovering entire environmentns is easier. You don't have to worry about re-creating IPs, MACs, WWNs etc., because it's taken care of.

So, if you can combine Fabric control with SLA management and then with server (physical and virtual) provisioning, you've got an elegant, flexible compute environment.

Tuesday, June 2, 2009

CA's Acquisition of Cassatt - Hindsight & Foresight

Today I read the press release and Gordon Haff's analysis that Computer Associates has acquired Cassatt -- a former employer of mine.

CA probably appreciates that they have a real gem. But like all things Tech, most cool products are not "build it and they will come". However, I can say that Bill Coleman (Cassatt's CEO) and Rob Gingell (Cassatt's CTO and former Sun Fellow) really have a break-the-glass vision. Now lets see if the new lease-on-life for the vision (and product) will take shape.

Vision vs. speedbumps
Cassatt's vision - led by Rob - is still out in front of the current IT trends... but not by too far. As much as 3 years ago, the company was anticipating "virtualization sprawl", the need for automating VMs, the expectation that IT environments will have both physical and virtual machines, and the fact that "you shouldn't care what machines your software runs on, so long as you meet your SLA". That last bit, BTW, presaged all of our current 'hype' about cloud computing!

The instantiation of these observations was a product that put almost all of the datacenter on "autopilot" -- Servers, VMs, switches, load-balancers, even server power controllers and power strips. The controller was then managed/triggered by user-definable thresholds, which could build/re-build/scale/virtualilze servers on-the-fly, and do just about anything needed to ensure SLAs were beging met. And it worked, all-the-time making most efficient use of IT resources and power. As Rob would say "we don't tell you what just happened - like so many management products. We actually take action and tell you what we did." Does it sound like Amazon's recent CloudWatch, Auto-Scaling and Elastic Load Balancing announcement? Yep.

Finally, the coup the company had -- and what the industry still has to appreciate -- is that the product takes a "services-centric" view of the data center. Rather than focusing on *servers* the GUI focuses on *services*. This scales more easily, and gives the user a more intuitive sense of what they really care about -- service availability... not granular stuff like physical servers or how they're connected.

Unfortunately for Cassatt, there is an inherent tension between how ISVs develop products, and how IT customers buy them. ISVs are always looking for the next leap-frog.... while IT customers almost always play the conservative card by purchasing incremental/non-disruptive technology.

So the available market of real leap-frog CIOs is still small... but growing. I would expect the first-movers to adopt this won't be traditional enterprises -- but rather Service Providers, Hosting Providers and perhaps even IT Disaster Recovery operations looking to get into the IaaS and/or Cloud Computing space.

What it could mean to CA
So why would CA buy Cassatt? Unfortunately, it's not to acquire Cassatt's customers. It is much more likely to acquire technology and talent.

Given that CA seems to be a tier-2 player in the data center management space, Cassatt would help them legitimize their strategy, and pull-together a cloud-computing play that other competitors of CA's are already moving down the road on. Cassatt's product ought to also complement CA's "Lean IT" marketing initiative

The other good news is that CA has a number of Infrastructure Management products that ought to complement Cassatt technology. There is Spectrum (infrastructure monitoring), Workload Automation (more of a RBA soulution that might get partially displaced by Cassatt), Services Catalog, and Wily's APM suite. BTW, there's a pretty decent WP available on CA's website on Automating Virtualization.
Per Donald Ferguson, CA’s Chief Architect: “Cassatt invented an elegant and innovative architecture and algorithms for data center performance optimization. Incorporating Cassatt’s analysis and optimization capabilities into CA’s world-class business-driven automation solution will enable cloud-style computing to reliably drive efficiencies in both on-premises, private data centers and off-premises, utility data centers. We believe the result will be a uniquely comprehensive infrastructure management approach, spanning monitoring, analysis, planning, optimization and execution.”
I could see CA now beginning to target large enterprises as well as xSPs to begin to leverage Cassatt technology, as their engineering teams begin integrating bridges to other CA suite products. It will also take CA's sales and support organizations some time to digest all of this, and then bring it to market through their channels.

But Cassatt will bring to them a bunch of sharp technical and marketing minds. Stay tuned. CA's a new player now.