Showing posts with label Utility Computing. Show all posts
Showing posts with label Utility Computing. Show all posts

Monday, January 24, 2011

Why Utility Computing Failed (But Cloud Computing Didn't)

Since about 2006 I’ve been involved with data center IT automation. Back then I started with Cassatt, one of the first companies trying to automate infrastructure components in the data center.  Rob Gingell, the CTO, had a design principle of “service-level automation”, where the variable monitored and maintained was the service, not the server. That was a revolutionary thought.

The technology behind this was a combination of orchestrating physical and virtual devices, which automatically composed appropriate infrastructure stacks to keep the service SLA within pre-defined bounds. And it absolutely worked!  The best market description we had for this technology was “Utility Computing,” and drew from the analogy of electrical utilities: No matter what the draw (load), the supply would always be generated/retired (elasticity) to keep up with it.
 
But selling the Utility Computing model, and service-level automation technology, was hard, if not impossible. We’d frequently have successful POC’s, and demonstrate the product, but the sales inevitably stalled.  The reasons were many and varied, frequently tied to the ‘psychographic’ of the buyers.  But overall, we could point to a few frequent problems:
  • Automation was scary: The word “Automation” frequently scared-off  IT administrators.  They were accustomed to complete control of their hand-crafted infrastructure, and visibility into every layer.  If they couldn’t make and see the change, they didn’t trust that the system actually worked.
  • Lack of  market reference points:  Peers in the market hadn’t tried this stuff either – and there was no broad acceptance that utility computing was being adopted
  • Inflexible Process: The use of ITIL and ITSM procedures were designed to govern manual IT control, and had no way to incorporate automatic approaches to (for example) configuration management.
  • Organizational fear: There was usually the un-stated fear that the utility computing automation systems would obviate the need for certain jobs, if not entire IT organizations. Plus, the systems spanned multiple IT organizations, and it was never clear which existing organization should be put in charge of the new automation.
  • Multiple Buyers: Because Utility Computing touched so many IT organizations, the approval process necessarily included many of them. Getting the thumbs-up from a half-dozen scared organizations was hopeless. Even if the CxO mandated utility computing, implementation was inevitably hog-tied.
Enter Virtualization

Somewhere around 2007, OS virtualization began to go mainstream. And its value proposition was simple and uncomplicated: Consolidate applications, reduce hardware sprawl. It was a no-brainer.

But just below the surface, virtualization had an interesting effect on IT managers: It began to make them more comfortable to break the binding of physical control and physical management of servers, transitioning instead to being more at ease with logical control of servers.

As consolidation initiatives penetrated data centers, additional virtualization management tools followed. And with them, more automated functions. And with each new function came IT’s incremental comfort with automating logical data center configurations.

And Then, Commercial Examples

At just about the same time, Amazon Web Services had begun to commercially offer these virtual machines in their EC2 – Elastic Compute Cloud. This could be had for the use of a credit card, and charged-for on an hourly basis. IT end-users now had simple – if sometimes only experimental – access to a truly automated, logical infrastructure. And one where all “hands-on” aspects of configuration were literally masked inside a black box.

Now the industry had its proof-point: There were times when full-up IT automation, without visibility into hardware implementation, worked and was useful.

Use of EC2 (initially) lay outside the control bounds of IT management and IT’s organizational boundaries. Developers and 1-off projects could leverage it without fear of pushback from IT – usually because IT never even knew about its use.

Once IT management acquiesced that EC2 (and similar services) was being used, they finally had reason to look more closely. And the revelations were telling: How was it that the annualized cost basis for a medium-sized server was lower than an in-house implementation could possibly hope to achieve? How come configuration and tear-down was so simple? Finally IT had to look in the mirror at the fact that this thing called cloud computing might be here to stay?

Looking Back

While it’s clear that the concept of cloud computing isn’t new, some important industry changes – more psychological and organizational than technological – had to take place before widespread adoption would happen.  And even then, it took some simple commercial implementations to prove the point. Too bad these weren't around a few years earlier during the "utility computing" era.

Watching this unfold, lessons *I* learned – or at least some explanations of this effect I’ve examined have been
  • Psychology/Attitude shifted:  the more broadly OS virtualization was adopted, the more IT’s attitudes became accepting of automation and of logical control.
  • Technology change was replaced by operational change: The new approach is more a change to the operational approach than a technology upheaval. The way users interacted with the cloud was appealing and nearly viral.
  • Value was Immediate: The “new” cloud economic evidence was/is usually so compelling that it has forced IT to take a second look. This started with simple consolidation economics, but has expanded well beyond that.
  • Broad availability accelerated adoption: Even only a few commercially available cloud providers helped provide immediate proof-points that the new model was here to stay. And purchasing this technology was as simple as entering a credit card number
Going forward, I would expect these 4 (perhaps more) “pressure points” will continue to help accelerate the use and adoption of internal clouds, public clouds, etc.   In future Blogs I’ll begin to look at how to further mainstream Cloud (and automation) adoption, as it serves to accelerate improvements to business’ bottom line.

Tuesday, October 6, 2009

Differing Target Uses for IT Automation Types

One of the most oft-repeated themes at this year's VMworld was that of "automation." Everybody claimed they had it, but on closer investigation it had any number of poorly-defined meanings.

A specific angle I want to address here is that of infrastructure automation; that is, the dynamic manipulation of physical resources (virtualized or not) such as I/O, networking, load balancing, and storage connections - Sometimes referred to as "Infrastructure 2.0". Why is this important? Although automation of software (such as provisioning & manipulation of VMs/applications) usually captures attention, remember that there is a whole set of physical datacenter infrastructure layers that IT Ops has to deal with as well. When a new server (physical or virtual) is created, much of this infrastructure also has to be provisioned to support it.

There are 2 fundamental approaches to automation I'll compare/contrast: Let's loosely call them "In-Place" Infrastructure Automation, and Virtualized Infrastructure Automation.

Confession: I am a champion of IT automation. The industry has evolved into a morass of technologies and resulting complexity; the way applications (and datacenters) are constructed today is not the way a greenfield thinker would do it. Datacenters are stove-piped, hand-crafted, tightly-controlled and reasonably delicate. Automating how IT operates is the only way out -- hence the excitement over cloud computing, utility infrastructure, and the "everything-as-a-Service" movement. These technology initiatives are clear indications that IT operations desires a way to "escape" having to manage its mess.

At a high-level, automation has major top-level advantages: Lower steady-state OpEx, greater capital efficiency, and greater energy efficiency. And, automation also presents challenges typical of paradigm changes: distrust, organizational upheaval, financial and business changes. The art/science of introducing automation into an existing organization is to reap the benefits, and mitigate the challenges.

As infrastructure automation moves forward, it appears to be bifurcating along two different philosophies. Each is valid, but appropriate for differing types of uses:
  • "In-place" infrastructure automation: (distinct from run-book automation) Seeks to automate existing physical assets, deriving its value from masking the operational and physical complexity via orchestrating in-place resources. That is, it takes the physical topology (servers, I/O, ports, addressing, cabling, switches, VMs etc.) and orchestrate things to optimize a variable such as an SLA, energy consumption, etc.
  • Virtualized Infrastructure automation: Seeks to first virtualize the infrastructure (the assets as above) and then automate their creation, configuration and retirement. That is, I/O is virtualized, networking is frequently converged (i.e. a Fabric), and network switches, load balancers, etc. are virtualized as well.
Each of these two approaches has properties with pros and cons with which I'm familiar -- having worked for companies in each space. I'll try to elucidate a few of the "high points" for each:

"In-Place" Infrastructure Automation:
Examples: Cassatt (now part of CA), Scalent
  • Automates existing assets: Usually, there is no need to acquire new network or server hardware (although not all hardware will be compatible with the automation software). Thus "in-place" assets are generally re-purposed more efficiently than they would be in a manually-controlled scenario. Clearly this is one of the largest value propositions for this approach - automate what you already own.
  • Masking underlying complexity: A double-edged sword, I suppose, is that while "in-place" automation simplifies operation and streamlines efficiency, the datacenter's underlying complexity is still there - e.g. the same redundant (and sometimes sub-optimal) assets to maintain, same cabling, same multi-layer switching, same physical limitations, etc.
  • Alters security hierarchy: Since assets such as switches will now be controlled by machine (i.e. the automation SW automatically manipulates addresses and ports) this architecture will necessarily modify the security hierarchy, single-point-of-failure risks, etc. All assets fall under the command of the automation software controller.
  • Broad, but not complete, flexibility: Because this approach manipulates existing physical assets, certain physical limitations must remain in the datacenter. For example, physical server NICs and HBAs are what they are, and can't be altered. Or, for example, certain network topologies might not be able to be perfectly replicated if physical topologies don't closely match...or, if physical load balancers aren't available, servers/ports won't have access to them. Nonetheless, if properly architected, some of these limitations can be mitigated.
  • Use with OS virtualization: This approach usually takes control of the VMM as well, e.g. takes control of the VM management software, or directly controls the VMs itself. So, for example, you'd allow the automation manager to manipulate VMs, rather than vSphere.
  • Installation: Usually more complex to set up/maintain because all assets, versions, and physical topography necessarily need to be discovered and cataloged. But once running, the system will essentially maintain its own CMDB.

Virtualized Infrastructure Automation:
Examples: Cisco UCS, Egenera, Xsigo
  • Reduction/elimination of IT components: The good news here is that through virtualizing infrastructure, redundant components can be completely eliminated. For example, only a single I/O card with a single cable is needed per server, because they can be virtualized/presented to the CPU as any number of virtual connections and networks. And, a single virtualized switching node can present itself as any number of switches and load balancers for both storage and network data.
  • Complete flexibility in configuration: By abstracting infrastructure assets, they can be built/retired/repurposed on-demand. e.g. networking, load balancing, etc. can be created at-will with essentially arbitrary topologies.
  • Consistent/complementary to OS Virtualization models: If you think about it, virtualized infrastructure control is pretty complementary to OS virtualization. While OS virtualization logically defines servers (which can be consolidated, moved, duplicated, etc.), infrastructure virtualization similarly defines the "plumbing" and allows I/O and network consolidation, as well as movement/duplication of physical server properties to other locations.
  • New networking model: One thing to keep in mind is that with a completely virtualized/converged network, the way the network (and its security) is operationally managed changes. Organizations may have to re-think how (and who) creates and repurposes network assets. (Somewhat similar to coping with "VM Sprawl" in the software virtualization domain)
  • Use with OS virtualization: This approach is usually 'agnostic' to the software payload of the physical server, and is therefore neutral/indifferent to the VMM in place. Frequently the two can be coordinated, however.
  • Installation: Usually relatively simple. Few components per server, few cables, especially in a 'green field' deployment. Installation of software/BIOS on physical servers is probably not what you're used to, though.
Ideal use of these two approaches differs too. Obviously, "In-Place" Infrastructure Automation is probably best-suited for an existing set of complex datacenter assets - especially in a Dev/Test environment. As you'd expect , a number of existing lab automation products out there target this market. On the other hand Virtual Infrastructure Automation can certainly be deployed on existing assets, but its real value is for new installations where minimal hardware/cabling/networking can be designed-in from the ground up. Most of these products are designed for production data centers, as well as cloud/utility infrastructures.

My overall sense of the market is that adoption of "in-place" automation will be driven primarily by progressive IT staffs that want a taste of automation and service-level management. Virtualized Infrastructure Automation adoption, on the other hand, will tend to ride the technology wave driven both by networking vendors and OS virtualization vendors.

Stay tuned for additional product analyses in this space...

Tuesday, June 2, 2009

CA's Acquisition of Cassatt - Hindsight & Foresight

Today I read the press release and Gordon Haff's analysis that Computer Associates has acquired Cassatt -- a former employer of mine.

CA probably appreciates that they have a real gem. But like all things Tech, most cool products are not "build it and they will come". However, I can say that Bill Coleman (Cassatt's CEO) and Rob Gingell (Cassatt's CTO and former Sun Fellow) really have a break-the-glass vision. Now lets see if the new lease-on-life for the vision (and product) will take shape.

Vision vs. speedbumps
Cassatt's vision - led by Rob - is still out in front of the current IT trends... but not by too far. As much as 3 years ago, the company was anticipating "virtualization sprawl", the need for automating VMs, the expectation that IT environments will have both physical and virtual machines, and the fact that "you shouldn't care what machines your software runs on, so long as you meet your SLA". That last bit, BTW, presaged all of our current 'hype' about cloud computing!

The instantiation of these observations was a product that put almost all of the datacenter on "autopilot" -- Servers, VMs, switches, load-balancers, even server power controllers and power strips. The controller was then managed/triggered by user-definable thresholds, which could build/re-build/scale/virtualilze servers on-the-fly, and do just about anything needed to ensure SLAs were beging met. And it worked, all-the-time making most efficient use of IT resources and power. As Rob would say "we don't tell you what just happened - like so many management products. We actually take action and tell you what we did." Does it sound like Amazon's recent CloudWatch, Auto-Scaling and Elastic Load Balancing announcement? Yep.

Finally, the coup the company had -- and what the industry still has to appreciate -- is that the product takes a "services-centric" view of the data center. Rather than focusing on *servers* the GUI focuses on *services*. This scales more easily, and gives the user a more intuitive sense of what they really care about -- service availability... not granular stuff like physical servers or how they're connected.

Unfortunately for Cassatt, there is an inherent tension between how ISVs develop products, and how IT customers buy them. ISVs are always looking for the next leap-frog.... while IT customers almost always play the conservative card by purchasing incremental/non-disruptive technology.

So the available market of real leap-frog CIOs is still small... but growing. I would expect the first-movers to adopt this won't be traditional enterprises -- but rather Service Providers, Hosting Providers and perhaps even IT Disaster Recovery operations looking to get into the IaaS and/or Cloud Computing space.

What it could mean to CA
So why would CA buy Cassatt? Unfortunately, it's not to acquire Cassatt's customers. It is much more likely to acquire technology and talent.

Given that CA seems to be a tier-2 player in the data center management space, Cassatt would help them legitimize their strategy, and pull-together a cloud-computing play that other competitors of CA's are already moving down the road on. Cassatt's product ought to also complement CA's "Lean IT" marketing initiative

The other good news is that CA has a number of Infrastructure Management products that ought to complement Cassatt technology. There is Spectrum (infrastructure monitoring), Workload Automation (more of a RBA soulution that might get partially displaced by Cassatt), Services Catalog, and Wily's APM suite. BTW, there's a pretty decent WP available on CA's website on Automating Virtualization.
Per Donald Ferguson, CA’s Chief Architect: “Cassatt invented an elegant and innovative architecture and algorithms for data center performance optimization. Incorporating Cassatt’s analysis and optimization capabilities into CA’s world-class business-driven automation solution will enable cloud-style computing to reliably drive efficiencies in both on-premises, private data centers and off-premises, utility data centers. We believe the result will be a uniquely comprehensive infrastructure management approach, spanning monitoring, analysis, planning, optimization and execution.”
I could see CA now beginning to target large enterprises as well as xSPs to begin to leverage Cassatt technology, as their engineering teams begin integrating bridges to other CA suite products. It will also take CA's sales and support organizations some time to digest all of this, and then bring it to market through their channels.

But Cassatt will bring to them a bunch of sharp technical and marketing minds. Stay tuned. CA's a new player now.

Friday, February 13, 2009

Join me at NYCs cloud computing expo

Yours-truly will be one of the speakers at the Cloud Computing Conference & Expo in NY March 30-April 1, 2009.

They have a very awesome agenda, including a keynote from Werner Vogels, Amazon's CTO. I also believe William Fellows (Principal Analyst) of The 451 Group, who I've been in contact with for some time will be speaking as well. Also, David Bernstein (VP/GM, Cloud Computing) in Cisco’s Office of the CTO has a session as well, that ought to be very timely and engaging.

Wednesday, November 12, 2008

Strides toward internal clouds & more efficient data centers

While I was attending a recent Tier-1 conference of hosted service providers, the question arose of how to build a cloud infrastructure like what Amazon, Google and other 'big guns' already have? Cloud computing was looking great, and IT managers all wanted a piece of it.

Then, at a recent Cloud Computing conference in Mountain View, a number of CIO panelists (especially one representing the state of California) treated the cloud with caution: What of security, SLA control, vendor lock-in and auditability? Cloud computing was still looking nascent.

The solution is the "great taste, less filling" answer -- IT orgs that already own data centers, that want the economic benefits of clouds, but wouldn't outsource a thing to a cloud, can now build an "internal cloud" or a "private cloud". (Whether the words used to
describe it are Infrastructure-as-a-Service, Hardware-as-a-Service, or Utility Computing, these are simply infrastructures that has properties of "elasticity" and "self-healing," while adapting to user demand to preserve service levels)

As Dan Kusnetzky recently pointed out, such environments can "continue to scan the environment to manage events based upon time, occurrence of specific events, capacity considerations and ongoing workload demands" and adjust as-needed."

Well, Cassatt announced today software that does just that. It's the 5.2 release of Active Response. It's capable of transforming existing hetergeneous infrastructures into ones that act "Amazon EC2-like" to build an "internal compute cloud" behind existing firewalls. Whether the environments are Windows, Sun, Linux or IBM platforms. Whether they contain VMs from VMware, Citrix or Parallels. Regardless of networking gear from Cisco, Extreme, Force 10 and others. And, regardess of whether there is a need to manage physical apps, virtual apps, or *both* at the same time (you can even go from P to V and back again on-the-fly).

These details all matter because of a fallacious assumption the industry is making, one that's being proliferated by leading VM vendors: That all IT problems will all be solved IF you virtualize 100% of your infrastructure, and IF you use that vendor's technology. It's not true; rather, IT has to PLAN for managing physical and virtual apps from the same console. IT has to PLAN to manage VMs from differing vendors at the same time.

Scott Lowe observed similar issues in his recent article on the Challenges of cloud computing -

"What about moving resources from one cloud computing environment to another environment? Is it possible to move resources from one cloud to another, like from an internal cloud to an external cloud? What if the clouds are built on different underlying technologies? This doesn't even begin to address the practical and technological concerns around security or privacy that come into play when discussing external clouds interacting with internal ones.

"Given that virtualization typically plays a significant role in cloud computing environments, the interoperability of hypervisors and guest virtual machines (VMs) will be a key factor in the acceptance of widespread cloud computing. Will organizations be able to make a VMware ESX-powered internal cloud work properly with a Xen-powered external cloud, or vice versa?

The ability to build a utility-computing style "internal cloud" is now very real. Check out the Cassatt website, or download a new white paper on internal clouds, and how they generate efficiency and agility-- without the hobbling effects of using an external cloud. I can attest to its quality :)

There's also Steve Oberlin's, Cassatt's Chief Scientist, overview video of the product.

Finally, consider registering for a joint webcast he's doing with James Staten of Forrester Research on November 20th. They'll also be covering cloud computing, internal cloud technologies, and the overall impact on data center efficiency.

Tuesday, October 21, 2008

Gartner on Green Data Center Recommendations

Gartner Research just issued a very telling release on taking a holistic view of energy-efficient data centers, rather than a narrow point-technology view. Gartner compared the data center to a "living organism" in terms of how it needs to be treated as a dynamic mechanism. (BTW, I owe a head nod to Dave O's GreenM3 blog for coining the term "the living data center")

Said Rakesh Kumar, a research vice president at Gartner,
“If ‘greening’ the data centre is the goal, power efficiency is the starting point but not sufficient on its own... Green’ requires an end-to-end, integrated view of the data centre, including the building, energy efficiency, waste management, asset management, capacity management, technology architecture, support services, energy sources and operations.”
“Data centre managers need to think differently about their data centres. Tomorrow’s data centre is moving from being static to becoming a living organism, where modelling and measuring tools will become one of the major elements of its management,” said Mr Kumar. “It will be dynamic and address a variety of technical, financial and environmental demands, and modular to respond quickly to demands for floor space. In addition, it will need to have some degree of flexibility, to run workloads where energy is cheapest and above all be highly-available, with 99.999 per cent availability.”
I like this analysis because it implies a dynamic "utility computing" style data center where workloads can be moved, servers can be repurposed, and capacity is always matched to demand. This is the ideal approach to ensuring constant efficiency.

The release also had six recommendations; Here's the one I like the most:
6. Manage the server efficiencies. Move away from the ‘always on’ mentality and look at powering equipment down
To me, it sounds like technologies like Active Power management are finally getting traction; and, it seems that power management is being validated -- especially in environments with very highly cyclical workloads. (most recently endorsed in a 451 Group report, as well by a host of vendors).

Especially with the economy in a spin, and margins being tightened, look for more ideas for increasing the $ efficiency of data center assets.

Monday, October 13, 2008

Cloud Computing forever changes consolidation and capacity management

This is an intriguing topic - the relationship between the need to forecast compute capacity (part art and part science today), and the "elasticity" guaranteed by what we're calling "the cloud."

So last week, when Michael Coté (an analyst with RedMonk) wrote about "How cloud computing will change capacity management" I thought it would be a good idea to expand on his observations and to dissect the issues and trends. Including my prediction that existing capacity management tool value will be overtaken by utility computing technologies.

First, terms: When I talk about the "cloud", I'm usually talking about Infrastructure-as-a-Service (a la Amazon EC2) rather than platform-as-a-service (e.g. Google app engine) or Software-as-a-Service. To me, IaaS represents that "raw" compute capacity on which we could provision any arbitrary compute service, grid, etc. (It's also what I consider the underlying architecture that's been called Utility Computing).

Michael was clear to define two other terms, Capacity Management, and Capacity Planning. Capacity management is the balancing of compute resources against demand (usually with demand data you have), while capacity planning is trying to estimate future required capacity (usually without the demand data you'd like).

Another related issue that has to be addressed is
Consolidation Planning - essentially "reverse" capacity planning -- estimating how to minimize overall in-use capacity while maintaining service and availability levels for virtualized applications.

So how does use of "cloud" (IaaS) impact capacity management/planning, as well as consolidation planning? In my estimation, there are two broad views on this:
  1. If you buy-into using the "public" cloud, then all the work you've been doing to estimate capacity and to plan for consolidation doesn't really matter. It's because your capacity has been outsourced to another provider who will bill you on an as-used basis. The IaaS "cloud" is elastic, and expands/contracts in relationship to demand.
  2. If you instead build an "internal cloud", or essentially architect a utility computing IaaS infrastructure, the story is a little different. You're taking non-infinite resources (your data center) and applying them in a more dynamic fashion. Nonetheless, the way you've been doing capacity management/planning, and even consolidation planning, will change forever.
I'll take #2, above, as an example, because its operation is more transparent. You start with your existing infrastructure (machines, network, storage) and use policy-based provisioning/controls to continuously adjust how it is applied. This approach yields a number of nice properties:
  • Efficiency: You only use the capacity (physical and/or virtual) you need, and only when you need it
  • Continuous consolidation: A corollary to above is that the policy engine can "continuously consolidate" virtualized apps (e.g. it can continually compute and re-adjust consolidated applications for "best-fit" against working resources)
  • Global view: Global available capacity (and global in-use capacity) is always known
  • Prioritization: You can apply policy to prioritize capacity use (e.g. e-commerce apps get priority during the holidays, financial apps get priority at quarter-close)
  • Safety net: You can apply policy to limit specific capacity use (e.g. you're introducing a new application, and you don't know what initial demand will be)
  • Resource use: It enables solutions for "resource contention" (borrowing from Peter to pay Paul); higher-priority applications can temporarily borrow capacity from lower-priority apps.
The net-net of the properties above is the long-term obviation of capacity planning, capacity management, and consolidation-planning tools. (Now take a deep breath)

Yes. Long-term, I would expect existing capacity management tools like PlateSpin PowerRecon, CiRBA's Data Center Intelligence, and VMware's Capacity Planner to be completely obviated with the appropriate internal IaaS architectures. Why? Well, let's say you do clever consolidation-planning for your apps. You virtualize them and cram them into many fewer servers. But a few months pass, and the business demand for a few apps changes... so you have to start re-planning over again. Contrast this against an IaaS infrastructure, where you let a computer continuously figure out the "best fit" for your applications. The current concept of "static" resource planning is destined for the history books.

Oh - and there are some nice side-benefits of allowing policy to govern when and where applications are provisioned in an internal IaaS ("internal cloud") architecture:

1) Simplified capacity additions: If capacity is allocated on a global basis, then the need to plan on a per-application basis is much less important. Raw capacity can be added to a "free pool" of servers, and the governing policy engine allocates it as-needed to individual applications. In fact, the more applications you have, the "smoother" capacity can be allocated, and the more statistical (rather than granular) capacity measurement can become.

2) Re-defined "consolidation planning": As I said above, the "static" approach to consolidation planning will give way to continuous resource allocation, essentially "continuous consolidation." Instead, you'll simply find yourself looking at used capacity (whether for physical or virtualized apps) and add raw capacity (as in #1) as-needed. The hard work of figuring out "best fit" for consolidation will take place automatically, and dynamically.

3) Re-defined capacity management: Just like #2 - Rather than using tools to determine "static" capacity needs, you'll get a global perspective on available vs. used raw capacity. You'll simply add raw capacity as-needed, and it will be allocated to physical and/or virtual workloads as-needed.

4) Re-defined
capacity planning for new apps: Instead of the "black art" of figuring-out how much capacity (and HW purchase) to allocate to new apps, you'll use policy instead. For example, you roll-out a new app, and use policy to throttle how much capacity it uses. If you under-forecast, you can "open the throttle" and allow more resources to be used -- and if it's a critical app, maybe even dynamically "borrow" resources from elsewhere until you permanently acquire new capacity.

5) Attention to application "Phase": You'll also realize that the best capital efficiency occurs when you have "out-of-phase" resource demands. For example, most demand for app servers happens during the day, while demand for backup servers happens at night -- so these out-of-phase needs could theoretically share hardware. So I would expect administrative duties to shift towards "global load balancing", and encouraging non-essential tasks to take place during off-hours. Much the same way Independent System Operators across the country share electric loads.

BTW, if all of this sounds like "vision" and vaporware, it's not. There are firms offering architectures like IaaS, Internal Clouds and utility computing technologies today, that work with your existing equipment. I know one of them pretty well :)

Thursday, October 9, 2008

Cassatt's chief scientist explains, simplifies

What's the sign of a really smart guy? The ability to take a complex topic and simplify it so that even your mom will understand it.

Steve Oberlin, Cassatt's Chief Scientist, had done that. He's taken a look at how data centers operate, the dynamics that drive them, and how existing technology can help simplify IT management's life. Simpler capacity management, service-level management, and overall energy efficiency. It's the basis behind utility computing, what will drive Infrastructure-as-a-Service (IaaS), the basis for building "internal cloud" infrastructures we're all talking about.



Oh. And what's the sign of an extraordinarily smart guy? That he can simplify-down these concepts -and- produce the entire video himself.



Wednesday, October 1, 2008

Postcards from SDForum - Cloud Computing and Beyond

I attended most of today's SDForum "Cloud computing and Beyond: The Web Grows Up (Finally)" in Santa Clara. Somewhere around 200 professionals from Silicon Valley showed to hear -- and to debate -- the relative maturity and merits of the thing we're calling the cloud.

The day was lead-off by James Staten, a friend and former colleague, and now with Forrester Research, who gave a fantastic keynote of "Is cloud computing the next revolution?" Just getting to a definition of terms, and mapping the taxonomy of this emerging market is tricky. But he's tracking this fast-maturing market rather closely. Both web-based services and Software-as-a-Service are becoming the norm; but the industry is also calling the lower-level services (PaaS, IaaS) cloud too. So be careful of terms when you enter into a cloud debate.

Another morning keynote (which I unfortunately missed most of) was delivered by Lew Tucker, Sun Microsystems' new CTO of Cloud Computing (and also a friend and former colleague). He's quite a visionary, and went so far as to suggest that computing resources of tomorrow will be brokered/arbitraged based on specializations, costs, etc.

One particularly lively panel was hosted by Chris Primesberger of E-Week, with panelists from Salesforce.com, Intacct, SAP, RingCentral and Google. There was some light discussion about cloud differentiation, interaction, and standard approaches to describing cloud SLAs. Most generally agreed that there would in fact be 3rd-party businesses brokering between providers at some point. The other enlightening discussion focused on capacity planning for the cloud -- what if a user scaled from ten to ten-thousand servers in a few days or weeks? Could services like Amazon handle this? In a consistent - and impressive - way, the panelists agreed that these sorts of scale issues were "a drop in the bucket" when you consider the vastness of what these large service provide on a daily basis.

In what drew the most spontaneous applause was a question asked to the panel (but probably directed to Rajen Sheth of Google) by a member of the audience. Essentially, how could we *not* assume there would be service lock-in, when Force.com had one platform model, and Google App Engine had another? (a good point elucidated by James Urquhart some time ago). The Google response focused on "providing the best possible service for customers" but was clearly a dodge. (BTW, the author herein suggests that SaaS and PaaS models will follow the same proprietary/fragmentary model as did Linux and Unix).

In an afternoon panel led by David Brown of AMR research, the main question addressed was whether (or to what degree) cloud computing was disruptive. The panel consisted of hardware, software and services vendors from Elastra, Egenera, Joyent and Nirvanix. The panel agreed that there were different types of disruption, depending on where you sit. From an infrastructure management perspective, internal cloud architectures can be disruptive to IT Ops, since it changes how resources are applied and shared, and the fundamentals of capacity planning. Cloud architectures can also be disruptive to traditional forms of hosting and outsourcing, due to their pay-as-you-go approach.

I will say that Jason Hoffman, Founder of Joyent, stood out in the panel clearly as a visionary in this field. Keep an eye on this guy. His take on disruption was that if "cloud" means Infrastructure-as-a-Service, then it's really just another form of hosting, and not very disruptive. But if how "clouds" are applied to support business needs using policy (i.e. to dynamically communicate SLAs, Geographic compute locations, costs, replication, failover,etc.) then they become very disruptive. IT administration would shift from scripting and fire-fighting, to policy-development and policy modification.

Finally, I will point out that many more folks showed-up who would use clouds and/or broker cloud services than who would actually *make* the clouds (IaaS) in the first place, again attesting to the point I made earlier this week that it's a lot harder to do, and only really sophisticated vendors will be taking that on.

Thursday, September 25, 2008

20 Cloud computing startups - analysis

I was pointed to John Foley's InformationWeek article earlier this week of "20 Cloud Computing Startups You Should Know." Aside from the fact I could only count 19, it was a great survey of what types of companies, ideas and ventures are getting on the bandwagon.

The quick-and-dirty chart above is mine; what I found so interesting is that 8 of the players are building solutions on top of other clouds (like Amazon's EC2 and S3) while another 7 are investing in essentially building hosted services.

However, only 4 (ok, maybe 4-1/2) are thinking/trying to bring "cloud" technologies and economics to the enterprise's own internal IT. This certainly attests to the difficulty in reworking IT's entrenched technologies, and building a newer abstracted model of how IT should operate.

Even though Cassatt wasn't mentioned in the survey (maybe we were supposed to be #20) we also play in the "build-an-internal-cloud-with-what-you-have" space.

This model -- that of an "internal cloud" architecture -- will ultimately result in more efficient data centers (these architectures are highly efficient) and ones that will be able to "reach out" for additional resources (if-and-when needed) in an easier manner than today's IT.

I'd look to see more existing enterprises considering building their own cloud architectures (after all, they've already invested lots of $$ in infrastructure) while startups and smaller shops opt for the products that leverage existing (external) cloud resources.

BTW, John also just posted a very nice blog of a "reality check" to curb some of the cloud computing hype.

Wednesday, September 17, 2008

Postcards from the Hosting Transformation Summit

Right across the street from VMworld was Tier 1's Hosting Transformation Summit. Roughly 400 folks -- mostly from Managed Service Providers (MSPs) -- attended to get the lowdown on where that industry is going. It's changing fast, given some of the recent "cloudy" offerings from Amazon, Mosso, OpSource and others. And part of the driver was the technology offered from VMware itself.

But First: The industry, and its growth, is compelling. Managed services hosting is growing in the U.S. at about 30%/y, and it will be a $10 billion industry by the end of 2008. While about 20% of that amount is represented by 13 of the largest firms, the remainder of the market is represented by hundreds if not thousands of smaller entities.

Dan Golding of Tier 1 pointed out that the categories called "web hosting" and "managed hosting" are colliding, given that so many apps are being delivered over http. He also pointed out that small/medium businesses are expectted to outsource even more of their own IT, as operating it themselves becomes more complex and expensive... also good for MSPs. In particular he noted that CRM, HR, Accounting, fileservers, utility storage, email and project management were expected to be the top managed SaaS applications.

Next, John Zanni, Microsoft's GM of worldwide hosting, gave a talk called "cloud computing - is virtualization enough?" Having seen Paul Maritz' VMworld keynote hours before, I couldn' t help but compare the two. Zanni's a really smart guy -- but vision-wise, his talk was a let-down. While he absolutely identified the same requirements of the "cloud" (which were surprisingly in-agreement with VMware's) Microsoft's vision was elementary in comparison to VMware, referencing Microsoft party-lines and products - and was weak on vision. Granted, the audience was not as heavily-laden with technologists as the VMware conference, but the vision that was sketched-out just didn't seem too fully-baked. One interesting side-note: John explicitly mentioned Microsoft management tools that would someday manage 3rd-party VMs such as VMware. Hmmm....

On day 2, Antonio Piraino (also of Tier 1) gave a really great talk on "virtualization and cloud computing" -- the guy really gets it, with respect to the MSP industry. His message to MSPs was pretty clear: Cloud is coming, and you (the MSP) will need to learn about it and get on board. The definition of "cloud" he gave to MSPs was
  • Server-based managed hosting
  • Virtualized offerings
  • Multi-O/S & DB support
  • Automated scalablity
  • Easy ordering of services
  • On-demand provisioning
  • Cross-service integration
  • Bill-for-use
  • SLAs were managed / managed-for
It's clear that the smaller MSPs out there will be jumping on the Utility Computing, Cloud, PaaS and SaaS bandwagon soon. That should begin to give folks like Mosso, Flexiscale, IronScale, OpSource etc. some competition. But i'm sure we're going to see the concept of abstracted-away hardware grow in popularity with frightening velocity.


Postcards from VMworld 2008 (with a twist)

I'm a bit late in reporting-back on day #1 of VMworld in Las Vegas. Word-on-the-floor is that there are over 14,000 attendees here. Definitely indicative of the hunger the industry has for this technology.

Rather than re-hash all of what CEO Paul Maritz had to say, I'd like to point out why VMware's vision is both on-the-mark -and- already available from sources other than VMware.... and showcase one such available product

Paul outlined 3 areas of vision:
  • Virtual Data Center O/S (VDC-OS)
  • vCloud (providing the ability to build internal/external clouds and federation between clouds)
  • vClient (providing end-client independence for services emanating from clouds
He emphasized, with a demo, how an "internal cloud" could reach-out to an off-premises (external) cloud for resources, say during peaking demand -- or perhaps as a failover scenario. The demo has 3 points to make: (a) the ability to provide "elastic" capacity, (b) the ability to provide self-healing in the form of replacing failed capacity, and (c) the fact that it was driven by policies based on SLAs. It was a demo of a non-commercially-available product, but it drew great applause from the audience.

Whenever the "big guys" show-off a concept/roadmap, you can be sure that there are already smaller guys who are paving the way for them; this is no different. Cassatt, for one, has been showing-off this type of demo (down to a similar GUI) for many months now. With a few key differences:
  • The product is shipping today
  • We don't require that there are "warm" hosts pre-provisioned as standby resources
  • We don't require that VMware is everytwere; in fact, we can already show the same demo but using Xen/Citrix (and soon, with other VM players)
  • We don't even require that Virtualization is used at all; our approach works with physical HW and O/Ss too (including x86, SPARC, Linux distros, Solaris, and others)

For those attending the keynote, perhaps the GUI above looks familiar; except it's Cassatt's Active Response 5.1

In the center is a chart indicating upper- and lower- SLA thresholds (SLAs can be arbitrarily defined and composed). If the upper SLA is breached, Active Response finds bare-metal resources in the "free pool" (again, defined how you like) and then automatically provisions those resources with whatever SW policy determined (read: either a physical server or a virtual server). The application "tier" grows automatically. If/when the lower threshold is breached, an instance on the "tier" is retired. This approach provides real-life SLA management, capacity-on-demand (elastic behavior), failover/availability, and many other nice-to-have properties -- automatically. And Today.

This set of properties were also discussed across the street today at Tier-1 Research Hosting Summit at the Mirage. Many MSPs in the audience wanted to know "how do I get some of that?" when discussion came to utility computing and cloud infrastructures. I'll post on that next :)

Monday, September 8, 2008

Inherent efficiencies of Cloud Computing, Utility Computing, and PaaS

Ever noticed that the two hottest topics in IT today are Data Center Efficiency and Cloud Computing? Ever wondered if the two might be related? I did. And it’s clear that the media, industry analysts – and most of all IT OPs – have missed this relationship entirely. I now feel obligated to point out how and why we need to make this connection as soon as we can.

Let me cut to the chase: The most theoretically-efficient IT compute infrastructure is a Utility Computing architecture – essentially the same architecture which supports PaaS or “cloud computing”. So it helps to understand why this is so, why today's efficiency “point solutions” will never result in maximum compute efficiency, and why the “Green IT” movement needs to embrace Utility Computing architectures as soon as it can.

To illustrate, I’ll remind you of one of my favorite observations: How is Amazon Web Services able to charge $0.10/CPU-Hour, (equating to ~$870/year) when the average IT department or hosting provider has a loaded server cost of somewhere between $2,000-$4,000/year? What does Amazon know that the rest of the industry doesn’t?

First: A bit of background

Data center efficiency is top-of-mind lately. As I’ve mentioned before, a recent EPA report to U.S. Congress outlined that over 1.5% of U.S. electricity is going to power data centers, and that number may well double by 2011. Plus, according to an Uptime Institute White Paper, the 3-year cost of power to operate a server will now outstrip the original purchase cost of that server. Clearly, the issues of high cost and limited capacity for power are currently hamstringing data center growth, and the industry is trying to find a way to overcome it.

Why point-solutions and traditional approaches will miss the mark

I regularly attend a number of industry organizations and forums on IT energy efficiency, and have spoken with all major industry analysts on the topic. And what strikes me as absurdly odd is that the industry (taken as a whole) is missing-the-mark on solving this energy problem. Industry bodies – mostly driven by large equipment vendors – are mainly proposing *incremental* improvements to “old” technology models. Ultimately these provide a few % improvement here, a few % there. Better power supplies. DC power distribution. Air flow blanking panels. Yawn.

These approaches are oddly similar to Detroit trying to figure out how to make its gas-guzzlers more efficient by using higher-pressure tires, better engine control chips and better spark plugs. They’ll never get to an order-of-magnitude efficiency improvement on transportation.

Plus, industry bodies are focusing on metrics (mostly a good idea) that will never get us to the major improvements we need. Rather, the current metrics are lulling us into a misplaced sense of complacency. To wit: The most oft-quoted data center efficiency metrics are the PUE (Power Use Effectiveness), and it’s reciprocal, the DCiE (Data Center Infrastructure Efficiency). These essentially say, “get as much power through your data center and to the compute equipment, with as little siphoned-off to overhead as possible.”

While PUE/DCIE are nice metrics to help drive overhead (power distribution, cooling) power use down, they don’t at all address the efficiency with which the compute equipment is applied. For example, you could have a pathetically low-level of compute utilization, but still achieve an incredibly wonderful PUE and DCIE number. Sort of akin to Detroit talking about transmission efficiency rather than actual mileage.

These metrics will continue to mislead the IT industry unless it fundamentally looks at how IT resources are applied, utilized and operated. (BTW, I am more optimistic about the Deployed HW Utilization Efficiency “DH-UE” metric put forth by the Uptime Institute in an excellent white paper, but rarely mentioned)


Where we have to begin: Focus on operational efficiency rather than equipment efficiency

So, while Detroit was focused on incremental equipment efficiency like higher tire pressure and better spark plugs to increase mileage, Toyota was looking at fundamental questions like how the car was operated. The Prius didn’t just have a more efficient engine, but it had batteries (for high peak needs), regenerative braking (to re-capture idle “cycles”), and a computer/transmission to “broker” these energy sources. This was an entirely new operational model for a vehicle.

The IT industry now needs a similar operationally-efficient re-engineering.

Yes, we still need more efficient cooling systems and power distribution. But we need to re-think how we operate and allocate resources in an entirely new way. This is the ONLY approach that will result in Amazon-level cost reductions and economies-of-scale. And I am referring to cost reductions WITHIN your own IT infrastructure. Not to outsourcing. A per-CPU cost basis under $1,000/year, including power, cooling, and administration. What IT Operations professional doesn’t desire that?

Punchline: The link between efficiency, cloud computing & utility computing architectures

What the industry has termed “utility computing” or Platform-as-a-Service (a form of “cloud” computing) provides just this ideal form of operational-efficiency and energy-efficiency to IT.

Consider the principles of Utility Computing (the architecture behind “clouds”): Only use compute power when you need it. Re-assign it when-and-where it’s required. Retire it when it’s not needed at all. Dynamically consolidate workloads. And be indifferent with respect to the make, model and type of HW and SW. Now consider the possibilities of using this architecture within your four walls.

Using the design centers above, power (and electricity cost) is *inherently* minimized because capital efficiency is continuously maximized. Always, regardless of the variation from hour-to-hour or month-to-month. And, this approach is still compatible with “overhead” improvements such as to cooling and power distribution. But it always guarantees that the working capital of the data center is continuously optimized. (Think: the Prius engine isn’t running all the time!)

On top of this approach, it would then be appropriate to re-focus on PUE/DCIE!

The industry is slowly coming around. In a recent article, my CEO, Bill Coleman, pointed out a similar observation: Bring the “cloud” inside, operate your existing equipment more efficiently, and save power and $ in the process.

I’m only waiting for the rest of the industry to acquiesce the inherent connection between Energy efficiency, operational efficiency, power efficiency, and the architectures behind cloud computing.

Only then will we see a precipitous drop in energy consumed by IT, and the economies-of-scale that technology ought to provide.

Monday, August 18, 2008

Creating a Generic (Internal) Cloud Architecture

I've been taken aback lately by the tacit assumption that cloud-like (IaaS and PaaS) services have to be provided by folks like Amazon, Terremark and others. It's as if these providers do some black magic that enterprises can't touch or replicate.

However, history's taught the IT industry that what starts in the external domain eventually makes its way into the enterprise, and vice-versa. Consider Google beginning with internet search, and later offering an enterprise search appliance. Then, there's the reverse: An application, say a CRM system, leaves the enterprise to be hosted externally as SaaS, such as SalesForce.com. But even in this case, the first example then recurs -- as SalesForce.com begins providing internal Salesforce.com appliances back to its large enterprise customers!

I am simply trying to challenge the belief that cloud-like architectures have to remain external to the enterprise. They don't. I believe it's inevitable that they will soon find their way into the enterprise, and become a revolutionary paradigm of how *internal* IT infrastructure is operated and managed.

With each IT management conversation I've had, the concept that I recently put forward is becoming clearer and more inevitable. That an "internal cloud" (call it a cloud architecture or utility computing) will penetrate enterprise datacenters.

Limitations of "external" cloud computing architectures

Already, a number of authorities have pretty clearly outlined the pros and cons of using external service providers as "cloud" providers. For reference, there is the excellent "10 reasons enterprises aren't ready to trust the cloud" by Stacey Higginbotham of GigaOM, as well as a piece by Mike Walker of MSDN regarding "Challenges of moving to the cloud”. So it stands that innovation will work around these limitations, borrowing from the positive aspects of external service providers, omitting the negatives, and offering the result to IT Ops.

Is an "internal" cloud architecture possible and repeatable?

So here is my main thesis: that there are software IT management products available today (and more to come) that will operate *existing* infrastructure in a manner identical to the operation of IaaS and PaaS. Let me say that again -- you don't have to outsource to an "external" cloud provider as long as you already own legacy infrastructure that can be re-purposed for this new architecture.

This statement -- and associated enabling software technologies -- is beginning to spell the beginning of the final commoditization of compute hardware. (BTW, I find it amazing that some vendors continue to tout that their hardware is optimized for cloud computing. That is a real oxymoron)

As time passes, cloud-computing infrastructures (ok, Utility Computing architectures if you must) coupled with the trend toward architecture standardization, will continue to push the importance of specialized HW out of the picture.
Hardware margins will continue to be squeezed. (BTW, you can read about the "cheap revolution" in Forbes, featuring our CEO Bill Coleman).

As the VINF blog also observed, regarding cloud-based architectures:
You can build your own cloud, and be choosy about what you give to others. Building your own cloud makes a lot of sense, it’s not always cheap but its the kind of thing you can scale up (or down..) with a bit of up-front investment, in this article I’ll look at some of the practical; and more infrastructure focused ways in which you can do so.

Your “cloud platform” is essentially an internal shared services system where you can actually and practically implement a “platform” team that operates and capacity plans for the cloud platform; they manage its availability and maintenance day-day and expansion/contraction.
Even back in February, Mike Nygard observed reasons and benefits for this trend:
Why should a company build its own cloud, instead of going to one of the providers?

On the positive side, an IT manager running a cloud can finally do real chargebacks to the business units that drive demand. Some do today, but on a larger-grained level... whole servers. With a private cloud, the IT manager could charge by the compute-hour, or by the megabit of bandwidth. He could charge for storage by the gigabyte, and with tiered rates for different availability/continuity guarantees. Even better, he could allow the business units to do the kind of self-service that I can do today with a credit card and The Planet. (OK, The Planet isn't a cloud provider, but I bet they're thinking about it. Plus, I like them.)
We are seeing the beginning of an inflection point in the way IT is managed, brought on by (1) the interest (though not yet adoption) of cloud architectures, (2) the increasing willingness to accept shared IT assets (thanks to VMware and others), and (3) the budding availability of software that allows “cloud-like” operation of existing infrastructure, but in a whole new way.

How might these "internal clouds" first be used?

Let's be real: there are precious few green-field opportunities where enterprises will simply decide to change their entire IT architecture and operations into this "internal cloud" -- i.e. implement a Utility Computing model out-of-the-gate. But there are some interesting starting points that are beginning to emerge:
  • Creating a single-service utility: by this mean that an entire service tier (such as a web farm, application server farm, etc.) moves to being managed in a "cloud" infrastructure, where resources ebb-and-flow as needed by user demand.
  • Power-managing servers: using utility computing IT management automation to control power states of machines that are temporarily idle, but NOT actually dynamically provisioning software onto servers. Firms are getting used to the idea of using policy-governed control to save on IT power consumption as they get comfortable with utility-computing principles. They can then selectively activate the dynamic provisioning features as they see fit.
  • Using utility computing management/automation to govern virtualized environments: it's clear that once firms virtualize/consolidate, they later realize that there are more objects to manage (virtual sprawl) , rather than fewer; plus, they've created "virtual silos", distinct from the non-virtualized infrastructure they own. Firms will migrate toward an automated management approach to virtualization where -- on the fly -- applications are virtualized, hosts are created, apps are deployed/scaled, failed hosts are automatically re-created, etc. etc. Essentially a services cloud.
It is inevitable that the simplicity, economics, and scalability of externally-provided "clouds" will make their way into the enterprise. The question isn't if, but when.

Monday, July 21, 2008

Postcards from the SF Datacenter Dynamics meeting

This was an awesome (and pretty intense) 1-day show in San Francisco this past Friday, covering all of the current IT operations and energy efficiency topics. It was one of a number of local/international shows run by the same Brits who also publish ZeroDowntime. And, it was one of their largest - they claim it drew ~ 800 folks, almost all of whom were directly involved in operating end-user datacenters. I definitely recommend attending one in your area.

I had great conversations with some of the authorities, and attended a handful of sessions that included the US DOE, a panel on datacenter econometrics, an end-user panel regarding datacenter automation, and some vendor presentations regarding upcoming technologies.

US DOE
This was definitely the most newsworthy session (see my previous blog entry). The DOE has been piloting their DataCenter assessment tool, "DC-Pro" lately - and their primary assistant, Lawrence Berkeley National Laboratories (LBNL), gave a walk-through of the tool, plus a roadmap of the overall goals and roll-out plans from now through 2011. I've now spoken with Bill Tschudi of LBNL a number of times; he's optimistic that the DOE will hit its goals for the tool, and for thousands of data center operators to make use of it in the next year or so.

IT Econometrics panel
Early in the morning there was also a decent panel covering topics of "green data center econometrics", essentially diving into a number of cost topics frequently overlooked in analyses. On the panel was Jon Haas of Intel, Mark Honeck of Quimonda (big DRAM manufacturer), and Winston Bumpus representing the Green Grid. Everyone agreed to "measure first", the same mantra that came out of the Uptime Institute earlier this year... in other words, measure power, temperatures, airflows and economics first, so to establish a baseline and a quantifiable goal for improvement. The other conclusion i'm happy they reached was to pursue projects that can get done quickly and show real benefits - pursue tactical initatives first.

Lastly (and by virtue of who was on the panel) came an interesting conclusion having to do with server-based power consumption: Memory is a *huge* power hog, made even worse by the move toward virtualization which typically requires a large memory upgrade for consolidating servers. One finding was that not all memory is created equal, and not all configurations consume equal power (16 1Gb SIMMs can consume four times as much power as 2 8Gb SIMMs)

On Datacenter Automation
This was a fantastic session given jointly by Cisco and OSIsoft. Cisco's primary speaker was in charge of their global lab compute capacity, and is trying to consolidate something like 200 separate labs around the globe. He clearly understood the organization differences between Facilities & IT operations, and the need to fill the gap - otherwise no useful efficiencies could be realized. Further, he predicted (if not asked to require) that IT automation systems (that govern compute, power and cooling resources) ultimately be integrated with building automation systems. In fact, he went a step further and posited that he'd like to see automation systems that interact globally. That means, he'd like to be able to dynamically push (compute) load to locations where capacity was economical -- a "follow-the-moon" strategy. This is counter to the traditional example whereby one pushes cooling to where the hot spots are; rather, push compute loads to where the cooling (and floorspace) is. This form of automation is right up my alley :) I'm happy to see other industry leaders as proponents.

I also had a chance to speak at length with Paul Marcoux, Cisco's VP of green engineering. (An interesting proposal from him here). He very much believes that the US will face carbon emissions capping/trading in the next few years... after it's incepted by the EC and others. Ergo, Cisco is taking the lead in comprehensive sustainability initiatives. And if you look at the number of sustainability organizations they're taking the lead in, you have to believe it.

Exhibitors
Aside from the sponsor/exhibitors, there were very few vendors at the show, and lots of time to interact & network with local peers -- something that's invaluable, and that I heard that time and again from attendees who've been in the past. What was also great what that most (but not all) of their pitches were truly education, with a minority being "commercials" for product.

Summary
This is a great show for data center managers to attend; it's only one day out of your schedule, and because it visits 7 US cities, minimal travel is usually involved. They've also got an international perspective because they visit 20 other cities around the globe.

Wednesday, July 16, 2008

Is an "internal" cloud an oxymoron?

By definition, "the cloud" lies external to the enterprise data center. And it's got great properties: in an Infrastructure-as-a-Service example, per-CPU operating costs are on the order of $800/year (see Amazon's EC2 price list), whereas CPU operating costs are typically $3k-$5k/year in the average data center.

It seems to me that the industry has become overly-fixated on hosted clouds (IaaS, PaaS, SaaS etc.) that are run by third parties which have all of those nice economies-of-scale.

But what about implementing an "Internal" cloud inside of corporate data centers? John Foley of InformationWeek just raised this question in talking about Elastra.
[they are] working on a version of Cloud Server for data center VMware environments, or what it refers to as "private clouds." That's an oxymoron since cloud computing, by definition, happens outside of the corporate data center, but it's the technology that's important here, not the semantics."
Semantics aside, what properties would an ideal "internal" cloud have? Clearly the same economics as a "traditional" cloud, but with some added benefits to avoid the current pitfalls of external clouds. The improved properties include --
  • Should work with existing physical & virtual resources in the data center (heterogeneous platforms & O/S's)
  • Should let you specify whether your apps are virtualized or not (but either way, provide capacity-on-demand)
  • Wouldn't require that sensitive data be hosted outside the enterprise; it would maintain internal auditability
  • Ought to adhere to internal security and configuration management processes
  • Would not disrupt existing software architectures
  • Would allow you to add additional capacity (compute resources) on-the-fly
  • Could be segregated to support both production & development environments
  • Would provide internal metering & billing for internal users and business units
Maybe an "internal cloud" needs a new name, but what it represents is essentially the basis of Utility Computing. Check out the Cloudy Times blog, that also references the potential of an Internal Cloud.

I'm guessing that, as cloud computing gains steam, IT organizations will want the same properties internally - through implementing an "internal" cloud leveraging utility computing infrastructure.

ProductionScale recently ruminated on this topic (calling it a private cloud, instead of internal):
"What is private cloud computing? To make a non-technical analogy, Private Cloud Computing is a little like owning your own car instead of using a rental car that you share with others others and that someone else owns for your automobile and transportation needs. Rental cars haven't completely replaced personal automobile ownership for many obvious reasons. Public Cloud Services will not likely replace dedicated private servers either and will likely drive adoption of private cloud computing".
Working for Cassatt, I'm biased toward believing that a market for Internal Cloud infrastructure providers will emerge.... and potentially help enterprises dovetail their internal clouds with public clouds. Any other opinions?

Friday, June 27, 2008

Sanity check: Data center energy summit

As I mentioned in yesterday's entry, the Silicon Valley Leadership Group's (SVLG) energy summit was fantastic - the first time I've seen actual implementations and data from innovations to help with energy efficiency. But on further reflection, there was something missing and rather alarming.

BTW, all of the the data was correct, and the conclusions were dead-on - the findings/predictions from the 2007 EPA report on data center efficiency were validated.

However, check out the list of projects in the Accenture report: 9 of them focused on site infrastructure, while only 3 of them focused on IT equipment.

Why weren't more projects aimed at making the IT equipment itself more efficient?

Now, if you look at where power is used in a data center, you'll find that with a "good" PUE, 30% might go toward infrastructure (with 70% getting to IT equipment), and that with a "bad" PUE, maybe 60% goes toward infrastructure (with the remaining 40% getting to IT equipment). In either case, the IT equipment is chewing-up a great deal of the total energy consumed... and yet only 1/4 of the projects undertaken had to do with curbing that energy.

This is like saying that a car engine is the chief energy-consuming component in a car, but that to increase gas mileage, scientists are focusing on drive trains and tire pressure.

My take is that the industry is addressing the things it knows and feels comfortable with: wires, pipes, ducts, water, freon, etc. Indeed, these are the "low-hanging fruit" of opportunities to reduce data center power. But why aren't IT equipment vendors addressing the other side of the problem: Compute equipment and how it's operated?

IT equipment is operated as an "always-on" and statically-allocated resource. Rather, it needs to be viewed as a dynamically-allocated, only-on-when-needed resource. More of a "utility" style resource. This approach will ultimately (and continuously) minimize capital resources, operational resources and (by association) power, while always optimizing efficiency. It is where the industry is going -- what's termed as cloud computing. This observation cuts directly to Bill Coleman's keynote (video here) earlier this week at O'Reilly's Velocity conference. It also alludes to Subodh Bapat's keynote where he outlined a continuously-optimized IT, energy, facilities and power grid system.

I certainly hope that at the SVLG's data center energy summit '09 next year, more projects focus on how IT equipment is operated, rather than on the "plumbing" that surrounds it. I can't wait to see the efficiency numbers that emanate from an "IT Cloud" resource.

Monday, June 23, 2008

Postcards from Gartner's IT Infrastructure, Operations & Management Summit 2008

Here I am in Orlando at the Gartner IT conference. The day's been insightful and validating, if you happen to be in the Real-Time Infrastructure (RTI) business.

Andy Kyte
The opening keynote was from Andy Kyte, Gartner VP and Fellow. He's a dynamic speaker, and focussed mostly on IT Modernization and strategy. He was careful to define strategy/strategic planning -- and accused just about every IT management organization of buying "puppies". That is, most orgs buy products because cute "here-and-now" reasons, without realizing that they're really signing-up to a 15-year-long relationship with a dirty, hairy, high-maintenance and expensive pet. His point was validated when he pointed-out all of the point-products that organizations have purchased that essentially only add to cost & complexity, rather than reduce it. He posited that more products have to be purchased with a long-term (7+ years) strategic vision, and that short-term economic validation was often to blame for the morass that IT finds itself within.

Donna Scott
Next was Donna Scott, speaking about IT Ops management trends -- and later on in the day, speaking about IT modernization and RTI. She led-off with a list of projects that enable business growth ... and that projects that don't enable business growth should be canceled.

But most interesting was her coverage of the "cloud" which she (and Thomas Bittman, next) predicted would be where IT is evolving. She suggested that IT ops will evolve into an "insourced hosting" model - where IT departments will be building "internal cloud-computing" style infrastructures to support business owners. We here at Cassatt salute you, since that's what we enable :)

What was also cool about Donna's presentations were her many polls from the audience (probably 1,000 plus). Her first question was "what grade would you give leading IT management providers" 70% of them (CA, BMC, HP and IBM) got a "C" or worse. Her conclusion was that they still don' t manage complexity (they may monitor it, though), they still support the point-solution mentality, and most focus on single homogeneous platforms.

Finally, and most validating, Donna listed the chief properties/components of a Real-time infrastructure system... which she feels is practically on the market. Her list:
  • IT services provisioning
  • IT services automation (starting & stopping applications as-needed)
  • Process automation & change management
  • Dynamic Virtualization management
  • Services optimization
  • Performance management, capacity management.
Personally, it sounds like a pretty familiar list. She outlined what RTI could enable; the list was also pretty familiar:
  • Service virtualization management
  • J2EE management
  • Oracle RAC management
  • Disaster Recovery - sharing & re-configuring assets
  • Managing a shared test environment
  • "loosely-coupled" HA - replacing failed nodes
  • Dynamic Repurposing nodes
  • Dynamic capacity on demand / capacity expansion
Thomas Bittman
Thomas gave a great talk on "Virtualization changes virtually everything"... and essentially outlined the path the industry will likely take towards cloud computing. He essentially pointed out where "automation" is going wrong today... that "automation" tools are focusing on components, rather than on service levels. Until that happens, IT will continue down its complexity path.

Then he hit on a concept that will IMHO be the next big thing: The Meta-O/S. Think of it as an O/S for the data center -- the O/S that enables RTI. For example, what if you started with VirtualCenter, made it work with any VM technology (Xen, MSFT, etc.), made it manage physical/native resources as well, and finally abstracted away the rest of your physical infrastructure? Then, what if it could be told to optimize resources for application service levels , and/or to minimize power or capital or some combination at all times?

We're probably closer to this vision than you think - and the more industry is comfortable with sharing resources, and more dissatisfied with vendor point-solutions, the more it will be accepting of this meta-O/S concept.

I sometimes use this analogy:
What if you walked into a data center and were told to manage it - - 10,000 servers, 100 different HW models, 5,000 applications, various O/S flavors and revs, multiple networks, etc. etc. Well, you *wouldn't* tell me that you'd hire 200 sysadmins, buy multitple software management tools and analysis packages, set up complicated CMDBs and change-management boards, and buy a bunch of pagers for after-hours fire-drills. But that's how it's done today.

Rather (given a clean slate) you'd say "I'd get a computer to figure-out how and when to run applications, and to govern what software was paired with what hardware when. It would prioritize resources, and continually optimize overall operating costs. That's the rational approach. and that's what the Meta-O/S will do.

Monday, May 19, 2008

A big step forward for self-managing data centers

Today there is a modest but hugely-meaningful announcement regarding data center operational efficiency coming out of Cassatt here in San Jose. It's about Cassatt Active Response 5.1 software, and it includes the ability to take actions on infrastructure based on application demand. That means that server power control, or even entire server repurposing and network provisioning, can be triggered by the demand (or service level) of a given application.

This means that if you run a development lab, idled machines might be automatically power-down after a period of time (say, after a test run) to save power. It means that a server farm can automatically trigger provisioning additional servers if service levels drop below a pre-determined threshold - saving time. It means as different application demands ebb and flow, the data center will adapt to demand by re-purposing (or retiring) bare-metal hardware -- making the best use of capital. And all of this whether-or-not virtualization is present, regardless of the underlying platform, and without adding software layers.

You'll find this concept filed under Gartner's Real-Time Infrastructure (RTI) category, under Forrester's Organic IT concept, or sometimes under Utility Computing. You've even seen big vendors predicting it as a vision. But I'm happy to point out that we're doing some of it already today. Think of this as another step toward a greener data center, because it really optimizes all forms of operational and capital efficiency...

The simplest application of this Demand-Based Policy Management is with Active Power Management of data center servers to curb rampant power waste. (Check out Who's Recommending Power Management) The concept has been long-used on desktops (check out 1E or Verdiem; there are over a half-million desktops under power management today). In IT environments such as dev/test, we've seen opportunities to cut gross power consumption by 30% or more in a few months. All this by simply monitoring the server activity and then gracefully shutting-down idled hardware
(where idleness can be defined any way data center managers prefer). Check out the endorsements of power management from the EPA's Andrew Fanara as well as from Jon Koomey in the Cassatt press release. You can also watch a webcast about Active Power Management

A more sophisticated use of demand-based policy is to automatically maintain the service-level of applications. Take the server farm example... or for that matter, a SOA service. In either example, demand on the service may be cyclical or unpredicatable. Instead of massively over-provisioning hardware, one could use a service level metric (again, of your determination) to provide control. If service level drops (say, due to increased demand, or perhaps because of equipment failure) the Cassatt system will simply power-up or re-purpose another piece of hardware and create a new server to increase compute capacity. And you don' t need a virtualization platform to do this. (Or you could. Previously we announced compatibility with VMware ESX as well as with Xen.) Check out the 10 min webcast on demand-based policy management.

The benefits here are massive: Besides the power saved when not using a piece of equipment, you're maximizing the use of capital because it's dynamically repurposed. Similar control is used for Cloud Computing infrastructure like Amazon Web Services (AWS) and they've achieved a compute price-point unheard-of in the industry: $0.10 per CPU per hour. Try that with your existing infrastructure.

My, my. I've overlooked the remainder of the Cassatt announcement:
  • a new interface to interact with external systems - either management systems, or equipment like Load Balancers. Take the example of using an F5 load balancer in an environment with dynamic repurposing. As new servers are provisioned, Active Response can communicate in real-time to the balancer and provide the new VIPs in seconds as the servers are brought online.
  • a new set of platform compatibilities, including power distribution units, used to remotely power-manage servers. This is in addition to a massive list of supported hardware, OSs, applications and VM technologies. This solution will work with what you have today :)
If you want more info, go to the Cassatt website, or tune into some on-demand webcast presentations of the various Policy-Based controls for IT infrastructure.