Thursday, December 20, 2007

A day at DOE's Save Energy Now workshop

Earlier this week I had the privilege of attending a workshop sponsored by the Department of Energy's Save Energy Now program , in collaboration with the EPA/EnergyStar group. The purpose was to help DOE define and engineer tools to help assess data center efficiency performance. They're reaching-out to industry to ensure they're using the right metrics, asking the right questions, and producing a product that will be useful in real-life. The tool will (presumably) help firms "score" the efficiency of their data center from 1-100, much the same way there are already tools to score the energy-efficiency of buildings, etc. BTW, see my earlier blog on some proposed metrics from The Green Grid and Uptime Institute.

In the room were mostly vendors such as AMD, Emerson Electric, HP, Intel, NetApp and Sun. Plus, folks from DOE, EPA, Lawrence Berkeley National Labs, the Green Grid, and the Uptime Institute attended as supporting resources to DOE. And first-off, it was great to see the overlap/interaction between these groups -- especially that DOE, Green Grid and Uptime Institute were cooperating.

The tool itself is expected to have 4 components of assessment: IT (servers, storage, networking, softwar), Cooling (chillers, AC units, fans), Power Systems (UPS, distribution) and Energy Sources (generation, etc.).

But the real core of the day's conversation is around what's meant by "Efficiency" - all agreed that at the highest level it was the ratio of "useful computing output" to the "energy supplied to the overall data center". That second number is sort of the easy part: it includes all of the power used for lighting, cooling, power distribution, etc. etc.; sometimes it's hard to measure empirically, but it all comes down to kWhs. The real issue, it turns out, is what's meant by the first number, "useful computing output".

In our afternoon IT breakout group, about 10 of us debated for a good hour or so about just that: how do we define "output"? Is it MIPS? Is it measured in SPEC units? Is it CPU utilization? And what about storage and network "output" as well? In the end, we agreed that we should define it the way any IT professional would: As an application service with an associated SLA. So for example, the useful output would be a service called "Microsoft Exchange" with SLA traits of X users, Y uptime, and Z latency. And most important, this approach makes the output independent of underlying technology and implementation. Thus, two data centers could implement this service/SLA combination in vastly different ways, with vastly different power requirements and therefore energy efficiencies.

When the DOE tool (or set of tools) is complete in mid-2008, it will represent a seminal event: Data Center operators will have a standard benchmark against which to self-rate -- and to (confidentially) compare their efficiencies against their peers. It will also begin to put pressure on IT vendors (!) to focus on the "gestalt" efficiency of their systems, not just of their components. (IMHO, this will make me *very* happy)

And, I hope, this benchmark will begin to accelerate (a) the move toward dynamically-managed shared compute resources, and (b) the technical & organizational bridge between IT system management and facilities management.

Wednesday, December 19, 2007

End-of-the-year holiday shopping for IT

The end of the fiscal year us upon most of us - and for some that means spending any remaining dollars on some quick-hit (read: fast ROI) initiatives.

So in the spirit of the holidays, here are some interesting links:

1.
Bridget Botelho at SearchDataCenter writes "Servers get no rest during the holidays" -- that most enterprises (and especially small/medium businesses) leave all servers on during the holidays, even if the company is closed. It's a waste of power/money, even though many solutions abound. (And did you even think about the security risk of keeping your IT on but essentially unsupervised?)

2. And, for those of you with money left in your budget, Rick Vanover from TechRepublic chimed-in to suggest 10 good ways to use your remaining IT budget before the end of the year - I particularly like the following:
  • #3: Purchase power management: Many new power management devices are available now that can be a good replacement for your limited power distribution units (PDUs). These PDUs can add management layers to individual power sockets for power consumption, naming, grouping, and power control. The new devices can also add more ports should you need to power more computer systems in your racks.
I would only add that in addition to switched PDUs, you should consider purchasing policy-based software to control them.

3. And, Rick went on to also cite an earlier blog of his,
10 things you should know about advance power management - another topic near-and-dear to my heart, especially:
  • #7: Turn off retired or unused devices: This will reduce your power consumption — and possibly accelerate your removal of the device so as not to overprovision power unnecessarily...
The net-net is the following: If you need a quick-hit, easy-to-implement (and money-saving) and pragmatic solution, pursue intelligent operation of your IT - no need for a technology refresh, no need to re-architect anything.


Monday, December 10, 2007

Server Power Management Myths - and more

An old friend, James Governor, in his GreenMonk blog recently got me thinking. First he pointed out that it's common in Japan to turn servers off at night. Then it wasn't so much as his follow-on blog (about turning servers off when you don't need them) as acomment he highlighted from Mike Gunderloy:
  • Has anyone looked at the labor costs of this? I know that even on my tiny little dozen-machine network, I am reluctant to power everything off at night simply because it takes so bloody long waiting for the damn things to boot up in the morning. Seems like actual working fast-boot technologies would go a long way to sell this initiative.
This is exactly the sort of objection or "urban myth" that we're trying to dispel. For example, many believe that application availabiltiy might be compromised if servers are shut down. However, there is a solution to this: policy-based control, whereby servers might be powered-up in advance of their need. That's the sort of work we're doing at Cassatt. And even in a small business, if servers in a closet are turned off nights and weekends, you're still talking about energy savings on the order of 40% or more over the duration of a year!

BTW, if you're interested in additional "Urban Myths" about server power control, check out the "Myths and Realities of Power Management" page.

And while you're at it, give us some feedback on how you feel about IT Energy Management and "Green IT": we're hosting a 5-minute survey this week (and, you could win a Wii if you take it).



Saturday, December 8, 2007

The Case for Energy-Proportional Computing

senior vice president of operations at Google and a Google FellowEnergy-proportional designs would enable large energy savings in servers, potentially doubling their efficiency in real-life use. Achieving energy proportionality will require significant improvements in the energy usage profile of every system component, particularly the memory and disk subsystems."

The two are making the case not only for server power management, but are calling on vendors to go a step further, to make computers adapt their consumptive ranges directly to the compute load consumed. This would be highly complementary to consolidation efforts currently underway.

In conclusion, the paper says,

  • Servers and desktop computers benefit from much of the energy-efficiency research and development that was initially driven by mobile devices' needs. However, unlike mobile devices, which idle for long periods, servers spend most of their time at moderate utilizations of 10 to 50 percent and exhibit poor efficiency at these levels. Energy-proportional computers would enable large additional energy savings, potentially doubling the efficiency of a typical server. Some CPUs already exhibit reasonably energy-proportional profiles, but most other server components do not.
  • We need significant improvements in memory and disk subsystems, as these components are responsible for an increasing fraction of the system energy usage. Developers should make better energy proportionality a primary design objective for future components and systems. To this end, we urge energy-efficiency benchmark developers to report measurements at nonpeak activity levels for a more complete characterization of a system's energy behavior
Among a few scholarly pieces from Google, the report also cites two great references; one, the US EPA's Report to Congress on Data Center Efficiency, and the other is one of many fine works by Jonathan Koomey, "Estimating Total Power Consumption by Servers in the U.S. and the World"

Tuesday, November 27, 2007

Postcards from the 2007 Gartner Data Center Conference

I attended Day #1 of the Gartner Data Center conference here in Las Vegas today - after making the strategic error of being dropped-off at the MGM Grand lobby, and having to walk what must have been 3/4 mile to the conference center...

Thomas Bittman opened the AM with a keynote on the Future of Data Center Operations. It had a pretty broad coverage of the state of DC Ops today. He had at least one memorable interjection -- What seemed as a warning to equipment vendors who have strangle-holds over customers... strongly urging customers to reject platform-specific IT technologies. He also predicted the emergence of the "meta-O/S" and the "cloud-O/S" which (I think) is a re-packaging of Gartner's Real Time Infrastructure (RTI) story. And that the meta-O/S had to be platform adn vendor-neutral. But this was the first time that I've heard Gartner pay specific attention to the emergence & legitimac of cloud computing (and the "O/S" to run it).

Next, Donna Scott gave an equally broad-ranging talk on IT Operations Management. Again, she conducted her now 5+ year-old survey of IT's biggest pressures. And once again "high rate of change", "cost containment" and "maintaining availability" took top-honors as the largest ulcer-producing pressures facing CIOs. Also true-to-form, she re-iterated that a shared infrastructure (RTI) is inevitable, breaking down the islands of technology in large data centers.


There were also some interesting vendor break-out sessions; take for example, a session on managing power and cooling from Emerson Network Power by Greg Ratcliff. The trend here is also toward an intelligent monitoring and infrastructure. He spoke of localized cooling (even within the rack) needed as rack power density increases. There was definitely reference to "adaptive cooling" and "adaptive power" -- again implying that efficiencies in large data centers can only be achieved through better use of technology, rather than throwing raw horsepower at the heat/power problem.


Finally, one last surprising (to me) datapoint: the general audience was asked who was using virtualization in production - and 1/2 to 2/3 of the audience raised their hands. This definitely drove-home the point that VMs are (and will be) everywhere. However, I combine this observation with the earlier point that data centers will need a management layer, an "O/S", which is vendor-neutral. At the moment, I don't see any of the existing large vendors stepping up to fill this virtualization managment need any time soon.









Thursday, November 15, 2007

Assessing the New Data Center Metrics

I've been reading-up on work that the Green Grid, Uptime Institute and others have been doing to define metrics around data center efficiency. The work is good but in my mind, misses the mark slightly. All of the metrics I've seen thus far are static - That is, they assume some steady-state aspect of the data center... steady compute loads, steady quantity of servers, etc. But that's not how the world works.

Even Detroit knows that autos get different efficiencies based on how & where they're driven... so the metric called "mileage" is actually measured & documented twice -- one for City, one for Highway. Data Centers need something akin to this as well.

Why? Because IT departments operate at greatly different levels; peak (maybe during the day) as well as off-peak (perhaps nights/weekends). Ideally, the data center should know how to adapt to these conditions: re-purposing "live" machines during peak hours; retiring and temporarily shutting-down idle servers during off-peak; removing power conditioning equipment when not needed; turning off specific CRAC units and chillers when not required (i.e. cold days and/or off-peak hours). We need an efficiency metric that indicates how data centers operate Dynamically.

Anyway, here's a quick survey course in what metrics I did find, and what I'd like to see:

The Green Grid on metrics:
  1. Data Center Infrastructure Efficiency,
    DCiE = (IT equipment power)/(total facility power).

    This is supposed to be a quick ratio showing how much power gets to servers, versus how much else is consumed by power distribution, cooling, lighting, etc. Driving this ratio up means you have less overhead wasting Watts. This wouldn't be too bad a metric if it was used and monitored 24x7, i.e. peak and off-peak.
  2. Power Usage Effectiveness,
    PUE = 1/DCiE (just the boring reciprocal)
  3. Data Center Productivity, (a metric to be adopted in the future)
    DCP = (useful computing work)/(total facility power)

    In theory, this is a great metric: It's like saying "how many MIPS per Watt" can you produce? (BTW, the human brain, the most powerful of all computers, consumes somethling like 25W). Anyway, DCP is a contentious metric... because each computing vendor wants to define "useful computing work" with their own (preferential) way of computing. Frankly, this is most useful to measure efficiency at the server level.
The Uptime Institute

In an excellent paper, the Uptime Institute discusses these in "Four Metrics Define Data Center "Greenness":
  1. Site Infrastructure Energy Efficiency Ratio:
    SI-EER which the Institute is currently working to re-cast in more intuitive and technically accurate terms. I suspect this is much like the Green Grid's DCiE, above
  2. Site Infrastructure Power Overhead Multiplier Which is essentially the same metric as the Green Grid's PUE, above
    SI-POM = (data center power consumption at the meter)/(total power consumption at the plug for IT equipment)
  3. Deployed Hardware Utilization Ratio:
    DH-UR = (qty of servers running live applications)/(total number of servers actually deployed)

    This speaks to the real-time utilization of hardware, and IMHO is one of the best metrics for a dynamic data center. It points to how many deployed servers are actually doing work, vs. those that are sitting "comatose". A very promising metric if it's used in conjunction with equipment that constantly optimizes how many servers are "on", and shuts down idled servers, constantly minimizing this metric.
  4. Deployed Hardware Utilization Efficiency
    DH-UE = (minimum qty of servers needed to handle peak load)/(total number of servers deployed)

    This is another great metric - it speaks to the capital efficiency of hardware - how many need to be provisioned and on the floor, relative to how many are being used actively.
In my ideal world, I'd like to see two things to get us to a "City/Highway" style approach:
  • A DH-UR that changes dynamically, constantly being minimized. This implies that only required servers are actually powered-up and active.
  • An SI-POM that was always driven toward a constant ratio, regardless of compute demand. Which implies that, as compute demand falls, servers are retired and other support equipment (power handling, cooling) also shuts down, keeping the efficiency ratio balanced.
I look forward to conversations with the Green Grid, Uptime Institute and EPA to consider these tweaks to their already fine work.




Wednesday, November 7, 2007

CIO Dialogue - Notes from the Real World

I felt the need to share the following conversation I recently had with the VP of Enterprise Operations for a major healthcare provider. On one hand, the conversation sounded like every stereotype I've heard in trade rags... except it's true. So read this, but be sure to get to the punch line at the end.

He's been in his job for 18 months, and is just now seeming to get his hand around turning the battleship. Which, I might add, "owns one of every imaginable platform and software type" and has perhaps 3,000+ apps on 12,000-15,000 servers, maybe 30%-40% is development. He's got lots of AIX and lots of Sun, but ultimately a mix of other vendors too.

When asked exactly what he owns, he says he doesn't really know... but they're planning a CMDB proje
ct soon. Also, they're quickly running out of data center space, and are pushing 95% of maximum UPS power in most locations. He's thrown-down the gauntlet and halted all new server purchases -- in favor of initiating a virtualization project (which, I might add, is getting upwards of 20:1 consolidation, although he knows that high ratio won't last). He's a risk-taker because he has to be.

So I asked him point-blank, what does he need to make this work. Without a flinch (or a smile) he said "Process and Automation." Process, a la ITIL, and automation -- both of the Run-Book style, as well as the operational style. " If I could have the automation vision that IBM was hawking a few years ago, I'd be thrilled. But it's still vapor".

The good news is that he's closely teamed with his Facilities manager to help him cope with power, real estate and cooling. The bad news is that the Facilities guy is also at wit's-end.

The punchline: This real-life vignette tells me that the traditional IT model is really broken. How come IT -- with all of its computers -- is actually the least automated and efficient arm of the company? I recently read a report from the Uptime Institute which talked about the Economic Meltdown of Moore's Law -- literally, for every $1 of compute asset, it currently costs $1.80 operate it; by 2009, electricity alone will cost triple ($3) what the box cost. What's wrong with this picture?

I know that my VP friend is not alone. But when will the treadmill of IT-being-slave-to-the-hardware end? I'd like to think that automation, active asset management, and the drive toward greater environmental efficiency will begin to influence vendors and managers alike.