Showing posts with label Big Data. Show all posts
Showing posts with label Big Data. Show all posts

Monday, May 20, 2013

What Clouds Will Form Around Data's Gravity?

The concept of Data Gravity posits that as data accumulates (whether it be stored, analyzed, used) it tends to attract even more similar data. And as the data amasses there is less likelihood that it will be moved/migrated elsewhere.  If you're not already familiar with the concept definitely check out Dave McCrory's excellent blogs, analyses and presentations on datagravity.org

I believe this data aggregation concept can also apply to attracting computing too. As I've mentioned in It will be a data-centric (cloudy) world, there are examples today where special-purpose compute clouds are already forming around special-use data sets... sometimes intentionally, sometimes organically. One example I frequently point out is the NYSE Capital Markets Community Platform - a special-purpose cloud computing environment formed with a massive market trading data set at its core.

I am increasingly asked by service providers and enterprises alike, what other businesses and special-purpose clouds might form around data?  What new clouds (and associated business models) might we build and monetize? How can we better serve vertical market needs in the Cloud?

Forming Community Clouds - Applied (vs. Theoretical) Data Gravitational Theory

After more thinking and conversations with experts on the topic, I wanted to offer some examples and ideas that I hope trigger further exploration by cloud- and service providers. Perhaps there are (or will be) new businesses based on some of these ideas of attracting data and computing.

Financial Services Community Cloud: as I've mentioned, the NYSE CMCP has at its core a huge database of stock market history.  It's natural attractor for trading firms and hedge funds to co-locate their compute loads near this data as they test and refine trading algorithms and prediction methods. High-performance processors with low-latency connections to Wall Street don't hurt the model either.  Perhaps there are other forms of gravitational financial data (other markets?) that could attract similar compute clouds?

Photography/Imagery Community Cloud: More and more companies (Shutterfly, SmugMug, EverPic stock photography companies etc.) are in the business of warehousing photos - mostly for simple monetization. But some innovative photo data collections might take advantage of this and provide a co-located compute platform for ISVs to provide higher-level photo identification, cataloging, enhancement and even geo-tagging services.  Perhaps the compute services could take advantage of knowledge about the larger database of images that have been previously tagged or otherwise cataloged within the larger community.  [Bonus thought experiment: create a shared medical imagery cloud].

CRM and Customer Insight Community Cloud: Consider the amount of customer data located on Salesforce.com and others. Now consider the amount of consumer behavior information collected by systems like Marketto and others.  What if one of these giants begins to acquire additional firms who house complementary marketing data - and begins to build valuable "big data" around customer behavior?  Much like force.com, the customer data would attract even more marketing and consumer behavior application workloads, again attracting more data and workloads.

The Retail Community Cloud: Start watching what Walmart Labs and Nielson are doing in the Big Data and retail analytics space.  It would be but a small jump for either to create a retail cloud - centered around a huge (but perhaps anonymized) database of consumer purchasing patterns, geographies, pricing and outlets. Monetize it by allowing co-location of marketing analytics workloads from marketing firms seeking insights into better forms of micro-marketing, associative/recommendation sales, and other forms of retail analytics engines. All retail firms great-and-small would want a piece of that action.

The Energy Community Cloud: What would it be worth to amass data about energy consumption -- at the customer level -- across the country? Perhaps associate those users with industry/SIC codes, zip codes, electricity prices and/or electricity source renewabilty (or carbon footprint)?  No single utility has this data, but firms such as Enernoc monitor consumption data across the country. What if they developed a cloud that encouraged co-location of workloads and businesses which take advantage of this data - such as monitoring which businesses are really "greenest", which vertical industries are growing fastest, or where alternative energy sources would be most attractive. Add to the database information such as energy efficiency programs or overlay it with data about alternative (wind, solar, geo) energy generation. The data at the core could attract compute workloads for use by other energy, efficiency, and economic monitoring businesses.

And more clouds: As I've mentioned before,
I could see this transforming both the cloud service provider ecosystem, as well as entire industry groups. Consider new Cloud Service Provider models:  What if NOAA formed the Weather and Atmospherics Community Platform? If healthcare companies created federated Medical Records Community Platforms? If the USGS formed the World Geologic Community Platform? If other brokerages created equivalent capital markets platforms? 

Building a Community Cloud with Gravity
The next natural question I wonder is how one might go about building a community cloud or "special purpose" data repository and associated compute cloud - be it around a vertical industry or specialized data type. In my opinion there are a few necessary properties each cloud (business) would have:
  • Data sets that become more valuable as they grow and become more diverse - and of course which generate additional gravity of their own
  • Business models that monetize the data - and perhaps generate additional derivative data. (In some instances the data may need to be anonymized).
  • Co-located workloads that need to be co-located near the large (gravitational) data sets due to their frequent access 
  • Privacy, security and regulatory controls specific to the industry and/or data type and globally provided/reinforced
  • Industry-specific sales & marketing - presumably each community cloud would have appeal to specific verticals, markets or industry groups. Driving demand / awareness within these markets is of course critical.
If you know of community clouds based on data gravity, please share. In my opinion we'll see dozens of these special-purpose clouds form around data sets in the coming years.


For Further Reading:

Wednesday, August 1, 2012

Enablers for Transformed IT - Placing My Bets

It's an odd time of the year to be making predictions. But recent conversations with start-ups, CEOs, CIOs and others have suggested areas in Enterprise IT that I  bet will be "hot" over the near/medium term.

Some of the areas are "Sexy" (high on the Hype Cycle) - and others are not. But in my opinion all are worth betting on. They are all interrelated and all are critical enablers to the goal of a transformed IT ecosystem.

Already Sexy: Integrated Cloud Infrastructure Management
Of course you expected this one... But the betting opportunity isn't exactly what you think.

There are currently hoards of point-products professing "cloud management" - including the OpenStack/Cloudstack alternatives - but surprisingly, these still lack in providing a comprehensive solution that a reasonably sophisticated SP or CIO needs to buy.

In other words, I don't just mean a product that offers an automated virtualization (server, storage, networking) layer.... No: The real opportunity I expect to see here is an integrated comprehensive solution, that includes security, compliance, monitoring, financial metering, end-user provisioning portals, etc.  The winners in the space will either integrate these features too, or offer a pre-integrated bundle of best-of-breed point-products. And it's going to happen soon.

Not Yet Sexy: The Integration Bus
If you believe that most of IT's infrastructure is "going cloud" and that many 3rd-party services will be SaaS, then the role of the CIO will begin to shift from being a technologist (who builds things from ground-up) to being a Supply-Chain manager (who integrates multiple services from multiple sources).

To execute on its new role in the enterprise, IT will therefore become an integration point for internally- and externally-generated services. It will need to provide core identity, compliance, data exchange, security, and access infrastructure to properly "broker" all of these diverse services, providers and APIs.

It seems to me that the notion of an integration bus will become crucial. Such a bus will provide the "glue logic" between all services, and avoids tedious hand-coded integration points. In many ways this bus is a core intersection point between the internal and external cloud - the hybrid nexus, if you will.

Remember, it's not enough to have a service in a cloud. There will be a huge need for coordinating the interactions into/out-of (and between) cloud-sourced services. The need for such a service bus will quickly elevate to that of a critical IT enabler. A good bet to take.

Getting Sexy: IT Business Management
Building on the concept of IT as a "supply chain manager" is the concept of IT as a Service Provider to the business. This is sometimes termed IT-as-a-Service, where IT begins to run itself as a "business". While it may not literally have a profit motive, it will be forced to become functionally competitive with external services, to market itself, and to price itself competitively.

As I've written before, in order to think like a business, one critical enabler is to know fixed and variable costs, as well as cost allocations and cost sensitivities.

Thus, the segments known as IT Financial Management (ITFM), IT Business Management (ITBM) and/or Technology Business Management (TBM) will necessarily have to expand in importance.  This segment looks at the granular cost basis of infrastructure (including cloud service costs) and assembles a composite cost structure at the application and even service level. This permits IT to understand cost sources, provide decision support and forecasting for infrastructure changes, and provide a "bill of IT" showback/chargeback to internal service customers.  All critical in the new Transformed view of IT.

Too Sexy: Big Data Business Models

Careful here... I'm not speaking about Big Data infrastructure or analytics products per se. I'm referring to the business models that will soon be based on mashing-up and analyzing large quantities of structured/unstructured data to uncover new revenue opportunities.

As I recently mentioned, data-based business models will begin to complement existing product/service based businesses.  While adoption of Big Data technology is already penetrating the infrastructure market, the real opportunity is when it penetrates the lines-of-business people such as marketers, sales, and business development.

Companies that employ marketing professionals paired with "data scientists" are the ones to watch. The new business opportunities presented by the promise of cracking big data will be the driving force behind the infrastructure and technology purchases.

When will this transition happen?  I predict when companies begin to hire "Big Data product managers" and team them with BizDev and data scientists. Now that's disruptive and sexy.


Sunday, February 5, 2012

The Growth of Data Growth: My Digital Contrail

Although I'm not a Big Data aficionado, I'd recently been struck by a few statistics from the IDC Digital Universe study:  It is estimated that in 2011, 1.8 Zettabytes of information was created, 75% of which comes from individuals. And by 2015, that number may grow to 7.9 Zettabytes.

So, just where does all of this data come from?  And the more I thought about it, the more I realized that each one of us (in developed countries, at least) kicks-off a near constant massive stream of data that gets stored somewhere, even if only transiently.

In our seemingly innocuous day, we surely generate more data than we consume, some of which is captured by others, and only some which we might be privileged enough to retain.

Thus I sat down to think about a typical "digital contrail" that I might generate. While I really can't quantify exactly how much stored data is created from each transaction, simply ball-parking the numbers would seem to support the Digital Universe claims.  And, if any of you reading this can help me quantify some of this, I'd be happy to append this post. Thanks in advance.

Sources of my "Digital Contrail"
  • I make a cell phone call:  Phone location tracking data (i.e. from towers) created and tracked by the carrier; phone log files; data created and stored by multiple mobile apps and their own hosting infrastructures 
  • I browse the web: Site tracking; clickstream storage; site analytics; Email storage, including replication on devices as well as replication in geographical-mirrored data centers.
  • Driving my car: Location-based tracking by RFID tags at toll booths; unique instrumentation data such from as OnStar systems
  • Go to the bank: data streams initiated and stored from a simple ATM withdrawals; security analysis of banking transaction patterns; audit and verification trails for individual transactions; mirrored/backed-up data within the bank's data center
  • Go to the store: data streams initiated and stored from a simple credit card transaction; product inventory changes; buying patterns stored and allocated to individual affinity discount programs
  • Browse an online store: All of the above, plus clickstream storage and analysis
  • Plan some travel: Airline reservations & pricing systems such as SABRE ticketing; airline tracking databases; TSA flyer database updates & analysis
  • At my home: electricity usage via smart metering data collection
  • Using entertainment: Uploaded Photography and Video; sales pattern data and DRM data 
  • Go to the doctor’s office: Medical imaging , EMR data, reports, other records
  • Somewhere in the background: With everything I do, there are surely security systems,  kicking-off background data processes and analytics DB’s
  • Also somewhere the background: Every service is sourced from a data center, where all data (including device data) is surely replicated and backed-up, including log files.
I finally thought through a simple habit I had, and how much storage space it spawned: I would receive an email with a PDF attachment - and carefully file the email in a folder while copying the PDF also into a separate folder related to the project. So I'd have 2 copies of the file on my PC, not to mention another copy on the Exchange server as well as one on the PC backup server file - 4 in all (assuming no deduplication system was in place). And if the email had been sent to others besides me... you get the point.  I was suddenly sensitized to data growth on a personal scale.

So, now I've convinced myself that "the data's out there". I've created scads of data in the past 24 hour stint... and fortunately (or unfortunately) it's all recorded in different repositories. But now I begin to wonder - what *if* some of these structured and unstructured data streams were re-constructed, mashed-up and analyzed?  That bit makes me both nervous (from a privacy and security perspective) and excited (from a Big Data and personalization perspective).  More later when I stop to think about that one.



    Thursday, January 19, 2012

    It Will Be a Data-Centric (Cloudy) World

    (Or, Where Tomorrow's Clouds Will Form)

    Move Mohammed to the mountain, or the mountain to Mohammed? 

    In the context of data, applications and cloud computing, this question takes on a new perspective - and the role of Mohammed and the Mountain may soon reverse.


    In the traditional application-centric (and static infrastructure) world, the Application is the immovable "mountain". Like a magnet, the app is permanently-located, attracting to it local data stores and peripheral support apps.  Administrators dote around it like worker bees around the queen.

    But for some uses and applications this may all change - altering with it the how-and-where compute and community clouds form.

    Observation #1: Apps are becoming mobile

    With increased use of a virtualization layer, migration tools, shared storage, fat network pipes, and virtual I/O and switching, we are all now realizing that where the executable application code resides is becoming far less important. Everything is just data - and can be moved/migrated. In mature virtual environments, VMs typically move between servers because of maintenance windows, because of capacity adjustments, etc.  But when VMs move move between physical data centers (separated by many miles or more) there is often a data movement as well. But there's no denying that the application is becoming more mobile.

    We also have the emergence of large data arrays and analytics appliances that embed internal servers that speed queries and analysis. VM's typically run on top of these servers being migrated in-and-out of the arrays as workloads and queries change. Hang on to this visual... we'll come back to it later.

    Observation #2: Big Data is becoming bigger

    When we start talking about hundreds of Terabytes - or even Petabytes - of structured/unstructured "big data", moving that data becomes increasingly physically difficult. Where it's generated is generally where it stays. Think about financial stock exchanges; retail data warehouses; medical imaging; geologic or climatological data. These stores are now becoming big and immovable.

    So, enterprises are now locating these data stores within critical data centers - within which they are co-locating the applications that require frequent access to that data. Sometimes that proximity is sufficient, and sometimes the analytics may even move within the array. But any way you look at it, those data stores are becoming the center of attention, around-which the applications now congregate.

    An interesting shift. But wait, there's more...
    ·     
    Outcome: Where Tomorrow's Clouds Might Form

    So let's expand this model from enterprise data centers to public clouds. Or even to "community clouds".

    Take the example of financial exchange data - NYSE’s Capital Markets Community Platform about which I blogged last year.  Here we have a special-purpose, "community cloud" - optimized for financial institutions, wherein they can locate trading and analytics applications. (Imagine a 3-person hedge-fund startup needing infrastructure). Operationally it's fit-for-purpose, with a high performance, low-latency compute backbone, with a common security/compliance envelope. But it's got another trait: At its core is a historic data warehouse of every tick for every trade. Now that's big data.

    If you think of it, the NYSE data store has become the "Mountain" around which applications (supplied by the cloud tenants) now congregate.They run their algorithms and analytics against the local data store. The data within the community cloud has become the anchor, the magnet. The apps are moved to be near the data.... not the other way around.  Remember that data array that had embedded VMs? Well, think of this model as that array on steroids.

    So, what might this mean if you want to build a differentiated cloud computing resource - say, targeting a specific industry vertical?  It says to me that the world will shift to a data-centric model. Focus on amassing and maintaining massive high-value data, all (presumably) requiring a similar security/compliance model. And then build a business by allowing tenants access by co-locating their applications in the same cloud as the data resides.

    I could see this transforming both the cloud service provider ecosystem, as well as entire industry groups. Consider new Cloud Service Provider models:  What if NOAA formed the Weather and Atmospherics Community Platform? If healthcare companies created federated Medical Records Community Platforms? If the USGS formed the World Geologic Community Platform? If other brokerages created equivalent capital markets platforms? 

    Cloud computing is shifting lots of conceptual IT models these days. But while you're considering what Cloud makes possible for applications, spend some time wondering what data makes possible for the Cloud.

    Other References

      Sunday, May 29, 2011

      It's All Just Data to Me

      Now that I’ve been with EMC for a few months, my relationship to storage, computing, and networking has once again shifted. And, in the context of the cloud computing operations model, my relationship to the physical location of data - and processing of that data - has shifted too.

      My new perspective starts with Computer Science 101: Where, at its heart, computing is simply data and instructions (stored on similar media) which are combined on a device (CPU) and produce an output.

      Since computing began, this model was consistent – but as the data and instructions grew in size and abstraction, the media changed to the point where instructions (code) and data, were each stored in physically separate locations.

      Until recently the data and instructions would be transported (over the network) to individual physical CPUs (with their own sets of OS) where they would be combined and executed. And then, the resulting data generally was transported back to its place of residence.

      Servers are Just Bits

      Now, enter the Virtual Machine.  At the heart of it, it's simply another file (e.g. VMDK) – in other words, just more data.

      So in the modern virtualized data center, what we have – at the extreme – is a model where not only the data and instructions are bits… but the servers are bits too. All they require are physical CPUs to execute.

      In the ‘traditional’ model, the data and instructions were brought to where the physical servers and O/S were.  But today, with pervasive farms of generic physical servers, we have the situation where *either* the data bits can be brought to the server, or the server bits can be brought to the data.

      Some of the implications you’ve probably already thought of – such as vMotion of a VM from one physical server to another, or using a DRS-style control to re-locate VMs from failed physical compute resources elsewhere.

      But consider another situation that’s happening with increasing frequency: The need to work with “Big Data” – such as running analytics on unstructured bits that could be on the Terabyte to Petabyte scale.  Here is a case where it makes sense to send Mohamed to the mountain than the other way around… To literally re-locate the servers (which are, after all just data themselves) closer to, or co-incident with, the data.

      Or, consider a “follow-the-moon” strategy for data center energy efficiency: where the most energy-efficient (and least expensive) physical servers are chosen to handle workloads. Once again, the data (which includes the virtual server, data and instructions) is simply transported to the optimal set of physical processing resources.

      Cloud Infrastructure and Data Management

      From where I sit, the importance of data storage, data management and data portability suddenly becomes paramount. It can reasonably be argued that physical servers are now merely execution platforms for the VM data bits, and that the network is simply becoming flatter and fatter.

      So the future data center and cloud model might be thought about as a data management problem. Where and how to locate bits, back-up bits, scale bits, operate on bits.   True, this is a data-centric view of the world. But it's also a healthy perspective from which to view the renewed importance of data and its dynamics, versus the other more static components of the data center.