As expected from conference schedule, Google conference turned out to be a technical event primarily focused on parallel programming and infrastructure scalability. At last minute, Google decided to merge two tracks in to one. Though, I got to attend all the sessions, they felt time-compressed and rushed. I was surprised to see lot of attendees who came from outside Seattle. I met quite a few people from Bay area, Canada and Europe. I enjoyed the sessions though some audience members commented about very technical nature of the conference compared to previous year. As Brian Bershad, Google commented in his welcome speech, the challenge is to find technologies and solutions to scale handling search queries from 600 million to 6 billion. And, I came away better informed on different challenges and potential solutions we may see down the road.
I also sat down and chatted with Robin Harris. We decided to forego making a video of our conversation. I am not a big fan of talking head videos or podcasts unless they leverage the unique values of these methods not available through written words or pictures. And who wants to listen to two storage bloggers chatting about nothing. I find them miserable myself so why put others through the same misery.
In my opinion, three sessions: CARMEN: a Scalable Science Cloud [PDF], GIGA+: Scalable Directories for Shared File Systems [PDF] and maidsafe stood out at the conference from infrastructure scalability perspective. Communicating Like Nemo was very entertaining. The common theme in audience questions on most infrastructure presentations was reliability, availability, scalability, and security of the offered solution. It is a good indication of what is on the mind of people when evaluating new infrastructure offerings. With the popularity of hashing in storage of data, speeding up hash lookup is becoming an interesting problem for scalability.
David Irvine's session on maidsafe was the only session where a speaker white-boarded most of the presentation. His confidence and knowledge was commendable. Not many speakers can pull off white-boarding 80% of presentation with 100s in audience. Comparing maidsafe with ant colony was an interesting way to show scalability and simplicity of solution. Maidsafe solution seems to be in same category as RevStor, Seanodes, Cleversafe, Oceanstor, Farsite and several others that are trying to leverage storage across 100s and 1,000s of distributed nodes in a peer-to-peer or quid pro quo network, a solution most likely attractive to players in cloud and web distribution market.
Showing posts with label Data Center. Show all posts
Showing posts with label Data Center. Show all posts
Monday, June 16, 2008
Sunday, March 16, 2008
Bandwidth, one hurdle in adopting Cloud Storage
This weekend, I read NY Times article Video Road Hogs Stir Fear of Internet Traffic Jam.
If you review the introduction and growth of various Amazon Web Services (AWS), a comparatively established cloud player, you will notice very limited use cases of Simple Storage Service (S3) on its own with clients outside the cloud. Most S3 usage is fronted by another AWS in the cloud such as Elastic Compute Cloud (EC2). Such combinations overcome the challenge of transferring large amount of data between storage cloud and an application/server outside the cloud over Internet. For cloud storage to be successful, it need to be in the same cloud with application/server or connected to application/server cloud with high speed link.
Any technology that can reduce the data transfer between the cloud services and clients outside the cloud will be the big beneficiary in this trend. Caching, Compression, and Data De-duplication will most likely benefit in the near term. And, the future seems to be very much like the past aka mainframe - Desktop Virtualization, Streaming, and On-the-Fly Visualization.
So, how will new cloud players like Nirvanix, EMC Mozy and Rackspace differentiate?
Last year, by one estimate, the video site YouTube, owned by Google, consumed as much bandwidth as the entire Internet did in 2000. …While reading the article, it occurred to me that isn't bandwidth going to be the main hurdle in adoption of storage in the cloud. When clients are not happy with 10/100/1000Mbps connection with application/server/data center, how can they be happy with DSL/Cable/T1/T3 connection to the cloud? I am sure everyone has felt the pain of trying to transfer large datasets over the Internet.
In a widely cited report published last November, a research firm projected that user demand for the Internet could outpace network capacity by 2011. …
Moving images, far more than words or sounds, are hefty rivers of digital bits as they traverse the Internet’s pipes and gateways, requiring, in industry parlance, more bandwidth.
If you review the introduction and growth of various Amazon Web Services (AWS), a comparatively established cloud player, you will notice very limited use cases of Simple Storage Service (S3) on its own with clients outside the cloud. Most S3 usage is fronted by another AWS in the cloud such as Elastic Compute Cloud (EC2). Such combinations overcome the challenge of transferring large amount of data between storage cloud and an application/server outside the cloud over Internet. For cloud storage to be successful, it need to be in the same cloud with application/server or connected to application/server cloud with high speed link.
Any technology that can reduce the data transfer between the cloud services and clients outside the cloud will be the big beneficiary in this trend. Caching, Compression, and Data De-duplication will most likely benefit in the near term. And, the future seems to be very much like the past aka mainframe - Desktop Virtualization, Streaming, and On-the-Fly Visualization.
So, how will new cloud players like Nirvanix, EMC Mozy and Rackspace differentiate?
Monday, July 16, 2007
Power Consumption by Google Services
Even though, Google doesn’t share a lot of details of their infrastructure, as we have seen from limited published information, they are obsessed with continuously monitoring, managing and improving the efficiency of their infrastructure.
Recently, Robin Harris attended the Google conference on scalability and then mused How Yahoo can beat Google. Few months ago, Google published results of their work on disk drive failure in paper Failure Trends in a Large Disk Drive Population [PDF]. It was extensively covered in blogosphere including by me in blog entries SMART not so smart in predicting disk drive failure and Google Findings of Disk Failures Rates and Implications and by Robin Harris in his blog entry Google’s Disk Failure Experience.
Google has done it again and presented results of their work on power consumption and provisioning in paper Power Provisioning for a Warehouse-sized Computer [PDF] at the ACM International Symposium on Computer Architecture, San Diego CA, June 9 – 13, 2007. In this work, Google researchers, Xiaobo Fan, Wolf-Dietrich Weber and Luiz Andre Barroso looked in to 15,000 servers running three different applications – Websearch, Webmail and Mapreduce for six months to determine the power usage characteristics at Rack, PDU and Datacenter levels.
Google Services
Websearch: A service with high request throughput and large data processing for each request.
Webmail: A disk I/O intensive service. Machines configured with large number of disk drives. Each request involves a relatively small number of servers.
Mapreduce: A cluster dedicated to running large offline batch jobs. Involve process terabytes of data using thousands of machines.
Key Findings
The key findings from this work are:
More details from this study later.
Recently, Robin Harris attended the Google conference on scalability and then mused How Yahoo can beat Google. Few months ago, Google published results of their work on disk drive failure in paper Failure Trends in a Large Disk Drive Population [PDF]. It was extensively covered in blogosphere including by me in blog entries SMART not so smart in predicting disk drive failure and Google Findings of Disk Failures Rates and Implications and by Robin Harris in his blog entry Google’s Disk Failure Experience.
Google has done it again and presented results of their work on power consumption and provisioning in paper Power Provisioning for a Warehouse-sized Computer [PDF] at the ACM International Symposium on Computer Architecture, San Diego CA, June 9 – 13, 2007. In this work, Google researchers, Xiaobo Fan, Wolf-Dietrich Weber and Luiz Andre Barroso looked in to 15,000 servers running three different applications – Websearch, Webmail and Mapreduce for six months to determine the power usage characteristics at Rack, PDU and Datacenter levels.
Google Services
Websearch: A service with high request throughput and large data processing for each request.
Webmail: A disk I/O intensive service. Machines configured with large number of disk drives. Each request involves a relatively small number of servers.
Mapreduce: A cluster dedicated to running large offline batch jobs. Involve process terabytes of data using thousands of machines.
Key Findings
The key findings from this work are:
- The difference between maximum power used by large number of computing devices, cumulatively, and their theoretical peak usage can be as much as 40% in datacenters.
- It may be more efficient to leverage power management techniques at datacenter level than at rack level.
- Nameplate ratings are of little use in power provisioning as they significantly overestimate actual maximum usage.
- CPU utilization as a measure of machine-level activity produces accurate results for dynamic power usage especially with large group of machines. The dynamic power range is less than 30% for disks and negligible for motherboards.
- Using maximum power draw of individual machines to provision the datacenter, will have some stranded capacity.
- A mix of diverse workload reduces the difference between average and peak power, an argument in favor of mixed deployment.
- Idle power is significantly lower than the actual peak power, but generally never below 50%.
- CPU dynamic voltage/frequency scaling may yield moderate energy savings (up to 23%) at datacenter levels.
- Peak power consumption at the data center level could be reduced by 30% and energy usage could be halved if systems were designed so that lower activity levels meant correspondingly lower power usage profiles.
More details from this study later.
Wednesday, June 20, 2007
Gear6 trailblazing Network Caching
Earlier this week, I had great conversation with Gary Orenstien and Jack O’Brien at Gear6. Here are the excerpts from our conversation.
How is Gear6 doing?
Gear6 seems to be doing well. Several units are currently in field being evaluated by various customers. No specific number of units provided, just a wide range between 10 and 100. Company has over thirty employees and financially all set in the near term. Company has started to focus CACHEfx on financial analytics, energy and animation segments and will expand focus by the end of the year.
What are the benefits of network based caching?
Network caching enables increased cache utilization, flexibility and scalability. Caching is moving from end devices to network and becoming a network resource.
What one factor is attracting customers to your caching solution?
By nature of caching, the obvious benefit to customer is performance. Most customers who come to Gear6 have performance problems, variable workload and demand certain Quality of Service. The success rate is very good with evaluations by customers as CACHEfx appliance doesn’t require forklift replacement.
How is Gear6 doing caching?
CACHEfx appliance doesn’t use any conventional mechanical disk storage internally, 100% RAM cache and is pass through to persistent storage. Robust single purpose appliance designed to do one job and do that job very well.
The caching is performed intelligently. The intelligence focus on how and where data is placed within the appliance. There are extensive built-in statistics. Most customers are impressed by network sniffer like capability.
In the past, cache was a constrained resource. Now, focus is on right-sizing cache. CACHEfx expands from quarter TB to multi-TB, can be preloaded with data from persistent storage and adjust to variable I/O profile.
What are the reliability, availability and scalability features of CACHEfx appliance?
It is a clustered appliance, scalable from quarter TB to multi-TB. The appliance can be expanded on the fly. Also, appliance only acknowledges writes only when persistent storage sends acknowledgment.
Is the CACHEfx installed at D.E. Shaw working with Solaris cluster?
Gary declined to comment on infrastructure details of customer. He claimed customer pleased with the solution.
Any plans to introduce network caching for block-level traffic? The present product seems to focus on NFS only.
The present focus is on NFS, market is large enough. The sweet spot is where customer is using 100+ concurrent clients accessing single dataset, most tend to be NFS. No firm plans for addressing CIFS or block level traffic. The primary industry focus on financial analytic, energy and exploration, electronic design, animation, biotechnology, and media, primarily HPC oriented tasks.
How does network caching stack up with parallel file systems and clustered storage?
Caching addresses I/O constrained systems rather than processing constrained. Parallel file systems and clustered storage solutions are capacity centric not performance centric, providing global namespace for ever expanding storage capacity. They are not low latency solution. Network caching is a complementary solution, capacity complemented by performance. Gear6 solution complements Netapp OnTAP GX, IBRIX, Isilon and Acopia.
Do you have any thoughts on potential application of CACHEfx in a Wide Area Filer Network environment?
The CACHEfx has enormous potential in variety of environment. But we are currently very focused on solving customer problems within the data center. We are open to partnerships in other areas.
How is Gear6 doing?
Gear6 seems to be doing well. Several units are currently in field being evaluated by various customers. No specific number of units provided, just a wide range between 10 and 100. Company has over thirty employees and financially all set in the near term. Company has started to focus CACHEfx on financial analytics, energy and animation segments and will expand focus by the end of the year.
What are the benefits of network based caching?
Network caching enables increased cache utilization, flexibility and scalability. Caching is moving from end devices to network and becoming a network resource.
What one factor is attracting customers to your caching solution?
By nature of caching, the obvious benefit to customer is performance. Most customers who come to Gear6 have performance problems, variable workload and demand certain Quality of Service. The success rate is very good with evaluations by customers as CACHEfx appliance doesn’t require forklift replacement.
How is Gear6 doing caching?
CACHEfx appliance doesn’t use any conventional mechanical disk storage internally, 100% RAM cache and is pass through to persistent storage. Robust single purpose appliance designed to do one job and do that job very well.
The caching is performed intelligently. The intelligence focus on how and where data is placed within the appliance. There are extensive built-in statistics. Most customers are impressed by network sniffer like capability.
In the past, cache was a constrained resource. Now, focus is on right-sizing cache. CACHEfx expands from quarter TB to multi-TB, can be preloaded with data from persistent storage and adjust to variable I/O profile.
What are the reliability, availability and scalability features of CACHEfx appliance?
It is a clustered appliance, scalable from quarter TB to multi-TB. The appliance can be expanded on the fly. Also, appliance only acknowledges writes only when persistent storage sends acknowledgment.
Is the CACHEfx installed at D.E. Shaw working with Solaris cluster?
Gary declined to comment on infrastructure details of customer. He claimed customer pleased with the solution.
Any plans to introduce network caching for block-level traffic? The present product seems to focus on NFS only.
The present focus is on NFS, market is large enough. The sweet spot is where customer is using 100+ concurrent clients accessing single dataset, most tend to be NFS. No firm plans for addressing CIFS or block level traffic. The primary industry focus on financial analytic, energy and exploration, electronic design, animation, biotechnology, and media, primarily HPC oriented tasks.
How does network caching stack up with parallel file systems and clustered storage?
Caching addresses I/O constrained systems rather than processing constrained. Parallel file systems and clustered storage solutions are capacity centric not performance centric, providing global namespace for ever expanding storage capacity. They are not low latency solution. Network caching is a complementary solution, capacity complemented by performance. Gear6 solution complements Netapp OnTAP GX, IBRIX, Isilon and Acopia.
Do you have any thoughts on potential application of CACHEfx in a Wide Area Filer Network environment?
The CACHEfx has enormous potential in variety of environment. But we are currently very focused on solving customer problems within the data center. We are open to partnerships in other areas.
Sunday, June 17, 2007
Bountiful Bandwidth Lagging Latency
Recently, I came across an interesting article published in 2004 comparing growth, reasons and handling imbalance between bandwidth and latency. Excerpts below are from Latency Lags Bandwidth, Recognizing the chronic imbalance between bandwidth and latency, and how to cope with it. By David A. Patterson, Communications of the ACM, October 2004/Vol. 47, No. 10.
In the time that bandwidth doubles, latency improves by no more than a factor of 1.2 to 1.4.Reasons for Bountiful Bandwidth
“There is an old network saying: Bandwidth problems can be cured with money. Latency problems are harder because the speed of light is fixed – you can’t bribe God” – Anonymous.Coping with Lagging Latency
Moore’s Law helps bandwidth more than latency.
Distance limits latency.
Bandwidth is generally easier to sell.
Latency helps bandwidth.
Bandwidth hurts latency.
Operating system overhead hurts latency.
Caching: Leveraging capacity to help latency.Marketing Latency Innovations
Replication: Leveraging capacity to again help latency.
Prediction: Leveraging bandwidth to again help latency.
The difficulty of marketing latency innovations is one of the reasons latency has received less attention thus far.
Perhaps, we can draw inspiration from the more mature automotive industry, which advertises time to accelerate from 0-to-60 miles per hour in addition to peak horsepower and top speed.
Tuesday, June 12, 2007
Where do you focus, Bandwidth or Latency?
Since my first post about Gear6, Gary Orenstein and I have been exchanging emails discussing various aspects of storage caching and Gear6. Recently, he commented in response to my request for pointers on storage caching market and implementations:
What problems does caching solve?
The major benefit of caching is in reducing the latency whether caching is part of the web, network, file system, storage device, processor or memory. What is latency? Any delay in response to a request.
Bandwidth Bias
One consistent theme struck me odd as I started studying caching is how often we suggest more bandwidth as a solution to the slow performance issues and how little focus we give to the latency side of the problem. What is bandwidth? The amount of data carried from one point to another in a given time.
Even in iSCSI world, we all hear how 10GbE will be the inflection point, indirectly giving the impression that bandwidth is the bottleneck in iSCSI adoption. What is the real bottleneck in iSCSI? Is it bandwidth or latency?
I guess it sounds more impressive "With 10GbE, the bandwidth will increase 10X so you will be able to push ten times of data but latency will only be reduced in half (approx)."
From the productivity aspects of users and applications, a predictable and quick response to a request seems to be considerably more important than the amount of data being transferred over a specified period. What good more bandwidth does if data needs to wait for processing? A balance between bandwidth and latency need to be considered in designing solutions.
In the end, my impression is that most of us tend to focus too much on bandwidth and too little on latency.
When I find interesting items related to caching I usually post on our blog. The thing is, there really hasn't been anyone promoting network-based caching until Gear6.With rising interest in flash memory and SSDs, I am finding storage caching quite intriguing. I decided to start from basics.
What problems does caching solve?
The major benefit of caching is in reducing the latency whether caching is part of the web, network, file system, storage device, processor or memory. What is latency? Any delay in response to a request.
Bandwidth Bias
One consistent theme struck me odd as I started studying caching is how often we suggest more bandwidth as a solution to the slow performance issues and how little focus we give to the latency side of the problem. What is bandwidth? The amount of data carried from one point to another in a given time.
Even in iSCSI world, we all hear how 10GbE will be the inflection point, indirectly giving the impression that bandwidth is the bottleneck in iSCSI adoption. What is the real bottleneck in iSCSI? Is it bandwidth or latency?
I guess it sounds more impressive "With 10GbE, the bandwidth will increase 10X so you will be able to push ten times of data but latency will only be reduced in half (approx)."
From the productivity aspects of users and applications, a predictable and quick response to a request seems to be considerably more important than the amount of data being transferred over a specified period. What good more bandwidth does if data needs to wait for processing? A balance between bandwidth and latency need to be considered in designing solutions.
In the end, my impression is that most of us tend to focus too much on bandwidth and too little on latency.
Tuesday, April 24, 2007
Storage Vendors to Watch: Gear6
Excerpts from Gear6 website:
… keeping frequently accessed data in a very large central memory pool … This enables high performance data access by avoiding time-consuming disk operations and accelerates applications due to dramatically decreased response times and increased data throughput.Actually, it was quite amusing at the conference. Most probably, I steered few end-users to Gear6 by suggesting to check out G6 caching appliance. These end users told me that they are using Oracle databases with NFS and performance being one of their pain points. I found three simple questions that can quickly tell whether someone may want to investigate G6 product.
This innovative approach complements existing NAS/NFS deployments and installs transparently in the data center without requiring changes to current applications or infrastructure.
- Are you using transaction databases?
- Do you use NFS mounts?
- Do you have performance issues?
BTW, you may want to hop over to Thoughtput blog maintained by Gary Orenstein at Gear6 for more caching related information. Also check out these presentations from Gear6.
Share your thoughts on Gear6 and its caching appliance approach.
Friday, March 02, 2007
Distributing Desperate Housewives to Ten Millions
Now the title and image caught your attention, the big let down is this post has no housewives to offer! It is about distributing episodes of ABC Television show “Desperate Housewives” over Internet … or may be not even that!!
After my previous post P2P powered Devices … coming soon?, Newell Edmond co-founder of GridNetworks forwarded me an interesting paper on Video Internet. And this paper led me to March 2, 2006 column by Robert Cringely Peering into the Future: Why P2P is the Future of Media Distribution even if ISPs have yet to Figure that out.
What type of storage infrastructure ecosystem will someone need to fulfill Ten million requests for distributing one episode of Desperate Housewives?
In my opinion, a storage infrastructure built around monolithic centralized storage most probably wouldn’t be practical. But this post is not about my opinion. It is about yours, so chime in with your thoughts on potential solution to this problem.
Show your design prowess or extol virtues of your favorite storage vendors with your storage ecosystem design. All responses are welcome.
After my previous post P2P powered Devices … coming soon?, Newell Edmond co-founder of GridNetworks forwarded me an interesting paper on Video Internet. And this paper led me to March 2, 2006 column by Robert Cringely Peering into the Future: Why P2P is the Future of Media Distribution even if ISPs have yet to Figure that out.
"Desperate Housewives," in its puny 320-by-240 iTunes incarnation, occupies an average of 210 megabytes per episode. A full-resolution version would be larger still. In theory, it would be four times as big, but practically it would probably come in at double the size or 420 megabytes. But let's stick with the little iTunes version for this example.Even though, Cringely was discussing bandwidth challenges of transferring one episode of Desperate Housewives, my mind wandered off to storage infrastructure side of the equation.
Twenty million viewers, on average, watch "Desperate Housewives" each week in about 10 million U.S. households. That's 210 megabytes times 10 million downloads, or 2.1 petabytes of data to be downloaded per episode. Fortunately for the download business model, not everyone is trying to watch the show at the same time or in real time, so iTunes, in this example, has some time to do all those downloads. Let's give them three days. The question on the table is what size Internet pipe would it take to transfer 2.1 petabytes in 72 hours? I did the math, and it requires 64 gigabits-per-second, which would require an OC-768 fiber link and two OC-256s to fulfill.
What type of storage infrastructure ecosystem will someone need to fulfill Ten million requests for distributing one episode of Desperate Housewives?
In my opinion, a storage infrastructure built around monolithic centralized storage most probably wouldn’t be practical. But this post is not about my opinion. It is about yours, so chime in with your thoughts on potential solution to this problem.
Show your design prowess or extol virtues of your favorite storage vendors with your storage ecosystem design. All responses are welcome.
Monday, February 26, 2007
Improving ROI of IT management
Most IT departments are having difficulties in addressing IT infrastructure monitoring and management requirements due to
Incorporating device configuration and control functions to unified management applications is considered the major hurdle in adoption of such applications. But, IT administrators typically prefer to work with feature-rich device-specific configuration and control applications instead of using a generic unified management product. Most device vendors also provide limited functionality beyond monitoring and basic reporting to unified management platforms limiting value delivered through them. Rightly so, it also plays in to lock-in strategy and defense against out-of-sight out-of-mind perception for device vendors.
As more and more devices being installed in customer environment with email capabilities, real-time monitoring and auto-support from vendors, these management functionalities are becoming quite burdensome for the IT administrators, if not potential security risks. Just think aboutseveral dozen devices with frequent out-bound emails, monitoring pings and SNMP alerts in a data center and the management headache these auto-support functions can create.
So, how can the IT management vendors increase the value delivered to the customers with monitoring, reporting and analysis tools and without device configuration and control abilities as well as without creating another layer of management headache? As mentioned in my last post Understanding the Web 2.0 Trends, this was a topic of my conversation with Scot French, VP of Marketing at Klir Technologies, a local startup, for last few weeks.

I particularly liked the Klir solution to IT monitoring, reporting and analysis platform market with scalability of Software-as-a-Service (SaaS) and collective intelligence of “2.0”. I believe their approach brings the ease of use, perpetual upgrades, contextual content and community approach to IT management that is not available from enterprise IT management software packages.
Your opinions are welcome on IT monitoring, reporting and analysis market, SaaS approach and leveraging “2.0” to enhance ROI of IT management tools.
- Limited resources
- Increasing complexity of infrastructure
- Growing number and type of devices in an environment, and
- Demands for better usage, performance and uptime reporting
Incorporating device configuration and control functions to unified management applications is considered the major hurdle in adoption of such applications. But, IT administrators typically prefer to work with feature-rich device-specific configuration and control applications instead of using a generic unified management product. Most device vendors also provide limited functionality beyond monitoring and basic reporting to unified management platforms limiting value delivered through them. Rightly so, it also plays in to lock-in strategy and defense against out-of-sight out-of-mind perception for device vendors.
As more and more devices being installed in customer environment with email capabilities, real-time monitoring and auto-support from vendors, these management functionalities are becoming quite burdensome for the IT administrators, if not potential security risks. Just think aboutseveral dozen devices with frequent out-bound emails, monitoring pings and SNMP alerts in a data center and the management headache these auto-support functions can create.
So, how can the IT management vendors increase the value delivered to the customers with monitoring, reporting and analysis tools and without device configuration and control abilities as well as without creating another layer of management headache? As mentioned in my last post Understanding the Web 2.0 Trends, this was a topic of my conversation with Scot French, VP of Marketing at Klir Technologies, a local startup, for last few weeks.
I particularly liked the Klir solution to IT monitoring, reporting and analysis platform market with scalability of Software-as-a-Service (SaaS) and collective intelligence of “2.0”. I believe their approach brings the ease of use, perpetual upgrades, contextual content and community approach to IT management that is not available from enterprise IT management software packages.
Your opinions are welcome on IT monitoring, reporting and analysis market, SaaS approach and leveraging “2.0” to enhance ROI of IT management tools.
Monday, February 19, 2007
SMART not so smart in predicting disk drive failure
Continuing from last blog post, Google report [PDF] also shares their analysis based on disk self-monitoring data and identifies important failure related SMART parameters.

Predictive models based on scan errors, reallocation count, offline reallocation count and probational count couldn’t predict more than half of the failed drives.
Scan Errors – Large scan error counts can be indicative of surface defects, and therefore indicative of surface defects.
Reallocation Count – When the drive logic believes that a sector is damaged it can remap the faulty sector number to a new physical sector drawn from a pool of spares. Reallocation count reflects the number of times this has happened, and is seen as an indication of drive surface wear.
Offline Reallocation – Offline reallocation are defined as subset of the reallocation counts in which only reallocated sectors found during background scrubbing are counted. In other words, it should exclude sectors that are reallocated as a result of errors found during actual I/O operations.
Probational Count – Disk drives put suspect bad sectors “on probation” until they either fail permanently and are reallocated or continue to work without problems. Probational counts can be seen as a softer error indication.
Seek Errors – Seek errors occur when a disk drive fails to properly track a sector and needs to wait for another revolution to read or write from or to a sector.
CRC Errors – CRC errors are detected during data transmission between the physical media and the interface.
- The drives with scan errors are 10x more likely to fail that the drives with no scan errors. 30% of the drives fail within the 8 months after first scan error. The failure probability is higher within first month of first scan error occurring in newer drives and then plateaus. With older drives, failure probability rises with time.
- The drives with reallocation count fail 3 – 6x more often than those with none. 15% of the drives fail within the 8 months after the first reallocation.
- There is no definite correlation between failure rates and seek errors.
- CRC errors are less indicative of drive failures than that of cables and connectors.
- There is no significant correlation between failures and high power cycle counts for drives less than two years old. For drives 3 years and older, higher power cycle counts can increase the absolute failure rate by over 2%.
Predictive models based on scan errors, reallocation count, offline reallocation count and probational count couldn’t predict more than half of the failed drives.
We conclude that it is unlikely that SMART data alone can be effectively used to build models that predict failures of individual drives. SMART parameters still appear to be useful in reasoning about the aggregate reliability of large disk populations, which is still very important for logistics and supply-chain planning.Glossary of Terms
Scan Errors – Large scan error counts can be indicative of surface defects, and therefore indicative of surface defects.
Reallocation Count – When the drive logic believes that a sector is damaged it can remap the faulty sector number to a new physical sector drawn from a pool of spares. Reallocation count reflects the number of times this has happened, and is seen as an indication of drive surface wear.
Offline Reallocation – Offline reallocation are defined as subset of the reallocation counts in which only reallocated sectors found during background scrubbing are counted. In other words, it should exclude sectors that are reallocated as a result of errors found during actual I/O operations.
Probational Count – Disk drives put suspect bad sectors “on probation” until they either fail permanently and are reallocated or continue to work without problems. Probational counts can be seen as a softer error indication.
Seek Errors – Seek errors occur when a disk drive fails to properly track a sector and needs to wait for another revolution to read or write from or to a sector.
CRC Errors – CRC errors are detected during data transmission between the physical media and the interface.
Sunday, February 18, 2007
Google Findings of Disk Failure Rates and Implications
Few months ago, I came to know that Google will be publishing a detailed report on disk drive failure rates in their environment. Prior to this report, best to my knowledge, there is very little information available from user perspective on this topic beyond limited work at Microsoft and Internet Archive.
Earlier, I thought of complimenting this report with a perspective from storage subsystem vendors. I was surprised to learn that either subsystem vendors don’t capture disk failure data effectively and in usable format or unwilling to share such data. And, there is very little published information on this topic from subsystem vendors.
In one case, a subsystem vendor recommended to contact disk drive manufacturers. Prior studies indicate that the actual drive replacement rate is 10 – 100x higher than failure rates published by disk drive manufacturers. Also, most likely disk drive manufacturers don’t have visibility in to deployment scenarios of failed drives for their data to be useful from the perspective of data centers and subsystem vendors. In the end, I decided to wait for Google report to come out to start this discussion.
Google Findings
I considered Google study to be unique because it looked at a very large sample of 100,000 disk drives. A summary of interesting results on age, manufacturers, read/write load and temperature from this study is listed below:
In my opinion, the second finding reaffirms the subsystem vendors’ stance on not match-and-mix drives of different model and manufacturers.
The third finding has wider implications. First implication is on the higher possibility of failure of new drive during RAID rebuild even though this study explicitly didn’t measure very young age in hours after operation. Second implication is on extending time period for an old drive set to be kept intact and in operation after a new drive set has been phased in. Third implication is on the higher possibility of losing disks in long-term archive storage where read/write load may not be high enough and the challenges in keeping disk read/write load high enough.
Earlier, I thought of complimenting this report with a perspective from storage subsystem vendors. I was surprised to learn that either subsystem vendors don’t capture disk failure data effectively and in usable format or unwilling to share such data. And, there is very little published information on this topic from subsystem vendors.
In one case, a subsystem vendor recommended to contact disk drive manufacturers. Prior studies indicate that the actual drive replacement rate is 10 – 100x higher than failure rates published by disk drive manufacturers. Also, most likely disk drive manufacturers don’t have visibility in to deployment scenarios of failed drives for their data to be useful from the perspective of data centers and subsystem vendors. In the end, I decided to wait for Google report to come out to start this discussion.
Google Findings
I considered Google study to be unique because it looked at a very large sample of 100,000 disk drives. A summary of interesting results on age, manufacturers, read/write load and temperature from this study is listed below:
- The failure rate varied from 1.7% for drives in their first year of operation to over 8.6% observed in their third year of operation.
- Confirmation of the fact from prior studies that failure rates are highly correlated with drive models, manufacturers and vintages.
- A complex correlation between high utilization, i.e. read/write load and higher failure rate instead of strong direct relationship as widely assumed.
- Surprising finding that lower temperatures are associated with higher failure rates and failures do not increase when the average temperature increases. The trend for higher failures with higher temperature is more pronounced for older drives.
First, only very young and very old age groups appear to show the expected behavior. After the first year, the AFR (annualized failure rate) of high utilization drives is at most moderately higher than that of low utilization drives. The three-year group in fact appears to have the opposite of the expected behavior, with low utilization drives have slightly higher failure rates than high utilization ones.
Overall our experiments can confirm previously reported temperature effects only for the high end of our temperature range and especially for older drives. In the lower and middle temperature ranges, higher temperatures are not associated with higher failure rates.
In my opinion, the second finding reaffirms the subsystem vendors’ stance on not match-and-mix drives of different model and manufacturers.
The third finding has wider implications. First implication is on the higher possibility of failure of new drive during RAID rebuild even though this study explicitly didn’t measure very young age in hours after operation. Second implication is on extending time period for an old drive set to be kept intact and in operation after a new drive set has been phased in. Third implication is on the higher possibility of losing disks in long-term archive storage where read/write load may not be high enough and the challenges in keeping disk read/write load high enough.
Friday, December 22, 2006
Success Factors for GridNetworks ... contd.
Continuing my thoughts on success factors for a video distribution infrastructure play like GridNetworks from previous post ...
As Michael Gersh commented in previous post, high quality video distribution will be viewer paid.
Who is going to collect payment from viewers? Will it be a content distribution infrastructure owner like Comcast or content distributor/aggregator like Netflix? Why is it important? IMO, it is the company in value chain that has most viewers captive benefits the most. And this is shown very clearly from some back of the envelope calculations for iTunes and Akamai.
Assuming weekly revenue of $10 million from analyst download estimates of $18.5 million songs per week, annual revenue of iTunes store, a content aggregator/distributor, is over $500 million+. Akamai, a distribution infrastructure provider to iTunes and with near monopoly in CDN, total revenues are barely in $400 million range.
With the success of iTunes, it is assumed that content aggregators/distributors are the ones who will be collecting payment from viewers. Distribution infrastructure owners like Akamai will be a service provider to iTunes for a fee.
Cost-side Success Factor
As GridNetworks (GN) will likely be paid by content distributors, it's profit-side success factor depend on number of content distributors using its delivery infrastructure and the revenue generated from each content distributor.
To attract paying content distributors and be a preferred delivery method, GridNetworks need to have the most expansive hybrid CDN P2P infrastructure. So the cost-side success factor for GN comes down to how quickly they can build 40 million quality nodes contributing to their delivery infrastructure and at what cost.
One method to achieve this goal is to freely distribute software for media sharing and playback. Once there are sufficient nodes established, harness those nodes and the brand recognition to make deals with content developers, owners and distributors. BBC deal with Azureus [pdf] will fall in to this category.
There are already enough high quality video content delivery startups trying to follow this route. Most with very little value differentiation originating primarily from the high profile and visible content deals. The success will belong to the ones with market/brand recognition, deep pockets and influence to make high-profile visible content deals.
Should GN follow the same path or there is another way to succeed? Chime in if you have any thoughts.
To be continued later ...
Housekeeping Notes
Recently, I noticed in Google Analytics stats that some traffic to my blog is coming from penny stock forums where link and text of my entries were posted. A caution note for readers from these forums: My posts are nothing more than personal rants and shouldn't be considered thoroughly researched analysis on prospects of any company, its stock, industry or market. Believe in my rants on your own perils.

This picture of a Pachinko in Tokyo seems quite appropriate after the above housekeeping note.
As Michael Gersh commented in previous post, high quality video distribution will be viewer paid.
Who is going to collect payment from viewers? Will it be a content distribution infrastructure owner like Comcast or content distributor/aggregator like Netflix? Why is it important? IMO, it is the company in value chain that has most viewers captive benefits the most. And this is shown very clearly from some back of the envelope calculations for iTunes and Akamai.
Assuming weekly revenue of $10 million from analyst download estimates of $18.5 million songs per week, annual revenue of iTunes store, a content aggregator/distributor, is over $500 million+. Akamai, a distribution infrastructure provider to iTunes and with near monopoly in CDN, total revenues are barely in $400 million range.
With the success of iTunes, it is assumed that content aggregators/distributors are the ones who will be collecting payment from viewers. Distribution infrastructure owners like Akamai will be a service provider to iTunes for a fee.
Cost-side Success Factor
As GridNetworks (GN) will likely be paid by content distributors, it's profit-side success factor depend on number of content distributors using its delivery infrastructure and the revenue generated from each content distributor.
To attract paying content distributors and be a preferred delivery method, GridNetworks need to have the most expansive hybrid CDN P2P infrastructure. So the cost-side success factor for GN comes down to how quickly they can build 40 million quality nodes contributing to their delivery infrastructure and at what cost.
One method to achieve this goal is to freely distribute software for media sharing and playback. Once there are sufficient nodes established, harness those nodes and the brand recognition to make deals with content developers, owners and distributors. BBC deal with Azureus [pdf] will fall in to this category.
There are already enough high quality video content delivery startups trying to follow this route. Most with very little value differentiation originating primarily from the high profile and visible content deals. The success will belong to the ones with market/brand recognition, deep pockets and influence to make high-profile visible content deals.
Should GN follow the same path or there is another way to succeed? Chime in if you have any thoughts.
To be continued later ...
Housekeeping Notes
Recently, I noticed in Google Analytics stats that some traffic to my blog is coming from penny stock forums where link and text of my entries were posted. A caution note for readers from these forums: My posts are nothing more than personal rants and shouldn't be considered thoroughly researched analysis on prospects of any company, its stock, industry or market. Believe in my rants on your own perils.
This picture of a Pachinko in Tokyo seems quite appropriate after the above housekeeping note.
Wednesday, December 13, 2006
GridNetworks, what’s my interest? Part two
Click here for the first part.
Third, GridNetworks piqued my interest with the use of word “Grid” in the company name. Grid is a powerful concept for utilizing idle resources and addressing infrastructure scalability. GridNetworks provides the extra scalability through P2P to fixed infrastructure of content delivery networks for superior user experience at the end of the last mile of Internet.
Few years ago, I researched a company for a client, Bycast and its StorageGRID product. During last bubble, Bycast was positioned as online video streaming startup. With its underlying technology, Bycast was successful in navigating turbulent times by repositioning in healthcare storage. Even though, Newell was quick to shoot down my comparison of GridNetworks with Bycast, I felt that the underlying technology of GridNetworks will find its ways in to other areas and applications too, not discounting its application in storage either.
Fourth, GridNetworks website listed Sujal Patel on its Advisory Board. Sujal is co-founder of Isilon Systems, a storage startup, soon to be public, and a veteran of RealNetworks. This piqued my interest to find out what connection Isilon or storage may have with GridNetworks.
In the end, it turned out to be nothing more than ex-RealNetwork links. Sujal is now acting more like angel investor. So, he doesn’t seem to be a good proxy anymore to find ties with Isilon or storage.
Fifth, the hints of P2P implementation in PowerGrid End User License Agreement piqued my interest as I saw a way to accomplish Distributed Storage Aggregation (DSA). And if you like me scan IEEE and ACM publications regularly, P2P way to utilize unused storage distributed across enterprise has started to become interesting.
Newell claimed about 50,000 nodes currently contributing on, average, 1GB of storage and 200kbps of bandwidth per node. His users are contributing almost 50TB of storage and 10Gbps bandwidth to deliver high quality video. When his vision of 40 million nodes comes to fruition, he may have access to 40 PB of storage without shelling out a single penny to any storage vendor.
The above reasons may explain what encouraged me to initiate contact with GridNetworks. One day, I expect to see the same technology being implemented with in the enterprise for distributed storage aggregation.
City view from Tokyo TV Tower. Mount Fuji is barely visible in the background.
Third, GridNetworks piqued my interest with the use of word “Grid” in the company name. Grid is a powerful concept for utilizing idle resources and addressing infrastructure scalability. GridNetworks provides the extra scalability through P2P to fixed infrastructure of content delivery networks for superior user experience at the end of the last mile of Internet.
Few years ago, I researched a company for a client, Bycast and its StorageGRID product. During last bubble, Bycast was positioned as online video streaming startup. With its underlying technology, Bycast was successful in navigating turbulent times by repositioning in healthcare storage. Even though, Newell was quick to shoot down my comparison of GridNetworks with Bycast, I felt that the underlying technology of GridNetworks will find its ways in to other areas and applications too, not discounting its application in storage either.
Fourth, GridNetworks website listed Sujal Patel on its Advisory Board. Sujal is co-founder of Isilon Systems, a storage startup, soon to be public, and a veteran of RealNetworks. This piqued my interest to find out what connection Isilon or storage may have with GridNetworks.
In the end, it turned out to be nothing more than ex-RealNetwork links. Sujal is now acting more like angel investor. So, he doesn’t seem to be a good proxy anymore to find ties with Isilon or storage.
Fifth, the hints of P2P implementation in PowerGrid End User License Agreement piqued my interest as I saw a way to accomplish Distributed Storage Aggregation (DSA). And if you like me scan IEEE and ACM publications regularly, P2P way to utilize unused storage distributed across enterprise has started to become interesting.
Newell claimed about 50,000 nodes currently contributing on, average, 1GB of storage and 200kbps of bandwidth per node. His users are contributing almost 50TB of storage and 10Gbps bandwidth to deliver high quality video. When his vision of 40 million nodes comes to fruition, he may have access to 40 PB of storage without shelling out a single penny to any storage vendor.
The above reasons may explain what encouraged me to initiate contact with GridNetworks. One day, I expect to see the same technology being implemented with in the enterprise for distributed storage aggregation.
Excerpts from PowerGrid EULA:
"Segment" means a small block of encrypted Content data, typically under one megabyte.
Encrypted, managed peer-to-peer "grid" network architecture.
The Content Segments stored on your computer may be portions of files you have viewed, or may be copied to your computer by GridNetworks host systems for later sharing with other GridNetworks authorized users.
The Segments contained therein are only accessible through the GridNetworks PowerGrid application and cannot be individually viewed, altered or deleted through your computer's file management tools.
These Content Segments are periodically removed and replaced with other segments and shall never exceed the maximum capacity of the Data Store, established by you.
Monday, December 11, 2006
GridNetworks, what’s my Interest? Part one.
As I mentioned in my post, Challenges of High Quality Video Delivery, last week, I spent an hour with Newell Edmond, Co-founder and Amy, Project Manager talking about GridNetworks, its technology and business.
Contrary to what few readers thought and queried, neither I met them for a job nor they rented me to blog about them. I discovered GridNetworks through John Cook’s blog post Getting 'Goodfellas' on the grid. It piqued my interest for several reasons.
First, I have been thinking of ways to highlight startups working on interesting infrastructure solutions. Consumer Internet and Web 2.0 startups are topic of discussion by bunch of bloggers. But the infrastructure startups that enable them are largely ignored.
There was nothing better than a local infrastructure startup to initiate my coverage. Expose your infrastructure startup, you know whom to contact and how!
Second, even though my initial impression was that it is yet another startup trying online video delivery. Further research showed their aspirations are more substantial than just becoming a video destination.
And after meeting with Newell, I am convinced that they are actually a video delivery infrastructure play, combining content delivery network and peer to peer network technologies, than just a torrent player. And this is also confirmed by Michael Gersh, a VP at Reeltime, a GridNetworks customer, in his email.
Contrary to what few readers thought and queried, neither I met them for a job nor they rented me to blog about them. I discovered GridNetworks through John Cook’s blog post Getting 'Goodfellas' on the grid. It piqued my interest for several reasons.
First, I have been thinking of ways to highlight startups working on interesting infrastructure solutions. Consumer Internet and Web 2.0 startups are topic of discussion by bunch of bloggers. But the infrastructure startups that enable them are largely ignored.
There was nothing better than a local infrastructure startup to initiate my coverage. Expose your infrastructure startup, you know whom to contact and how!
Second, even though my initial impression was that it is yet another startup trying online video delivery. Further research showed their aspirations are more substantial than just becoming a video destination.
And after meeting with Newell, I am convinced that they are actually a video delivery infrastructure play, combining content delivery network and peer to peer network technologies, than just a torrent player. And this is also confirmed by Michael Gersh, a VP at Reeltime, a GridNetworks customer, in his email.
We are currently using Grid's tools as part of our end-to-end solution that is streaming DVD quality video over the web, to customers all over the world.Second part coming soon.
Wednesday, December 06, 2006
Challenges of High Quality Video Delivery
Since Google acquired YouTube, online video sharing and delivery segment has been hot topic of discussion. A good overview of online video viewing is presented by Scott Kirsner in As online viewing booms, the amateurs give way to big media. As more and more big media content coming online, a new technology challenge is emerging for content distributors.
The file size for 90 minute full length HD quality video can range from 2 to 6 GB. And it may be easier, convenient and hassle-free for viewer to get it delivered via next day Sneakernet than to wait 3 to 18 hours in downloading it online. There is nothing instant about downloading full-length high quality video.
Streaming
The next best online alternative is to stream video to viewers in real-time. For smooth playback, video streaming needs to have some buffering and a 12 second buffering of HD quality video requires a 5 to 15 MB of video that need to be always available to user. And, if streaming video is originating from a centralized distribution infrastructure, the number of viewers are first limited by the processing capacity of central system and then available bandwidth.
Content Delivery Network (CDN)
To address the central system processing capacity limitation, the option is to have multiple delivery nodes. And, to address bandwidth limitation, these content delivery nodes should be located as close to the viewer as possible. This is the model used by CDN providers who share the processing capacity of these nodes and available bandwidth among multiple content distributors to maximize utilization of their edge nodes.
Content delivery nodes work great as most viewers reside on the last mile that extends beyond the Internet spiderweb and as long as nodes are not overloaded by too many viewers trying to download and watch the latest video simultaneously.
Peer-to-peer (P2P) Networking
P2P networking is one scalable alternative proposed to deliver high quality video and bring the viewers in to the Internet spiderweb. Among it's other vices, from delivery perspective, P2P is infested with freeriders. You know the ones who want to download content but don't want to allow their system to be used to deliver content to others. And, the whole P2P premise fails if there are not enough P2P nodes participating in content distribution.
GridNetworks – New Kid on the Block
Today, I had the pleasure of meeting and talking with Newell Edmond, co-founder and the technical brain behind a local startup in online delivery of high quality video, GridNetworks. Newell and his team worked hard for last two years to develop a hybrid solution combining the best of Content Delivery Network (CDN) and Peer-to-Peer Networking (P2P) for delivery of high quality video.
More about it next time.
P.S. This is my first effort to reach out and highlight early stage startups in Seattle area working on innovative infrastructure solutions. If you are one of them, get in touch.
How to deliver a high quality full-length video instantly to multiple viewers on their big screen simultaneously and securely?Sneakernet
The file size for 90 minute full length HD quality video can range from 2 to 6 GB. And it may be easier, convenient and hassle-free for viewer to get it delivered via next day Sneakernet than to wait 3 to 18 hours in downloading it online. There is nothing instant about downloading full-length high quality video.
Streaming
The next best online alternative is to stream video to viewers in real-time. For smooth playback, video streaming needs to have some buffering and a 12 second buffering of HD quality video requires a 5 to 15 MB of video that need to be always available to user. And, if streaming video is originating from a centralized distribution infrastructure, the number of viewers are first limited by the processing capacity of central system and then available bandwidth.
Content Delivery Network (CDN)
To address the central system processing capacity limitation, the option is to have multiple delivery nodes. And, to address bandwidth limitation, these content delivery nodes should be located as close to the viewer as possible. This is the model used by CDN providers who share the processing capacity of these nodes and available bandwidth among multiple content distributors to maximize utilization of their edge nodes.
Content delivery nodes work great as most viewers reside on the last mile that extends beyond the Internet spiderweb and as long as nodes are not overloaded by too many viewers trying to download and watch the latest video simultaneously.
Peer-to-peer (P2P) Networking
P2P networking is one scalable alternative proposed to deliver high quality video and bring the viewers in to the Internet spiderweb. Among it's other vices, from delivery perspective, P2P is infested with freeriders. You know the ones who want to download content but don't want to allow their system to be used to deliver content to others. And, the whole P2P premise fails if there are not enough P2P nodes participating in content distribution.
GridNetworks – New Kid on the Block
Today, I had the pleasure of meeting and talking with Newell Edmond, co-founder and the technical brain behind a local startup in online delivery of high quality video, GridNetworks. Newell and his team worked hard for last two years to develop a hybrid solution combining the best of Content Delivery Network (CDN) and Peer-to-Peer Networking (P2P) for delivery of high quality video.
More about it next time.
P.S. This is my first effort to reach out and highlight early stage startups in Seattle area working on innovative infrastructure solutions. If you are one of them, get in touch.
Monday, August 28, 2006
Power - The Theme of Data Center Outages
As soon as I finished my last rant (See, Infrastructure Failure - Be Paranoid), a reader pointed out a similar incident close to home, at Fisher Plaza in Seattle.
Fisher Plaza is considered to be a premium data center housing numerous high profile clients and ten different telecom carriers bringing fiber to the buildings. I visit the facility time-to-time as several customers are co-located there.
It seems Fisher Plaza experienced an outage due to electrical power equipment failure few weeks ago knocking KOMO TV and KOMO 1000 News Radio stations offline (See, Unsinkable Data Center Crashes in Seattle). There were also several similar incidents reported in other cities (See, Data Center Outages Bring Headaches, Headlines and InterNAPPing?)
The moral of these incidents for customers are very simple:
Fisher Plaza is considered to be a premium data center housing numerous high profile clients and ten different telecom carriers bringing fiber to the buildings. I visit the facility time-to-time as several customers are co-located there.
It seems Fisher Plaza experienced an outage due to electrical power equipment failure few weeks ago knocking KOMO TV and KOMO 1000 News Radio stations offline (See, Unsinkable Data Center Crashes in Seattle). There were also several similar incidents reported in other cities (See, Data Center Outages Bring Headaches, Headlines and InterNAPPing?)
The moral of these incidents for customers are very simple:
- There is no 100% uptime SLA. Anyone who promises you this is smoking something they shouldn't be. In a AFCOM survey, 20% had been hit with at least five failures in past five years (See, Five Predictions: 9 of 10 Companies Face Failures)
- When planning mission critical operations, look beyond just redundant power supplies, UPS and WAN connections. Facility is as important as your equipment in keeping your operations up. "Many data centers just can't handle new technologies coming out," Comment by a presenter at AFCOM Data Center World Conference (See, Five Predictions: Relocations and Outsourcing)
- Perform due diligence not only on tenant happiness and satisfaction but also crisis handling and management.
- Murphy's Law is alive and kicking even for well planned, thought out, mission critical activities.
Wednesday, August 16, 2006
Infrastructure Failure – Be Paranoid
Just after writing my previous blog post Data Center Power Consumption and Heat Generation, I came across an interesting blog post with 400+ comments. It details the failure of hosting infrastructure at DreamHost (See, Anatomy of a(n ongoing) Disaster..) In my opinion, it is a recommended reading for everyone who manages or designs IT infrastructure for living.
Here are excerpts from the post with my commentary and takeaways.
No incumbent makes competitor entry easy and painless. Get specific deliverables when you have the leverage. Be prepared to work alone and have a Plan B considering total non-cooperation from vendor.
Here are excerpts from the post with my commentary and takeaways.
Ironically, all the recent disasters stem somewhat from us attempting to take some proactive steps to head off any sort of future power outages like the kind we experienced last year.Instead of narrowly focus on preventing something from happening again, use the event as wake up call. Assess your environment for other potential risks and develop comprehensive plan to address them. Also, be aware of new problems that may arise while solving another problem. I like to use the Chess analogy - Further you anticipate moves, better your chances of prevailing.
We're now basically 95% of their data center.Consider how important you are to your vendor and leverage your position to negotiate better deal.
The Garland Building is supposed to be an excellent place for data centers. There are more than a dozen in the building. Companies like iPowerWeb, Media Temple, BroadSpire, and even MySpace (now the most popular website in the whole US!) are in there.Do your own due diligence even if you think vendor passed the due diligence by a larger well known company. Their failure may be critical to your business and a drop in the bucket for these "other" companies. It is not uncommon for vendors to offer sweet deals to attract high profile companies.
Around last June though, the building informed all its data center tenants that they had essentially run out of power!Don't wait for other shoe to fall before taking actions. Be paranoid.
After months of searching and negotiating with Alchemy, we still had to get Switch and Data to allow us to put a cross-connect in from their data center over to their competitors down the hall.Finding an alternate data center down the hall may seem quick and easy fix to existing power problem. But such short-sighted and point solutions fail to address other lingering issues such as Disaster Recovery / Business Continuity. I guess the company plan is to wait until problem occurs before addressing them.
No incumbent makes competitor entry easy and painless. Get specific deliverables when you have the leverage. Be prepared to work alone and have a Plan B considering total non-cooperation from vendor.
Wednesday, August 09, 2006
Data Center Energy Consumption and Heat Generation
This morning, I read interesting news about LBNL, Sun and others demonstrating DC power distribution in data center as a way to reduce power consumption and heat generation (See, Engineers: DC Power Saves Data Center Dough). Six months ago, I did come across few data center people interested in ways to reduce power and space requirements and decrease heat generation (See, Happy New Year & Food for Your Brain).
It looks like technology industry has been making some efforts in this area. DC power sounds an interesting alternative. I wonder what are the drawbacks of using 380-volt DC power instead of AC. The eWEEK article focuses on the good side only. I am sure there is bad side to it too. Aren't there reasons for preference of AC over DC in power distribution? Can't recall exact details but something to do with distribution losses.
On facility wide basis, 15% reduction in energy consumption does sound significant. I am looking forward to reading the complete report when and if available. I expect that efforts focusing on reducing energy consumption at rack level may produce more significant results.
This area is expected to be pretty Hot "literally" in near future considering projections indicating cost of powering up and cooling down exceeding cost of equipment.
Last year, Luiz Andre Barraso of Google also published an interesting article on the same topic in ACM Queue (See The Price of Performance) where he described cost trends of large IT infrastructure such as Google's with couple of interesting graphs. Some of the key points mentioned are:
It looks like technology industry has been making some efforts in this area. DC power sounds an interesting alternative. I wonder what are the drawbacks of using 380-volt DC power instead of AC. The eWEEK article focuses on the good side only. I am sure there is bad side to it too. Aren't there reasons for preference of AC over DC in power distribution? Can't recall exact details but something to do with distribution losses.
On facility wide basis, 15% reduction in energy consumption does sound significant. I am looking forward to reading the complete report when and if available. I expect that efforts focusing on reducing energy consumption at rack level may produce more significant results.
This area is expected to be pretty Hot "literally" in near future considering projections indicating cost of powering up and cooling down exceeding cost of equipment.
Last year, Luiz Andre Barraso of Google also published an interesting article on the same topic in ACM Queue (See The Price of Performance) where he described cost trends of large IT infrastructure such as Google's with couple of interesting graphs. Some of the key points mentioned are:
- Every gain in performance has been accompanied by a proportional inflation in overall platform power consumption. The result of these trends is that power related costs are an increasing fraction of the TCO.
- The energy costs of that system today would already be more than 40 percent of the hardware costs. (The system is a x86 server worth $3,000 consuming 200 watts on average).
- If performance per watt is to remain constant over the next few years, power costs could easily overtake hardware costs, possibly by a large margin.
Sunday, April 16, 2006
Gatekeepers Babysitters Datacenters
For the past few weeks, I wasn't able to post regularly due to a very hectic schedule. So I am trying to write a series of quick short catchup posts summarizing the things I have been working on, thinking, reading and researching about. No indepth analysis, here. This is the first in this series:
In last couple of weeks, I visited about a dozen datacenters. What I noticed most is the inefficiencies, in the name of security, at datacenters for both services provider personnel and datacenter employees! During each visit, I spent 15 to 30 minutes going through the gatekeeper's entry procedures and datacenter employees spent two to four hours babysitting me while I was inside the datacenter. That is some serious productivity loss for both service providers and datacenter owners.
Another thing, I observed is the lack of security cameras inside the datacenters. Security cameras were everywhere outside and inside the buildings except in datacenters!
Aren't there better ways to handle gatekeeping and babysitting work at datacenters?
In last couple of weeks, I visited about a dozen datacenters. What I noticed most is the inefficiencies, in the name of security, at datacenters for both services provider personnel and datacenter employees! During each visit, I spent 15 to 30 minutes going through the gatekeeper's entry procedures and datacenter employees spent two to four hours babysitting me while I was inside the datacenter. That is some serious productivity loss for both service providers and datacenter owners.
Another thing, I observed is the lack of security cameras inside the datacenters. Security cameras were everywhere outside and inside the buildings except in datacenters!
Aren't there better ways to handle gatekeeping and babysitting work at datacenters?
Subscribe to:
Posts (Atom)