For almost a month, blogging and blog reading took a back seat to hosting friends and family from four different countries, playing golf and Wii (addictive!), rigors of the day job and enjoying the summertime.
I most probably will still be on blogging hiatus for few more weeks if not for the battle of the words among storage bloggers Robin Harris, Barry Burke, and others. I guess I got embroiled by making an observational comment on Robin's Blog post resulting in pen-lashing from Storagezilla.
Overall, I am very surprised with the emotional responses and the offense both bloggers and commentators are taking. Be nice, boys! There is some great feedback in these blog posts and comments for everyone to take home. Whether, like it or not, just saying thanks for the feedback, will go a long way.
After storage blogging for almost four years, I would be saying thanks for the feedback, if I was receiving such thoughts about my blog from someone. May be, I am not cut out for EMC DNA mutation. BTW, thanks Chuck for acknowledging the storage trenchtrolls.
This war of words also reminded me of an incident, I was involved in, about two years ago. At that time, Hu was a class act in responding to my criticism of his blogging. In retrospect, I was out-of-line. Despite this incident, I was invited and hosted at HDS Executive Briefing Center by Jeremiah, previously HDS community evangelist. For this very reason, I admire and have great respect for both Hu and Jeremiah.
My offer to storage bloggers, readers, industry insiders and outsiders remains same. If you ever visit Seattle area, get in touch. I will love to pick your brain over a beer/coffee/lunch/dinner. I will do the same when I am in your town, just send me your contact info.
Of course, any feedback through comments, email and phone is always welcome too.
Friday, August 17, 2007
Tuesday, July 24, 2007
Tales of Storage IPOs: What happened to Quality?
A comparison of upcoming IPOs, Netezza, BladeLogic, Voltaire and Compellent, with recent storage IPOs. If VMware is to be shown on this chart, it will be outside of upper-right quadrant.
Monday, July 16, 2007
Power Consumption by Google Services
Even though, Google doesn’t share a lot of details of their infrastructure, as we have seen from limited published information, they are obsessed with continuously monitoring, managing and improving the efficiency of their infrastructure.
Recently, Robin Harris attended the Google conference on scalability and then mused How Yahoo can beat Google. Few months ago, Google published results of their work on disk drive failure in paper Failure Trends in a Large Disk Drive Population [PDF]. It was extensively covered in blogosphere including by me in blog entries SMART not so smart in predicting disk drive failure and Google Findings of Disk Failures Rates and Implications and by Robin Harris in his blog entry Google’s Disk Failure Experience.
Google has done it again and presented results of their work on power consumption and provisioning in paper Power Provisioning for a Warehouse-sized Computer [PDF] at the ACM International Symposium on Computer Architecture, San Diego CA, June 9 – 13, 2007. In this work, Google researchers, Xiaobo Fan, Wolf-Dietrich Weber and Luiz Andre Barroso looked in to 15,000 servers running three different applications – Websearch, Webmail and Mapreduce for six months to determine the power usage characteristics at Rack, PDU and Datacenter levels.
Google Services
Websearch: A service with high request throughput and large data processing for each request.
Webmail: A disk I/O intensive service. Machines configured with large number of disk drives. Each request involves a relatively small number of servers.
Mapreduce: A cluster dedicated to running large offline batch jobs. Involve process terabytes of data using thousands of machines.
Key Findings
The key findings from this work are:
More details from this study later.
Recently, Robin Harris attended the Google conference on scalability and then mused How Yahoo can beat Google. Few months ago, Google published results of their work on disk drive failure in paper Failure Trends in a Large Disk Drive Population [PDF]. It was extensively covered in blogosphere including by me in blog entries SMART not so smart in predicting disk drive failure and Google Findings of Disk Failures Rates and Implications and by Robin Harris in his blog entry Google’s Disk Failure Experience.
Google has done it again and presented results of their work on power consumption and provisioning in paper Power Provisioning for a Warehouse-sized Computer [PDF] at the ACM International Symposium on Computer Architecture, San Diego CA, June 9 – 13, 2007. In this work, Google researchers, Xiaobo Fan, Wolf-Dietrich Weber and Luiz Andre Barroso looked in to 15,000 servers running three different applications – Websearch, Webmail and Mapreduce for six months to determine the power usage characteristics at Rack, PDU and Datacenter levels.
Google Services
Websearch: A service with high request throughput and large data processing for each request.
Webmail: A disk I/O intensive service. Machines configured with large number of disk drives. Each request involves a relatively small number of servers.
Mapreduce: A cluster dedicated to running large offline batch jobs. Involve process terabytes of data using thousands of machines.
Key Findings
The key findings from this work are:
- The difference between maximum power used by large number of computing devices, cumulatively, and their theoretical peak usage can be as much as 40% in datacenters.
- It may be more efficient to leverage power management techniques at datacenter level than at rack level.
- Nameplate ratings are of little use in power provisioning as they significantly overestimate actual maximum usage.
- CPU utilization as a measure of machine-level activity produces accurate results for dynamic power usage especially with large group of machines. The dynamic power range is less than 30% for disks and negligible for motherboards.
- Using maximum power draw of individual machines to provision the datacenter, will have some stranded capacity.
- A mix of diverse workload reduces the difference between average and peak power, an argument in favor of mixed deployment.
- Idle power is significantly lower than the actual peak power, but generally never below 50%.
- CPU dynamic voltage/frequency scaling may yield moderate energy savings (up to 23%) at datacenter levels.
- Peak power consumption at the data center level could be reduced by 30% and energy usage could be halved if systems were designed so that lower activity levels meant correspondingly lower power usage profiles.
More details from this study later.
Subscribe to:
Posts (Atom)