Your verification ID is: guDlT7MCuIOFFHSbB3jPFN5QLaQ Big Computing: EC2
Showing posts with label EC2. Show all posts
Showing posts with label EC2. Show all posts

Tuesday, June 21, 2011

What is the Cloud?

At Sifma last week I had a meeting with a major cloud vendor who felt the biggest issue they were dealing with was a lack of understanding of what the cloud actually was. I was surprised by this, but I had seen a similar problems with the open source project Hadoop when it became a popular term, but few people actually understood what it was.

Yesterday I ran across a thread on Linkedin that asked how would a person describe the cloud to their co-workers. The answers where interesting and diverse. I will share a few here:

"Cloud = Commoditization of IT" 

" cloud, in essence is a Fabric which at its core supports the application stack "

"For the end user, I would use the example of google documents, google calendar, gmail. Most people are familiar with Google and you can demo it as well. All the user need is internet access and a browser and they can basically access their "desktop" from anywhere with any internet connected hardware. "

"It is the internet"

Let not forget the Wikipedia definition of the cloud 

When I think of the cloud I do not really consider those things like Salefoorce.com or Quickbooks.com delivering enterprise applications through a simple Website. I really limit my thinking to the ability to  access remote computer time to run applications. A concept that really started with Amazon's EC2 in 2006.  This was a great idea! Amazon had built massive infrastructure to handle the huge computing volumes of their business. However, they noticed that their business volumes were seasonal and a lot of their computers remained idle for much of the year. EC2 allowed them to rent out that capacity, and create another revenue stream for themselves.  It was a true win/win. It was such a great idea that others vendors like Rackspace, Google and Mircosoft with Azure have entered the business. 

The basic idea is sound and companies can save significant money by outsourcing their peak computer usage rather than maintaining the internal infrastructure to support that need. Because of economies of scale of these large cloud suppliers some companies may save money by outsourcing all of their hardware needs.

However, I do see a problem in recent history and on the horizon. Amazon correctly named their service EC2 ( elastic cloud). It was excess capacity that companies could use. What happens if cloud usage no longer is elastic?  In January 2011 Netflix launched on amazon's EC2. Their volume has grown to 20% of total internet usage at any given time. This load along with the overall increase in usage of EC2 has resulted in problems including service EC2 being interrupted. I believe these sorts of periodic volume constraints will continue and increase in cloud computing. In the long term I believe it will be addressed just like it has in the past with a priority system and many levels of service by the cloud provider, but in the cloud provider's case backed by a pricing model.  

Monday, May 2, 2011

What did I learn for R/Finance 2011

The R/finance 2011 meeting was a huge success! All the talks were just great. I do not have the time to go through each talk one by one but I do feel there were a couple of themes that ran through the entire conference. The opening speaker, Mebane Faber, and the keynote speaker,  John Bollinger, touched on two topics near to my heart. The first is that in many cases the simplified model does nearly as well as the more complex one and in some case with fewer pitfalls. The second is that models are our attempt to describe reality, but they are not reality. Therefore there is always the possibility that the model is a bad fit for the reality that it is trying to model or there exists a deviation from the  model to the reality it is describing. Both phenomenons can be exploited for advantage. Never get blindly enamored with a model and approach things with an opening mind. These ideas carried pretty consistently throughout the conference.

Parallel or High Performance Computing for R are becoming a more and more important factor in analytic computing. I am not sure if it is because to the continue growth of data in general, the enterance of HPC into general awareness through the "cloud",  or because the really cool problems seems to exist on the edge of our current capability. I believe with the exposure of more users to HPC tools for R it is time to update the various pros and cons of each approach and to benchmark them against each other with a set of set typical data set and models. I do wonder if the recent problems on Amazons EC2 could will slow down the growth of cloud computing? Lost time is one issue here but the users that lost their data could be much more reluctant to take that risk in the future.

I was also amazed at the traction that Rstudio had among this group of experienced R users. I have always held the belief that experienced users of any software package shy away for IDEs and GUIs and prefer the simple interaction of command line coding. I felt IDE were the tool for new or mid-level users. In this case, I was wrong. Rstudio appears to provide benefit to the very experienced R user to the point they are willing to change away from what they are currently doing and learn this model tool.

I thought JD Long's Dr Seuss inspired talk was the most entertaining of the confernece. It takes some talent to do that and even more to do it well. His Segue for R package is pretty cool too. Flash talks are a great format, and I wish they were used more often