Your verification ID is: guDlT7MCuIOFFHSbB3jPFN5QLaQ Big Computing: open source software
Showing posts with label open source software. Show all posts
Showing posts with label open source software. Show all posts

Tuesday, June 7, 2011

Looking for a free statisitical software?

I have been following a Linkedin discussion for a couple of weeks on what open source analytical software to use for a company with little money. I put up a recommendation for R, but it was really was wild to see the range of suggestions. One was just to use SAS because it really was not that expensive when you considered what you got. Others recommended R, Rattle, Knime, and a whole bunch that I have never even heard of.  Overall a very entertaining discussion, but I an not sure it provided any real benefit for the person who posed the question. I will be the first to admit I have an R basis so my knee jerk reaction to the question was to reply R without fully understanding this guys needs. Therefore, another solution might be superior to R in his particular case.

One person responded with a link to The Impoverished Social Scientist's Guide to Free Statisticial Software and Resources by Professor Micab Altman of Harvard. First what a great title! Second what a fine resource. Yes it is a little dated with a last update in 2008, but I think it is still pretty on target even three years later. So if you are an Impoverished Social Scientist take a look. If you are simply a person wondering what open source tools are available to address you needs this is a good place to start.

Friday, May 6, 2011

Heritage Health Prize goes against Open Source

Today on KDnuggets I read the Heritage Health Prize recently modified the License agreement to make the work product the sole property of Heritage Health. I think this is wrong. If you want to develop a proprietary algorithm go hire someone to do it, but to claim all the work product submitted in the competition even the ones that do not win and therefore are not paid for is just wrong.

Heritage Health can not have their cake and eat it too. Kaggle has been very clear that their site has been the develop cheaper, faster analytic tools for its customers ( the contest sponsors) at a lower cost than they could do otherwise. That is fine and the contest sponsors should use and implement the models submitted to the contest. However, what we have seen is a collaborative approach wins these competitions, and a sharing of how they did win with the larger community sometimes on the Kaggle site itself makes future models even better. If predictive analytics is going to makes the leaps forward that it really needs to do it can only happen in a open collaborative environment which not only encourages but demands the sharing of information, algorithms and approaches. If we do not, analytics will cease to progress at the rate that it has been in recent history, and we will return to the bad old days of investment companies jealously guarding their superior infinite random walks from the other investments houses.

It is no coincidence that predictive analytics took off with the advent of open source software. The R environment is a shining example of that which also wins most of the Kaggle contests. It is better than what came before and will continue to improve because of the collaborative contributions of its dedicated users.

Tristan has called for a boycott of this contest. The thread bring out some other outlandish and real issues of concern.