When I was a sophomore in college I could not find an apartment so a friend of mine let me stay in his apartment for most of the first term. The same guy helped me write a forecasting program that was my senior project two years later. He was a great guy who was always there. His name is Todd Pack. When I meet him the only thing he wanted to do was build robots. He spent much of his free time either writing code on his computer ( named "Bear") or building robots in the lab. It is what he loved to do.
Todd got his PhD out of Vanderbilt with a thesis on IMA, and took a job at iRobot. Yup, they make the Roomba vacuum cleaner. It is my understanding the reason the Roomba was created was because in the early 1990s there were no government contracts for robots so they needed a consumer product to survive.
I haven't seen or heard from Todd in a number of years. Then about a month a ago I went to a talk by Yann LeCun on vision and learning algorithms for robots. The videos of the robots reminded me of the work Todd Pack did in the labs back at Cornell all those years ago. Yesterday I read an article saying that the US was sending Robots to Japan to help out at the Fukaushima Nuclear Power Plant (article). It turns out the robots come from Dr Pack's company iRobot, and I am sure he had a hand in their design. I am so happy that Todd has created a career doing what he loves, and his work absolutely saves lives. I have my guess why they call them Packbots, but it is only a guess.
Now if we could only make robots that do not look like Johnny number 5.
Here is a recent video on MSNBC featuring the iRobot's Packbots.
I blog about world of Data Science with Visualization, Big Data, Analytics, Sabermetrics, Predictive HealthCare, Quant Finance, and Marketing Analytics using the R language.
Showing posts with label predictive. Show all posts
Showing posts with label predictive. Show all posts
Wednesday, April 20, 2011
Friday, April 15, 2011
Would we have found out San Diego Basketball through Analytics
This week the San Diego State Basketball program had a number of its players arrested for shaving points in a game in February of 2010 and trying to do the same thing in a game against UCR in February of 2011 (News Story). It appears that these guys were caught by human intelligence. This type of cheating is bad for both the sports organizations (NCAA, NFL, NBA, MLB) and the major betting community. It is bad for the sports organizations because who, except WWE fans, are going to watch fixed match. It is bad for the betting community because they really make their money on the Juice they charge gamblers. The betting community wants fair games that split the money evenly over the spread. Anything that shifts that is a problem for their business model. People believing that the games are fixed could reduce the amount of money bet on games. Also bad for the Bookies.
In the book Freakomonics by Levitt and Dubner, they expose match fixing in sumo matches using statistical analysis. It was a fun read and showed that analytics have the ability to expose cheating in sports without human intelligence. It was also a safe sport to look at because Americans do not really care about sumo nor do they bet on it.
It is interesting to me that there exists so much data and analysis with a goal of prediction on sport, but I have found nothing on using predictive analytics to discover point shaving or game fixing. I realize that doing this kind of work in team sports would be more completed than something like sumo. I just feel that it might be another tool to add to the effort to deter this kind of problem. At first pass it seems the most likely times there is potential cheating is when a players statistics in a game are an outlyer, and the team did not beat the spread. The problem that I see is the sparsity of known point shaving in games. For example I only know of one alleged fixed game in the NCAA basketball season in 2010 out of something like 5,000 games. I do not think that is enough to be useful. Sad to say that if there were more fixed games we might be able to build a better model. Someone suggest to me that I would get better data on game fixing if I looked at Italian football. However, I call it soccer and care Italian football about as much as I do about sumo.
So I have no data to present or model to put forth. I just hate cheaters, and this story has bothered me since it came out on April 11th.
In the book Freakomonics by Levitt and Dubner, they expose match fixing in sumo matches using statistical analysis. It was a fun read and showed that analytics have the ability to expose cheating in sports without human intelligence. It was also a safe sport to look at because Americans do not really care about sumo nor do they bet on it.
It is interesting to me that there exists so much data and analysis with a goal of prediction on sport, but I have found nothing on using predictive analytics to discover point shaving or game fixing. I realize that doing this kind of work in team sports would be more completed than something like sumo. I just feel that it might be another tool to add to the effort to deter this kind of problem. At first pass it seems the most likely times there is potential cheating is when a players statistics in a game are an outlyer, and the team did not beat the spread. The problem that I see is the sparsity of known point shaving in games. For example I only know of one alleged fixed game in the NCAA basketball season in 2010 out of something like 5,000 games. I do not think that is enough to be useful. Sad to say that if there were more fixed games we might be able to build a better model. Someone suggest to me that I would get better data on game fixing if I looked at Italian football. However, I call it soccer and care Italian football about as much as I do about sumo.
So I have no data to present or model to put forth. I just hate cheaters, and this story has bothered me since it came out on April 11th.
Labels:
analytics,
basketball,
fixed,
Freakomonics,
point shaving,
predictive,
San diego,
sumo,
UCR
Tuesday, April 12, 2011
Analytics, Sabermetrics, Data Mining...Why can't we all just get along?
Sabermetrics was a term coined by Bill James to describe the analysis of baseball through objective evidence. Saber, or more accurately SABR, stands for the Society for American Baseball Research. With Bill James as its advocate. Sabermetrics has changed the way baseball is played. No easy task in a sport so encumbered by tradition. Baseball probably collects more data during a game than any other sport and each team plays at least 162 games a year. Rich data territory compared to the 16 regular season games played in the NFL. Sabermetrics has taken a hard look at the core beliefs of what statistics make a good baseball player or team and runs them against the cold judgement of analytics. The results showed that some previously treasured statistics like batting average were not as important statistics as once thought, but others like on base percentage were better indicators. This is predictive analytics at it best. So it is time to call Sabermatrics what it is analytics.
It is funny for all the impact Sabermetrics has had on baseball I believe it is still limited by the traditions of baseball. Let me give you some examples.
The Blog Sabermetic Research talks about Buck Showalter changing the way his base runners play to gain 5 runs per year which he claims is worth $10 million dollars. Makes sense if the data he is using is good, but the key here is the decision is claimed to be made solely on the numbers.
Pitching is another story. In baseball a starting pitcher must pitch five full innings in order to earn a decision (win/loss). Many talk about the difference between ERAs of starting versus relief pitchers. The data clearly shows that relief pitchers, even when they are the same person, have an overall ERA .50 lower than starting pitcher or better. Tango on Baseball touches on the subject in this article. My question is that if relief pitchers have a better ERA than stating pitchers, and starters are generally accepted to be better pitchers than relievers why aren't starters being used like relievers? The impact would be huge! A quick pass says this .50 ERA reduction in starting pitchers would result in 40 less runs allowed by a team over the course of a season! Using Showalter math that is $80 million dollars. I believe the reason that this is not looked at as a solution is because of tradition. If starting pitchers where used like relievers they would never pitcher 5 innings, and therefore would never get a decision. This would be a fundamental change in the way baseball is played.
In defense of Sabermetricians, there has been some discussion that ERA, like BA, is not a very useful statistic. This would mean that conclusions drawn from those statistics may not be as useful as they appear. I have not seen anything on starters versus relievers in terms of CERA, dERA, DICE or DIPS.
It is funny for all the impact Sabermetrics has had on baseball I believe it is still limited by the traditions of baseball. Let me give you some examples.
The Blog Sabermetic Research talks about Buck Showalter changing the way his base runners play to gain 5 runs per year which he claims is worth $10 million dollars. Makes sense if the data he is using is good, but the key here is the decision is claimed to be made solely on the numbers.
Pitching is another story. In baseball a starting pitcher must pitch five full innings in order to earn a decision (win/loss). Many talk about the difference between ERAs of starting versus relief pitchers. The data clearly shows that relief pitchers, even when they are the same person, have an overall ERA .50 lower than starting pitcher or better. Tango on Baseball touches on the subject in this article. My question is that if relief pitchers have a better ERA than stating pitchers, and starters are generally accepted to be better pitchers than relievers why aren't starters being used like relievers? The impact would be huge! A quick pass says this .50 ERA reduction in starting pitchers would result in 40 less runs allowed by a team over the course of a season! Using Showalter math that is $80 million dollars. I believe the reason that this is not looked at as a solution is because of tradition. If starting pitchers where used like relievers they would never pitcher 5 innings, and therefore would never get a decision. This would be a fundamental change in the way baseball is played.
In defense of Sabermetricians, there has been some discussion that ERA, like BA, is not a very useful statistic. This would mean that conclusions drawn from those statistics may not be as useful as they appear. I have not seen anything on starters versus relievers in terms of CERA, dERA, DICE or DIPS.
Monday, April 11, 2011
The Most Boring day in the Last hundred Years
I was driving in to work this morning and listening to the radio because I feel that distractions make me a better driver. A news article comes on telling me that a Cambridge researcher has found out that April 11, 1954 is the most boring day in the last 100+ years. I had a good laugh thinking that it isn't only the US government that gives out silly grants for pointless research ( remember the which came first the chicken or the egg paper).
So I dig a little deeper. Turns out this is a software guy who just launched his "smart" search engine. I could not find any specifics on how the engine works which would have been cool, but then I thought this is maybe even cooler than another algorithm. Here is a guy who got this new web site rolled out as a news story in three major newspapers and NPR just by asking an interesting question about the most boring day. It would be really amazing if the search engine gave him the question after he queried "the most likely answer to a question to be picked up by major news organizations".
The Most Boring Day Article
True Knowledge
Tunstall-Pedoe and his search engine
So I dig a little deeper. Turns out this is a software guy who just launched his "smart" search engine. I could not find any specifics on how the engine works which would have been cool, but then I thought this is maybe even cooler than another algorithm. Here is a guy who got this new web site rolled out as a news story in three major newspapers and NPR just by asking an interesting question about the most boring day. It would be really amazing if the search engine gave him the question after he queried "the most likely answer to a question to be picked up by major news organizations".
The Most Boring Day Article
True Knowledge
Tunstall-Pedoe and his search engine
Friday, March 18, 2011
How does beer relate to Scrapping Data?
Sometimes I find myself involved in projects as a result of having just one too many rounds of beer after a meetup. Such is the case with the beer predictor. In an overt attempt to ingratiate ourselves with the local craft brewers we thought it would be a good idea to write some analytics on craft beers to get noticed by those brewers so that they would in turn supply us with free beer. At least we had a reasonable goal in mind. I will update that project as we go along. The first part of the project was to scrap the data from the Beer Advocate. I had not started doing this because it was St Paddy's day, and I felt the proper way to study beer on that day is consumption. Others wrote code, and finished that portion of the project. I did come across this blog post on scrapping data in R with XML versus in Python with Beautiful Soup which I thought was interesting.
http://thelogcabin.wordpress.com/2010/08/31/using-xml-package-vs-beautifulsoup/
http://thelogcabin.wordpress.com/2010/08/31/using-xml-package-vs-beautifulsoup/
Labels:
analyticds,
beautifulsoup,
beer,
datamining,
predictive,
Python,
r,
scrapping,
XML
Subscribe to:
Posts (Atom)