Your verification ID is: guDlT7MCuIOFFHSbB3jPFN5QLaQ Big Computing: MLB
Showing posts with label MLB. Show all posts
Showing posts with label MLB. Show all posts

Tuesday, August 16, 2011

The Log Cabin goes all SABR in Colorado

One of the blogs I read all the time is The Log Cabin by Ryan Elmore. I particularly like his sports stuff because he looks at questions I had never even thought of. He recently did a talk at the Rocky Mtn SABR Meeting which features sports, R and ggplot2. That is about my vision of a perfect night so I am posting the link below. Enjoy


Slides from Rocky Mtn SABR Meeting

Friday, July 22, 2011

Friday Musing - Hot temperatures results in ejections in MLB, but what happens to the game?

In a previous post I wrote about a research paper that found a relationship between hot temperatures and batters getting hit by pitches. Yesterday I came across an old article that finding that baseballs fly farther in hot temperatures ( about 2%). Finally this morning I saw a post that Managers are more likely to get ejected when the temperature rises. All this is very interesting to read when the East Coast is suffering from an oppressive heat wave. In fact I am trying to think of a way to get myself ejected to somewhere cooler than here right now. However, it got me to wonder if the games themselves where different from games played in less draining conditions.

So I looked at total runs scored in the American League East during the hottest month of the year which is July in 2011 compared to the rest of the year. Total runs scored in July was 9.7 versus 8.9 the rest of the year. This heat wave may be a factor. It absolutely a factor in why I have been watching the games on TV with the AC blasting instead of going to the ballpark.

Tuesday, July 5, 2011

Does Attendence effect Winning in MLB

I started messing with this idea over the weekend because I like to watch TV and play with the computer at the same time. I was looking at various stats on MLB on the ESPN website, and I ended up looking at the various teams home field attendance numbers. There really are few surprises in the top ten, half are American League and half are National League with all except the Twins, Cubs and Dodgers having good seasons.

The bottom ten was more of a surprise to me. with 70% of those teams with poor attendance from the American League and Division Leader Cleveland Indians right there at 26th place. The Indians draw poorly both home and anyway.

This got me thinking if there is an impact of the mythical 11th man (or 10th man in the National League). So I thought I would look at if variations in attendance impacts the Indians performance on the field. The spreadsheet below shows the results for the first half of the 2011 season. It is interesting to note that the Indians have a much better record when attendance is less than 10,000 (6-1) or more than 30,000 (7-1) than they do when attendance is between 10,000-20,000 (10-7) or 20,000-30,000 (4-2). Correlation is not causality, but at least for this year the Indians do better when the 11th man shows or stays home, and worse when he kinda shows.


Attendance Runs scored Runs Allowed Run Diff
8726 7 1 6
9025 3 1 2
9076 8 2 6
9523 8 4 4
9650 9 4 5
9722 7 2 5
9853 3 8 -5
10594 1 0 1
10714 8 3 5
13017 4 2 2
13551 5 4 1
14164 5 4 1
15224 7 8 -1
15278 4 6 -2
15336 2 11 -9
15397 4 7 -3
15498 1 0 1
15568 9 5 4
15849 2 3 -1
15877 3 4 -1
16336 2 8 -6
16346 8 2 6
17568 4 3 1
18107 4 7 -3
19225 3 2 1
20261 0 2 -2
23752 2 4 -2
26408 2 14 -12
26433 3 2 1
26833 12 4 8
27458 0 4 -4
30023 5 2 3
31622 5 4 1
31865 5 1 4
33774 5 4 1
38549 5 1 4
40631 2 1 1
40676 6 3 3
41721 10 15 -5 

Tuesday, June 21, 2011

The Red Sox season is a case of Dr. Jekyll and Mr Hyde

In the first 36 games of the season the Red Sox went 17-19 while scoring 4.22 and allowing 4.47 runs per game. In the following 35 games the Red Sox went 26-9 while scoring 6.54 and allowing 3.91 runs per game. A 50% increase in runs scored over the same time period is usual.

Even more strange is if you look at the first 35 games the Red Sox never scored more than nine runs in a game. In the following 36 games they have scored more than nine runs 8 times. If I plot the run total frequency of the first 35 games of the season versus the following 36 games you get two totally different distributions.

The Standard Deviation for the first 35 games was a relatively tight 2.72 while for the following 36 games the Standard Deviation has ballooned 4.31. I have never seen a so flat a distribution of runs as the Red Sox have had in the last 36 games. Most of the runs scored distributions I have seen look like Chi Squared distributions

Wednesday, June 1, 2011

A look at Batting orders

There is one Blog I read on sports statistics religiously and the is Phil Birnbaum's Sabermetric Research. It is a great read, and he looks at many aspects of lots of different sports as opposed to just baseball. If you have not looked at his stuff before check it out.

One of his recent postings dealt with a paper written by Nobuyoshi Hirostu who looked at if using expected runs was always the best way to determine the batting order or could a lineup with a lower expected runs produce more wins because of lower volatility. Nobuyoshi used a cut down version of the game to calculate the expected runs and ran a MC calculation to determine the winners of each potential matchup. For this experiment he used the 2007 season.

Out of the 600,000 potential matchup guess how many instances he found where the lineup with the lower expected runs won more than 50% of the games? 13! I was surprised there were not many more than that. I expected there would be a fair number of lineups of high batting average singles hitters that might have a lower expected number of runs but wins against a lineup of power hitters who score more runs on average but have great volatility due to lower batting averages.

Based on Nobuyoshi's approach to this problem I think the results are surprising, but correct. However, I can see some potential problems with how he constructed his model for analysis. First, by building a cut down model for expected runs he may have reduced the volatility of the various lineups and made the winning potential for a lower expected run lineup less likely. Second, the lineups were based on the player makeup of the various teams. For whatever reason, ( in baseball I usually assume tradition) most MLB have a lineup consisting of Power Hitters and Reliable Hitters. I think a very interesting question to ask is if this type of lineup is optimal. What type of lineup gets the highest expected wins, and does it do it with highest expected runs or some balance between high expected runs and lower volatility? 

Monday, May 23, 2011

Sabermetrics Seminar

I went to the Sabermetrics Seminar at Harvard this weekend. It was a charity event, and all the speakers came and talked on their own dime. I just want to thank those speakers for giving up their time for such a great cause.

The Seminar itself was an eye opening experience for me. The last seminar I went to was the R/Finance in Chicago. That Seminar, like most that I go to, is for hard core statisticians and computer scientists. I believe of the hundreds of attendees to R/finance I am one of the few without a PhD.  The presentations with the possible few exceptions of JD Long's honoring of Dr Suess were of a highly technicial level. The Sabermatrics Seminar was totally different. The audience varied from the Head of the Harvard Statistics Department and an eminent physicist to people with very limited mathamatical education. The presentations also ran the gambit from something that would be taught in a high school physics class to some fairly high level stuff. The great unifier in the room was these people loved baseball and where using mathamatics to expand their understanding of the game and increase their enjoyment. One Speaker, Dan Duquette, former GM of the Boston Red Sox, reminded us of the words of Flippe Alou to "remember to enjoy the game". Tom Tippet, Director of Baseball Information Systems, gave a great Q&A on the state of Sabermetrics in MLB today. I have included a link to a summary of the seminar here.

Sabermetrics is different than the other fields I work in. In Pharma, the models are widely shared, but the data is highly confidential. In Finance the models are confidential, but the data is basically public. MLB analysts seem to strongly guard both their models and there data viewing both as propietary. While I think this makes it a great opportunity for consulting, I believe it may hinder the rate of refinement. Kaggle has shown in a very public way that open collaboration on data and models yields astounding improvements in prediction.

Friday, May 20, 2011

It is a Sabermetrics Weekend so todays post is Sabermetrics

This weekend I am going to the Sabermetrics Seminar in Boston. Some might think that it is strange that I am excited about this given that I never played baseball, and I do not watch many games. However, the analytics being done is baseball is developing and expanding at such a rapid pace there is no way you can enjoy analytics and not be interested. Recently a friend of mine ran into Prof. Bertsimas and asked if he could have a copy of the now famous paper that he wrote predicting the Red Sox would win 100 games this year. Prof Bertsimas asked "are you a fan of baseball?" to which he responded "No, I am a fan of statistics". 

The development of Sabermetrics in the last 30 years has been to look at existing data and try to build predictive models out of that data. It was a good first step and produced some good results. This work revealed that some of the historical statistics, like ERA, were not good predictors of anything so Sabermetricians created statistics that were better predictors. This is all great, and it has taken Sabermetrics to where it is today.

The problem with the data that has been used today in baseball is that it is all result based data. The pitcher threw a strike or a ball, the batter got on base, etc. That is all changing. Welcome to the world of physical data in Baseball. This post on Beyond the Boxscore is a good example. It has taken the improvement of a players performance back to the physical location of his pitch not just that more of his pitches resulted in ground balls, but an attempt to answer why based on data not opinion. The technology exists not only to track data of a baseball as it crosses the plate but within the entire ballpark. First this is going to create an unbelievable amount of data that needs to be in studied in ways not currently used in baseball because of shear volume. Second this data is collected in real time which means the models could be updated in real time. Billy Bean may have had his 3X5 note card in front him, but the manager of the future may be holding his iPad with feedback on up to the last pitch and the suggested options with predicted results of those options.

One of the companies doing this physical data collection in baseball is Trackman. They also recently posted for an R developer. I can not wait to see what is coming!