Happy New Year! Here is an article from the Freakonomics blog tracking American’s (un)happiness.
Cheers
Happy New Year! Here is an article from the Freakonomics blog tracking American’s (un)happiness.
Cheers
So the Jets blew it. They were 8-3 and they finished the season 1-4 to miss the playoffs at 9-7.
I was reading this article on the Freakonomics blog.
Attempts/completions: 175/98
Passing yards: 1,011
Touchdowns: 2
Interceptions: 9
Sacks: 9
Passer rating: 55.4
This is terrible. Make the guy retire. He WAS great. WAS.
So anyway, the Jets probably should have made the play-offs, but they just lost too many games. The Patriots, on the other hand, probably should have just made the play-offs. I know there are rules and criteria, but do the Cardinals (who got massacred by the patriots) or the Chargers (who started the season 4-8) really deserve to be in the play-offs over the Patriots?
Probably not, but the rules are what they are. There just aren’t any good teams (even mediocre teams) in either the AFC or NFC west this year.
This got me thinking. Don’t they always talk about parity in the NFL? How many commentators ever week do you hear throwing around parity this and parity that. Well where was the parity this year? Is there less parity in the league now than in past years? How can we measure parity?
Let’s start with what a good measure of parity would be. If there were perfect parity in the league every team would finish 8-8. This would be like flipping a coin to decide every game. (Not exactly the most compelling sports league.) The opposite (teams deviate from a record or 8-8), a lack of parity, can be thought of as entropy. Most notably over the last two season we have had a 16 win team and a 0 win team. Clearly, these teams were significantly better and worse, respectively, than the other teams in the league. So how can we measure parity. Lets put all 32 teams into a 32 by 2 contingency table. 32 rows, one for reach team and 2 columns, on for wins and one for losses. (This leads to fixed row and fixed column totals.)
We wish to test the null hypothesis that there is no association between the rows and columns. (ie the team you play for has nothing to do with the number of wins you get). Clearly this is not true and we will always reject the null hypothesis of no association, but we can compare by how much we reject the null hypothesis, namely the p-value. The smaller the p-value, the more parity in the league. While this is not a perfect measure, its and interesting start.
Since we have fixed row and column totals, normally this would lead to using Fisher’s exact test. However, with a 32 by 2 table this is computationally very intensive. Thus as an alternative we can use the Pearson Chi-Square test statistics and the G-squared test statistics. Here I report both statistics. I am going to rely more heavily on the G-squared statistics because of its close relationship with entropy.
95.62900 1.640269e-08 2008 32 7.95362
97.11005 9.714734e-09 2007 32 8.138756
69.63735 8.525854e-05 2006 32 4.704668
94.40966 2.518507e-08 2005 32 7.801207
79.79948 3.535543e-06 2004 32 5.974935
76.59059 9.912913e-06 2003 32 5.573824
56.98146 3.010793e-03 2002 32 3.122683
86.44988 2.244115e-07 2001 31 7.042142
79.80431 2.107774e-06 2000 31 6.198153
71.86353 2.721014e-05 1999 31 5.189674
(Does anyone know how to make nice tables in wordpress blogs?)
This jumbled mess summarizes the results of the G-squared test. Columns one and two are the G-squared test statistic and the respective p-value and column three is the year in which this regular season took place. (Note that the 2009 Super Bowl would correspond to the 2008 regular season.) Column four is the number of teams in the league that year and the last column is the number of standard deviations above the mean that the test statistic is. (p-value would be the best way to compare seasons, but the scale of p-value is difficult to visualize, so the graphs use standard deviation. Also the test statistics cannot be directly compared to one another because they have distributions with differing degrees of freedom.)
Over the past ten years, using the p-values of the G-squared statistic the years with the most entropy were 2005, 2007, and 2008. The years with the most parity were 2002, 1999, and 2006.
Time for some pictures.
The first graph is a plot of NFL season versus what I am calling entropy (Number of standard deviations above the mean of the distribution of the test statistic.) I have also labeled each year with the super bowl champion and the number of wins they had in the regular season. Notice that over the last four years we observe the three highest amounts of entropy.

Further note 2002. Lets take a closer look at this year. There were no teams with 13 wins and every team had at least two wins. This is compared to 2007 when there were FOUR teams with at least THIRTEEN wins (Dallas, Green Bay, Indianapolis, and undefeated New England) and one team (Miami) team with only one win. The histograms of wins in the 2002 and 2007 seasons are below.


Look how tightly bunched the 2002 teams are in the middle and compare that with the 2007 season.
One last picture. The histogram of the 2008 NFL season.

Conslusions:
According to my measure of entropy, the level of parity in the NFL over three of the last 4 seasons has been very low. There does not, however, appear to be any upward trend in the amount of parity; rather, it seems as if the level of parity in the league varies trendlessly from year to year.
2002 was a season with unusually high parity with many teams finishing with similar records.
One final thought: I would argue that parity is bad for the league. When there is no standout team, it is difficult to market exciting games. Isn’t it more compelling to watch a play-off game featuring teams who absolutely dominated the regular season (Think Green Bay, Dallas, New England, and Indianapolis from 2007), than a slug fest of mediocrity between two teams that made it into the play-offs by default (I’m looking at you Arizona and San Diego). When parity is high, everyone is mediocre, but someone has to win by default. When there is high entropy good teams exist. I’ll take the latter any day of the week.
Within the past few years I’ve started golfing fairly regulary. Last year, a few friends of mine and myself, started tracking our progress on Oobgolf. We enter out scores, and they automatically track our handicaps. (I’m a 20.4 by the way).
Anyway, at the end of the season, we had a big single elimination tournament with our handicaps. At some point during the tournament we got to talking about how two people could have the same handicap and be entirely different players.
Here is an extreme example:
Handicap is calculated using the best 10 scores from your last 20 rounds. Golfer one could play 20 rounds and shoot 90 in all of them. Golfer two could shoot 90 ten times and then 110 the other ten times. Both these golfers would have the same handicap, but if you were going to play for money (or in our big end of the season tournament) you’d rather play golfer two, even thought golfer one and golfer two have the same handicap.
We had a brief discussion about how you could quantify this disparity. Apparently, however, some other people have thought a lot harder about this.
This article in Chance from 2001 discusses how a “Steady Eddie” has an advantage over a “Wild Willie”. (The chance article uses order statistics so you might want to check out this link for a brief description.)
Then this article proposes a measure called Anti-Handicap which mesaures your worst ten scores in your last twenty rounds. By comparing a golfers handicap and anti-handicap, some measure of variability in a golfers game can be assessed. As before in our extreme example, gofler one would have a handicap of 18 and an anti-handicap of 18. However, golfer two would have a handicap of 18, but an anti-handicap of 38.
Cheers.
Finals are over and I hope to post more regularly again. Here is a quick picture.

This picture, by French engineer Charles Joseph Menard, graphically depicts Napolean’s fateful march to Russia. The width of the line represents how many troops Napolean had at each point on his way to Russia and what makes this graphic so great is just how many different variables are all displayed at once.
Edward Tufte says in his book, The Visual Display of Quantitative Information, “Minard’s graphic tells a rich, coherent storhy with its multivariate data, far more enlightening than just a single number bouncing along over time. Six variables are plotted: the size of the army, its location on a two-dimensional surface, direction of the army’s movement, and temperature on various dates during the retreat from Moscow”.
In the last line of the description below the graph, Edward Tufte says, “It may well be the best statistical graphic ever drawn,” which, in my opinion, may be the best claim ever made about a statistical graphic ever.
Cheers.
For Stanley:
Complete caption of the graphic:
“This classic of Charles Joseph Minard (1781-1870), the French engineer, shows the terrible fate of Napolean’s army in Russia. Described by E. J. Marey as seeming to defy the pen of the historian by its brutal eloquence, this combination of data map and time-series, drawn in 1861, portrays the devastating losses suffered in Napolean’s Russian campaign of 1812. Beginning at the left on the Polish-Russian border near the Niemen River, the thick band shows the size of the army (422,000 men) as it invaded Russia in June 1812. The width of the band indicates the size of the army at each place on the map. In September, the army reached Moscow, which was by then sacked and deserted, with 100,000 men. The path of Napolean’s retreat from Moscow is depicted by the darker, lower band, which is linked to a temperature scale and dates at the bottom of the chart. It was a bitterly cold winter, and many froze on the march out of Russia. As the graphic shows, the crossing of the Berezina River was a disaster, and the army finally struggled back into Poland with only 10,000 men remaining. Also shown are the movements of auxiliary troops, as they sought to protect the right flank of the advancing army. Minard’s graphic tells a rich, coherent story with its multivariate data, far more enlightening than just a single number bouncing along over time. Six variables are plotted: the size of the army, its location on a two-dimensional surface, direction of the army’s movement, and temperature on various dates during the retreat from Moscow. It may well be the best statistical graphic ever drawn.”
Here is a good ESPN article from 2002 about a baseball statistic called runs per game developed by Harvard professor of statistics Carl Morris.
Cheers.
Cassell to Moss. TOUCHDOWN. With only seconds left the Patriots had completed a drive started deep in their territory to pull within one point of the Jets. They kicked the extra point, went to overtime, and lost. (After having the Jets at 3 and 17.) If the Patriots had gone for two after the touchdown, they could have won the game right there. So should Belicheck have gone for two?
Endgame Technologies has developed a simulator for football games called ZEUS. According to these simulations Belicheck should have gone for two at the end of regulation instead of kicking the extra point and sending the game to overtime.
I’ve been interested in decision making in football for a long time, especially the decision to go for two points after a touch down instead of kicking the extra point. The article “Refining the Point(s)-after-touchdown decision” by Harold Sacrowitz is an excellent article on the subject. In his results he develops a table for when to go for two or kick the extra point in order to maximize a teams chances of winning.
More recently, an article “Do firms maximize? Evidence from professional football” by David Romer, investigates NFL teams decision about going for it on fourth down. He argues that NFL teams are kicking (punting and going to field goals) too often and they would increase their probabilty of winning by going for it on fourth down more often.
And here is a guest post by Ian Ayres on the Freakonimcs blog asking the question “Why don’t sports teams use randomization?”
Cheers.
I like stats. I also like baseball. So what could I love more than baseball statistics.
href=”http://en.wikipedia.org/wiki/Bill_James”>Bill James. He is widely considered to be the father of baseball statistics, or sabermetrics.
Bill James came up with a whole bunch of very clever ways to analyze baseball using statistics, including runs created, range factor, and win shares.
Another one of his stats is Pythagorean expectation:

This statistics works very well for predicting wins, but it doesn’t really make an sense. Why does it work? I was wondering how this statistics would compare with a multiple regression based on the same data.
Using the 2008 MLB baseball data of wins, runs scored, and opponent runs scored by team. I compared the predictions for expected wins made by Pythagorean expectation versus a simple multiple regression model of the form Wins=Runs scored+Opponent runs scored+error. The root mean squared error for the pythagorean expectation was 7.37. The root mean squared error for prediction for the regression model was 4.18. While Pythagorean expectation does a very good job predicting win percentage, a multiple linear regression does a better job. Also, the regression model has coefficients that can be interpreted practically while Pythagorean expectation works well, but offers very little reasoning as to why it works well.
The model was: predicted wins=79.7416+.1025*Runs scored-.1088*Opponent runs scored. This model has R-squared=.8528. So on the average, approximately every extra ten runs a team scores is worth a win and every extra ten runs given up is equal to a loss.
What happens if you build a regression model with more than just runs and opponent runs as predictor variables. Using the same 2008 data, a model for predicting wins in 2008 is:
Predicted wins=60.75-71.15*WHIP+.11768*SB+271.133*OBP+.10984*HR
It seems that you can break down winning baseball games into four factors:
1.) Pitching
2.) Speed
3.) Contact hitting
4.) Power hitting
I realize that’s not a shocking revelation, but it’s neat to see it even with this small data set.
I was a little bit surprised to see that SB shows up because a common theory is that stealing bases is not worth the risk, but it shows up very strongly in this model.
So I looked it up:
Top 5 teams in SB for 2008
1.) Tampa Bay Rays
2.) Colorado Rockies
3.) New York Mets
4.) Philadelphia Phillies
5.) Los Angeles Angels
Bottom 5 in SB in 2008
30.) San Diego Padres
29.) Pittsburgh Pirates
28.) Arizona Diamondbacks
27.) Atlanta Braves
26.) Detroit Tigers
3 of the top 5 and 5 of the top 7 teams made the playoffs. Interesting.
I’ll end with a quote from the greatest base stealer of all time: “It took a long time, huh? [Pause for cheers] First of all, I would like to thank God for giving me the opportunity. I want to thank the Haas family, the Oakland organization, the city of Oakland, and all you beautiful fans for supporting me. [Pause for cheers] Most of all, I’d like to thank my mom, my friends, and loved ones for their support. I want to give my appreciation to Tom Trebelhorn and the late Billy Martin. Billy Martin was a great manager. He was a great friend to me. I love you, Billy. I wish you were here. [Pause for cheers] Lou Brock was the symbol of great base stealing. But today, I’m the greatest of all time. Thank you.
—Rickey Henderson’s full speech after breaking Lou Brock’s record
Cheers
Voter turnout for this last election was as high as it has been in the last 60 years. Click here for a good graphical display of voter turnout since 1948 on Andrew Gelman’s blog.
Also check out the United States Election Project where they have data about past election going back about 50 years.
A friend of mine told me about this problem, so I went and looked it up. This is stats in the wildest of the wild.
So, during World War II, the Allies were trying to estimate the number of a certain kind of German tank. They needed this information to better plan their attacks and invasions. There were two sets of estimates made, one by intelligence and another by a group using statistical methods.
Estimates made using statistical methods in June 1940, June 1941 and August 1942 of the number of a certain type German tank were, respectively, 169, 244, and 327. The intelligence estimates for each of those same three periods of time were, respectively, 1000, 1550, and 1550. (from :Number of German tanks)
These estimates are drastically different, and depending on which estimate was believed, it is possible that battle plans may have been significantly affected. So who made the better estimates?
In most situations when we estimate something, we can never actual know what the true value is. However, as it turns out, after the war was over, German records became available and the actual number of tanks that they had at each of those three points in time became available. The actual number of tanks that the Germans had at the three points in time (June 1940, June 1941 and August 1942) when the estimates were made were, respectively, 122, 271, and 342. (Recall that the statistical estimates were 169, 244, and 327 for those three time periods.) The statistical estimates are astonishingly close. (As well as the intelligence estimates being alarmingly inaccurate.) So how did they do it?
The statistical group looked at the serial numbers of tanks that had been captured or destroyed by Allied troops, and they assumed that the serial numbers of the tanks were ordered from 1 to T where T is the number of tanks that the Germans had. So they assumed that if the Allies found a tank with serial number 200, that the Germans had at least (and almost surely more than) 200 tanks.
So if we assume that each serial number has equal probability of being observed our maximum likelihood estimate (our best guess) of T is simply the maximum serial number that we encounter on a destroyed tank. However, using the maximum encountered serial number to estimate T turns out to be an unbiased estimator. (If we always used the largest serial number as our estimate of T, we would be systematically underestimating T, because our largest observation is usually not the actual largest value.) So what we need is an unbiased estimator for T.
As it turns out the expected value of our estimator of T (the maximum observed serial number) is n/(n+1)*T (hence biased). So on the average the largest observed value will be smaller than actual T. To correct for this we simply multiply the largest observed value by (n+1)/n. This will give us an unbiased estimate for the number of tanks the Germans had, and this is how they reached their statistical estimates.
Example:
Say we observe 50 tank serial numbers and the largest observed serial number is 245. With all of the above assumptions, our unbiased estimate as to the number of tanks is 51/50*245=249.9.
If we observe 25 serial number and the largest is 110, our best guess is 26/25*110=114.4.
Here is a link to another blog post about the German tank problem.
Modern note: I saw online that someone was using this approach to try to estimate the number of servers that Google has. (More to come on that)
References:
Ruggles, R., and Brodie H. (1947) An empirical intelligence in World War 2. Journal of the American Statistical Association, 42:72-91.
Goodman, L. A. (1954), “Some Practical Techniques in Serial Number Analysis,”
Journal of the American Statistical Association, 49, 97–112.
Accoriding to this article from www.readwriteweb.com, they claim that “Errors By Bloggers Kill Credibility & Traffic, Study Finds”. Interesting. So how did they reach this conclusion?
From the article:
“The company [goosegrade.com] asked a demographically diverse group of respondents on Amazon’s Mechanical Turk website to fill out the survey and published the results today on the goosegrade.com company blog. The bulk of respondents spent some time reading blogs but were people who remained dependent on ‘mainstream sources’ for most of their news.”
(For an explanation of Mechanical Turk, the Wikipedia article is here.
Comment: How does goosegrade know these people were demographically diverse? The only people they asked were Mechanical Turk workers. That seems like a very specific group of people. So you should only be able to make inference about that group of people. They hardly speak for internet users in general, but goosegrade.com uses them to make inference about “internet users” when they should just be making inference about “mechanical turk workers who are being paid by gooseGrade.com”. Those two groups are drastically different.
gooseGrade.com says on their site (http://www.goosegrade.com/reader-perception-survey-results)
“Readers want gooseGrade. Here’s proof.
175 People polled.
ABSTRACT: It appears that grammar, spelling, factual, and other errors do affect reader opinion as well as how likely they are to share or link to an article. These errors also seem to dictate the readers opinion of the author’s skills as a writer. 65.86% of internet users say that a tool like gooseGrade would increase their confidence in the content they are reading. Filtering further shows that 9 out of 10 newspaper readers say that a tool like gooseGrade would increase their confidence in author’s content. This merrits further investigation of newspaper readers and could show a path for new media to take more market share.”
As I said before, I’m not sure the opinions of 175 (more on this below) mechanical turk workers are sufficient to make inference on all internet users. Furthermore, remember that all of these respondents were paid by goosegrade.com (although it was probably only a few cents.)
A note on their sample size: They claim a sample size of 175 internet users, but an examination of the raw data shows that there are only 161 unique IP address. 9 IP addresses are repeated twice and 1 IP address is repeated 5 times. These should be thrown out of the sample because it is likely that they are the same person.
The readwriteweb.com article concludes with:
“Below are a few of the charts, you can see the rest on the GooseGrade blog. The lesson here? It seems pretty clear. We bloggers are harming our own credibility and traffic with our inattention to details, not just in the facts, but in the basics of our writing. Let’s do better!”
Here is a promise I am willing to make. I’ll write better and make less grammatical errors if you apply statistics more fairer. (LOL)
Cheers.