• Playoff Picture:

    AFC
    1. NewEngland
    2. Pittsburgh
    3. Indianapolis
    4. Kansas City
    5. NYJets
    6. Baltimore

    NFC
    1. Atlanta
    2. Philadelphia
    3. GreenBay
    4. Seattle
    5. NewOrleans
    6. TampaBay

    Estimated Probabilities of making the playoffs/conference champion/super bowl champion:

    AFC.Playoff.Teams

    Baltimore Cleveland Denver Houston Indianapolis

    0.9706 0.0002 0.0046 0.0024 0.6040

    Jacksonville KansasCity Miami NewEngland NYJets

    0.3584 0.6874 0.0282 0.9950 0.9840

    Oakland Pittsburgh SanDiego Tennessee
    0.0830 0.9568 0.2334 0.0920

    NFC.Playoff.Teams

    Arizona Atlanta Chicago GreenBay NewOrleans

    0.1446 0.9996 0.5944 0.8642 0.7922

    NYGiants Philadelphia SanFrancisco Seattle StLouis

    0.1460 0.9740 0.0332 0.6892 0.1338

    TampaBay Washington
    0.5752 0.0536

    AFC.Champion

    Baltimore Indianapolis Jacksonville KansasCity Miami

    0.1668 0.0182 0.0046 0.0054 0.0004

    NewEngland NYJets Pittsburgh SanDiego Tennessee

    0.3490 0.2510 0.2018 0.0012 0.0016

    NFC.Champion

    Arizona Atlanta Chicago GreenBay NewOrleans

    0.0002 0.5206 0.0246 0.0926 0.0704

    NYGiants Philadelphia SanFrancisco Seattle StLouis

    0.0042 0.2354 0.0002 0.0062 0.0004
    TampaBay Washington
    0.0434 0.0018
     

    SB.Champion

    Atlanta Baltimore Chicago GreenBay Indianapolis

    0.2730 0.0854 0.0046 0.0308 0.0048

    Jacksonville KansasCity Miami NewEngland NewOrleans

    0.0010 0.0006 0.0004 0.2220 0.0214

    NYGiants NYJets Philadelphia Pittsburgh SanDiego
    0.0012 0.1452 0.0832 0.1116 0.0002

    Seattle TampaBay Tennessee Washington
    0.0014 0.0124 0.0006 0.0002

    Cheers.

  • Here are the StatsInTheWild rankings of the NFL teams after Week 10.
    Rank Teams Wins
    1 NewEngland 7
    2 Atlanta 7
    3 NYJets 7
    4 Pittsburgh 6
    5 Baltimore 6
    6 Miami 5
    7 Philadelphia 6
    8 NewOrleans 6
    9 TampaBay 6
    10 GreenBay 6
    11 Indianapolis 6
    12 Tennessee 5
    13 Cleveland 3
    14 NYGiants 6
    15 Chicago 6
    16 Jacksonville 5
    17 KansasCity 5
    18 Oakland 5
    19 Seattle 5
    20 SanDiego 4
    21 Houston 4
    22 Washington 4
    23 Denver 3
    24 Minnesota 3
    25 Cincinnati 2
    26 Arizona 3
    27 StLouis 4
    28 SanFrancisco 3
    29 Dallas 2
    30 Detroit 2
    31 Buffalo 1
    32 Carolina 1

    The two teams that pop out are the Giants and the Browns. The Browns are 3-6 with wins over the Bengals, Saints, and Patriots. Their losses are to Tampa Bay, Kansas City, Baltimore, Atlanta, Pittsburgh, and the Jets. Those teams average 6.33 wins each and all hold at least a share of the lead in their respective division.

    The Giants have losses to Indianapolis, Tennessee, and Dallas. They have beaten Carolina, Chicago, Houston, Detroit, Dallas, and Seattle. The six teams they have beaten average only 3.33 wins. A vast difference in schedule.

    I also wrote some code to simulate the rest of the NFL season based on what has already happened. So I used the games that have already occurred as data for a model (a very simply model). This model predicts the probability of a win for a given team. Then the rest of the season is simulated 5000 times based on the estimated probabilities of winning a game.
    Here are the results of that:

    Probability that a team wins its division:
    Division Winners:
    AFCEast
    Miami NewEngland NYJets
    0.0080 0.6004 0.3916

    AFC North
    Baltimore Cleveland Pittsburgh
    0.4712 0.0004 0.5284

    AFC South
    Houston Indianapolis Jacksonville Tennessee
    0.0020 0.6158 0.1058 0.2764

    AFC West
    Denver KansasCity Oakland SanDiego
    0.0450 0.5646 0.2138 0.1766

    NFC East
    NYGiants Philadelphia Washington
    0.1850 0.8118 0.0032

    NFC North
    Chicago GreenBay Minnesota
    0.2066 0.7886 0.0048

    NFC South
    Atlanta NewOrleans TampaBay
    0.9150 0.0432 0.0418

    NFC West
    Arizona SanFrancisco Seattle StLouis
    0.1726 0.0374 0.6958 0.0942

    Conference Champions:
    AFC
    Baltimore Cleveland Indianapolis Jacksonville KansasCity Miami
    0.1524 0.0002 0.0378 0.0010 0.0008 0.0078
    NewEngland NYJets Oakland Pittsburgh Tennessee
    0.3418 0.1844 0.0138 0.2586 0.0014

    NFC
    Atlanta Chicago GreenBay NewOrleans NYGiants Philadelphia
    0.5882 0.0280 0.0484 0.0556 0.0084 0.2106
    Seattle TampaBay
    0.0156 0.0452

    Super Bowl Champion:

    Atlanta Baltimore Chicago GreenBay Indianapolis KansasCity
    0.3112 0.0862 0.0040 0.0126 0.0094 0.0002
    Miami NewEngland NewOrleans NYGiants NYJets Oakland
    0.0032 0.2154 0.0154 0.0016 0.1068 0.0022
    Philadelphia Pittsburgh Seattle TampaBay Tennessee
    0.0640 0.1522 0.0016 0.0134 0.0006

    So, right now SITW is predicting a Atlanta Falcons – New England Patriots Superbowl with Atlanta winning.

    Cheers.

  • So I’m in a poker league. We play three out of four Thursday nights in a month for a total of fifteen events in a season. Each event you earn a certain number of points. You’re 10 best finishes based on points out of the fifteen events are counted. After the grueling fifteen week season, there is a finale where you start with an amount of chips proportional to the number of points you earned during the season. Anyway, another player (Shaun) and I took over the scoring for the league this season. Our scoring system is pretty basic (50 points for showing up, 50 points for everyone you beat, 600-300-150-75 bonus for cashing (finishing top 4)). On top of this I’ve devised a fairly reasonable ranking system separate from the points based on your average finish and how many events you have played. One criticism of my ranking system is that I’m not accounting for the strength of the field, I’m just looking at average percentile of finish.

    So I usually car pool to and from events with Shaun and we’ve been talking about ranking systems. Last night he mentioned that he had been doing some research on how Xbox live does their rankings. Each player has some level of ability and an uncertainty associated with their skill level. After each game players rankings, skill level and uncertainty, are updated. He described it a little bit more, and I mentioned that it sounded Bayesian to me. Turns out it is!

    Here is an introduction to the TureSkill ranking system and here is a more detailed description. For those of you who are interested in all the details, here is the paper “TrueSkill(TM): A Bayesian Skill Rating System” where they propose the system.

    Cheers.

    P.S. A big SITW congratulations to James T. O’Connor of Belchertown, MA for passing the CT bar exam.

    P.P.S. I finished second in the regular season last year, but won the finale. This year I was briefly in first place until last night. I am now in second place by 150 points with 4 events to play.

  • This was forwarded to me by S.J. It’s quite a long journey from the stuff in this video to some of the present day statistics like spatial aggregate fielding evaluation (S.A.F.E).

    “[Recorded: circa 1959] โ€“ โ€œThe Electronic Coachโ€ is a short film made by IBM describing the use of computers in the management of a university basketball team. The film features computer science legend Don Knuth, then a senior at Case Institute of Technology. For all four of his undergraduate years at Case (1956-60), Knuth was manager of the basketball team and sought ways to improve his teamโ€™s play by analyzing a series of special statistics he captured during games. The scoring method was unusual in the weightings it gave to activities not necessarily associated with traditional coaching but Knuthโ€™s insights into basketball, combined with his computerization of the reams of data he collected, helped Caseโ€™s coaching staff make their basketball team a winner. The computer used is an IBM 650. ”

    Cheers.

  • I’m teaching my first class this fall, and I’ve been preparing my notes for class this past week. I wanted to use keno as an example of how to compute probabilities. So I was computing some probabilities and checking them against the posted “odds” on masslottery.com. I couldn’t get my computed odds to match with what the lottery had posted, which led to a brief period of panic that I wasn’t qualified to be teaching this class. Turns out, I’m not computing anything wrong. It’s just that what the lottery is calling “odds” are actually probabilities. Take a look again at masslottery.com and look at the posted odds for a one spot game. For the one spot game they say that the odds are 1:4. This is incorrect. The probability of winning this one spot game is \frac{1}{4}=.25, which would make the odds of winning \frac{.25}{.75}=1:3. Likewise, the odds against winning are 3:1. Generally, if the probability of an event is p, the odds of this even occuring are \frac{p}{(1-p)}.

    So what the lottery is referring to as odds are actually probabilities of winning. They actually get this correct that the bottom where they say “Probability of winning a prize in this game = 1:4.00”. The mistake is that they aren’t making any distinction between the probability of winning and the odds of winning when, in fact, these are different.

    Cheers.

  • I’ve been in Vancouver the past week for the Joint Statistical Meetings (JSM) . Here is a collection of my thoughts and comments from the few days I was at the conference.

    On Monday I went to the section on Survey Research Methods and saw Meena Khare, of the National Center for Health Statistics (NCHS) and Laura Zayatz of the United States Census give talks. They both spoke about measures that their institutions go through to release data to the public. The NCHS looks for uncommon combinations of variables that could be used to possible de-identify the data. Both organizations first remove obvious identifiers and then go on to make the released data more private. For instance, if the number of observations with a unique combination of variables in a data set is n and the number of observations with the unique combination is N, they would consider any combination of variables where n/N<.33 at risk for disclosure. At the U.S. Census, they have used something called data swapping to protect public release data sets in the last two Censuses (2000 and 2010). Along with data swapping, the Census will also be using partially synthetic data to maintain confidentiality to protect individual privacy in public release data in 2010.

    Several things strike me about this.
    -The methods that these organizations use to protect confidentiality are certainly going to increase privacy compared to a release of raw data, however, there doesn’t seem to be any way to know that what is being done is providing “enough” privacy.
    -It’s clear that many different government organizations have issues which require some use of disclosure limiting techniques, however, it seems that each organization is creating its own rules and there is limited discussion going on between organizations to conceive of a standard policy for data sharing.
    -There doesn’t actually seem to be any definition of what is considered a disclosure. For instance, if government data is released and I discover through some technique that someone definitely has HIV, then clearly a disclosure has taken place. However, if I use the same data to discover that someone definitely does not have HIV, a disclosure has still taken place, but the consequence is much less damaging. Furthermore, consider a situation where prior the the data release, I know a particular individual has a 50 percent chance of having HIV. After the data release, I can infer that there is a 99 percent chance that they have HIV. Clearly, I would consider this a disclosure. But what if the probabilities shift from only 50 percent (pre-data release) to 75 percent (post-data release) or 50 percent to 55 percent. At what point is “too much” information being released. It seems as if this issue receives less attention than is warranted.
    -Finally, I believe that the ultimate solution to the disclosure problem is a careful combination of policy and disclosure limiting techniques. Policy issues include defining how much privacy must be maintained by a given technique, as well as, legal consequences for knowingly disclosing private information. Statistics has an obligation to provide increasingly improving statistical disclosure techniques along with metrics for measuring the privacy of a given technique.

    Later on Monday, I saw the tail end of the talk by David Purdy titled “Statisticians: 3, Computer Scientists: 35”. The abstract for the talk was:
    “John Tukey and Leo Breiman warned us that a day would come when statistics would need to focus more on computing, or risk losing good students to computer science. The Netflix Prize provides many examples of how our field needs to do more.

    In the top 2 teams, participants with a computer science background vastly outnumbered those with a statistics background. There are a number of lessons that the field of statistics can learn from the fact that undergraduates in CS were well equipped to compete, while statisticians at all levels were not well prepared to implement advanced algorithms.

    In this talk, I will address methodological issues arising with such a large, sparse dataset, how it demands serious computational talents, and where there is ample room for statistics to make contributions.”

    I only saw the end of the talk, but I feel like I got the point. He notes how programs in statistics need to expose students to more aspects of computing. One quote from his talk that particularly shocked me was from a prominent statistician referring to the Netflix prize data set: (I’m paraphrasing) “I can’t do anything with the data, there is just too much of it.” (If anyone knows the actual quote, I would love to have it). Too little data may often be a problem, but too much data should be a blessing, rather than a curse.

    When he is talking about computing he is referring to implementing complex algorithms to analyze the data, however, in my experience I have seen people struggle with simply managing data of this size. This is a simple problem to deal with, but, in my experience, I have both had this happen to myself and seen it happen to others. When I was in grad school working towards my master’s degree right out of undergraduate, we were given a problem in a consulting class with a “large” (several thousand observations) amount of data (well, “large” to someone with no experience managing data.) We (my group) knew exactly what we wanted to do with the data, but we are unable to manage the data in a way that would make it useful for analysis. So we did nothing. The moral of the story here is that, while we were taught well the techniques which were useful for analyzing the data, we were never taught and had never learned any useful data manipulation techniques, rendering our statistical educations useless. It was not until I got my first job that I learned, out of necessity, data management techniques including SAS data steps, SAS macros, and SQL.

    When I returned to school to pursue a Ph. D., I saw many students with no work experience struggling through all of the same problems that I had with managing data. The same old “I know exactly what I want to do, but I can’t organize the data.” Often times in grad classes, books or teachers will describe a data set as “large” when it has several hundred or several thousand observations. This seems inadequate preparation for working in industry, as my first jobs often dealt with data sets with millions of observations and, later, a summer consulting project involved billions of observations.

    Currently, there are no required computing or data management classes in my program for earning a Ph. D. in statistics. I think there should be a required class in every statistics program covering data management issues and, at least, a solid introduction to programming.

    After, David Purdy’s talk, Chris Volinksy (Follow on Twitter) and he took questions. One interesting question that came up was about a second Netflix prize. However, Chris noted that this had to be cancelled because of privacy concerns. I’ve written before (or at least posted on Twitter) about some researchers who claim to have de-anonymized the data from the Netflix prize and, as a result, a lawsuit has been filed. (Netflix’s Impending (But Still Avoidable) Multi-Million Dollar Privacy Blunder) Whether you agree with canceling the prize over privacy concerns or not, it is clear that disclosure limitation is currently a big issue that certainly cannot be ignored.

    On Tuesday, I went to one of the sports research sections and saw two talks before I left to go see a talk about partially synthetic data in longitudinal data. The first speaker, Shane Jensen, spoke about evaluating fielders abilities in baseball using a method he proposed called Spatial Aggregate Fielding Evaluation (SAFE). The previous link explains how their evaluation of players works and gives measures of performance for each player. Probably, the most shocking result of his work is that, averaged over 2002-2008, SAFE evaluated Derek Jeter as the worst shortstop who met the minimum number of balls in play (BIP). Alternatively, SAFE rates Alex Rodriguez as the second best shortstop over this same period, even though he now plays third to allow Jeter to play SS.

    The next speaker was Ben Baumer, statistical analyst for the Mets (and native of the 413 area code). He spoke about his paper, “Using Simulation to Estimate the Impact of Baserunning Ability in Baseball“. One of the interesting things I took away from his talk is that he claims that players’ speed used to break up a double play is one of the important aspects of base running, but this is often largely or completely ignored as an evaluation tool of a players base running ability.

    Before I end, I’d like to say thanks to all the speakers that I saw speak this past week and, finally, I’ll leave you with a view of Vancouver from the convention center.

    Cheers.

  • According to infochimps.org these are the 25 most used emoticons on twitter.com. Download the whole data set yourself here.

    1 13458831 ๐Ÿ™‚
    2 3990560 :d
    3 3182129 ๐Ÿ˜ฆ
    4 2935301 ๐Ÿ˜‰
    5 2082486 ๐Ÿ™‚
    6 1461383 =)
    7 1439234 :p
    8 1013758 ๐Ÿ˜‰
    9 979947 (:
    10 669086 xd
    11 656784 :/
    12 595140 =d
    13 527391 =]
    14 490897 :]
    15 398246 ๐Ÿ˜ฆ
    16 367208 ๐Ÿ˜ฎ
    17 350291 d:
    18 332427 ;d
    19 321328 =(
    20 310343 =/
    21 252914 =p
    22 247794 ):
    23 240355 :-d
    24 217052 ๐Ÿ˜
    25 179184 ^_^

    Info about the data:
    “This data comes from a scrape of the Twitter social network conducted by the Monkeywrench Consultancy. The full scrape consists of 35 million users, 500 million tweets, and 1 billion relationships between users.
    This dataset is a corpus of tokens collected from tweets sent between March 2006 and November 2009. A โ€œtokenโ€ is either a hashtag (#data), a URL, or an emoticon (smiley face โ€” ;)). Think about comparing this data to the stock market, new movies, new video games, or even trendingtopics.org. For example, use it to look at the social networking adoption of Google Wave on the rate of its mentions.”

    Actually, all that I got in the free download was emoticon counts for the period between March 2006 and November 2009. So I got to thinking about what you could possibly do with this data in a useful way. What I was thinking about doing is trying to get a break out of emoticon usage by day or hour over the last few years. Then try to look for spikes in smileys, or frowns, or winks and see if these spikes are related to anything. Do you think we can measure world happiness or sadness based on emoticon use? (Probably not, but it’s an interesting thought, right?)

    Cheers.

  • The Joint Statistical Meetings (JSM) are coming up in the first week of August in Vancouver. StatsInTheWild will be in attendance.

    Here are the StatsInTheWild suggestions for interesting talks to attend.

    Disclosure Limitation and Confidentiality:
    A New Approach to Protect Confidentiality for Census Microdata with Missing Values Yajuan Si, Duke University; Jerome P. Reiter, Duke University. Monday, August 2, 2010 10:35 AM

    Disclosure Avoidance for Census 2010 and American Community Survey Five-Year Tabular Data Products
    Laura Zayatz, U.S. Census Bureau; Paul Massell , U.S. Census Bureau; Jason Lucero, U.S. Census Bureau; Asoka Ramanayake, U.S. Census Bureau. Monday, August 2, 2010 11:15 AM

    Multiple Imputation Method for Disclosure Limitation in Longitudinal Data
    Di An, Merck & Co., Inc.; Roderick Joseph Little, University of Michigan; James W. McNally, University of Michigan
    11:35 AM

    Balancing Individual Privacy with Access to Data for Policymaking
    Panelists:
    Stephen E. Fienberg , Carnegie Mellon University
    Nancy M. Gordon, U.S. Census Bureau
    Michael Lee Cohen, Committee on National Statistics
    Tom Krenzke, Westat
    Ed J. Christopher, Federal Highway Administration
    Stephen Gunnells, The Planning Center
    Wednesday, August 4, 2010 : 2:00 PM to 3:50 PM

    Sports:
    Spatial Modeling of Fielding in Major League Baseball Shane Jensen, The Wharton School, University of Pennsylvania. Tuesday, August 3, 2010 10:35 AM

    Exploring the Count in Baseball Jim Albert, Bowling Green State University. Tuesday, August 3, 2010 11:35 AM

    False Starts and Alternative Hypotheses Michael Rotkowitz, The University of Melbourne

    Multiple Imputation:
    Multiple Imputations for Survey Sampling and Their Diagnostics – Invited – Papers. Sun, 8/1/10, 2:00 PM – 3:50 PM
    This entire section is worth attending.

    Law:
    Formal Statistical Analysis Provides Sounder Inferences Than the U.S. Government’s ‘Four-Fifths Rule’: Examining the Data from Ricci v. DeStefano. Weiwen Miao, Haverford College; Joseph L. Gastwirth, Washington University. Wednesday, August 4, 2010 11:15 AM.

    See you in Vancouver.

    Cheers.

  • I attended the New England Statistics Symposium (NESS) last Saturday, and I’ve been meaning to write about one of talks I saw. After lunch, I went to the Columbia section so I could see the talk about multiple imputation using chain equations. The MI talk was the second in the section, so I sat through the first talk presented by Tian Zheng (Tian’s Blog) which turned out to be very interesting. The talk was about using social networks to learn about at risk populations.

    My understanding of this is that a survey could be given asking people questions about who they know rather than about themselves. For instance, instead of asking “Is your name Michael?” and “Are you homeless?” ask “How many Michaels do you know?” or “How many homeless people do you know?” Then using the responses to these questions, researchers can estimate how large at risk populations are. And they can do this without ever asking people who are in the at risk population! Really neat.

    Why is this useful?
    This excerpt from this flyer that was created to describe the method to a general audience says it very well:
    “AT-RISK POPULATIONS: At-risk populations can be hard to access (eg. homeless) or reluctant to admit their status for fear of others finding out (eg. HIV/AIDS, drug abusers, sex workers). Statisticians learn about these populations through their friends and acquaintances. Instead of asking if a person uses IV drugs, ask ‘How many IV drug users do you know?’ and use social structure to learn about the person using IV drugs.”

    Really, really neat stuff.

    Cheers.

  • StatsInTheWild NCAA Basketball top 25:
    (Boldindicates team is still in the NCAA tournament, italics indicate change from last week)
    1. Kansas 0
    2. Kentucky 0
    3. Syracuse 0
    4. West Virginia 0
    5. Duke +2
    6. Cornell NR
    7. Kansas State +2
    8. Purdue +5
    9. New Mexico -3
    10. Northern Iowa +13
    11. Butler +6
    12. Baylor -1
    13. Temple -5
    14. Tennessee +1
    15. Texas A&M -1
    16. Ohio State +3
    17. Villanova -12
    18. Xavier +6
    19. Georgetown -9
    20. Pittsburgh -8
    21. Michigan State NR
    22. Missouri NR
    23. Gonzaga NR
    24. Maryland -3
    25. San Diego State NR

    Other Notables: Saint Mary’s (26), Washington (33)

    Cheers.