• SITW NFL Rankings after week 1.

     Team Preseason Rank  After Week 1 Rank
     New England  1 1
     Green Bay  3  2
     New York Jets  5  3
     Baltimore  6  4
     Pittsburgh  2  5
     Atlanta  4  6
     Chicago  7  7
     Philadelphia  10  8
     New Orleans  9  9
     Tampa Bay  8  10
     New York Giants  11  11
     Miami  13  12
     Indianapolis  12  13
     Detroit  16  14
     San Diego  15  15
     Jacksonville  18  16
     Kansas City  14  17
     Minnesota  17  18
     Oakland  20  19
    Buffalo  24  20
    Washington  23  21
    Cincinnati  26  22
    Cleveland  19  23
    Tennessee  21  24
    Dallas  25  25
    Seattle  22  26
    Houston  29  27
    San Francisco  28  28
    St. Louis  27  29
    Arizona  30  30
     Denver  31  31
     Carolina  32  32

    Cheers.

  • Auto-complete for search “Rick Perry ” on Google over the last couple of weeks. The last row is the polling percentage based on Real Clear Politics polls.

    8-17-2011 8-30-2011 9-6-2011 9-9-2011 9-12-2011
     for president  for president  for president for president gay
    gay  gay  gay gay for president
     wiki  wiki  wiki wiki wiki
     for president website  2012  prayer prayer prayer
     2012  for president 2012  2012 galileo secession
    18.4 23 29 29 31.8

    Auto-complete for search “Rick Perry is ” on google over the last couple of weeks. The last row is the polling percentage based on Real Clear Politics polls.

    8-17-2011 8-30-2011 9-6-2011 9-9-2011 9-12-2011
    gay gay gay gay gay
    an idiot an idiot an idiot an idiot an idiot
    a rino a rino crazy crazy crazy
    evil evil nuts nuts scary
    not a conservative not a conservative stupid stupid evil
    18.4 23 29 29 31.8

    Cheers.

  • Republican presidential debate, Obama addressing the nation, AND the start of the NFL season. It’s almost too much to handle.

    Before we get to any NFL predictions, I’ll make a presidential prediction.  Mitt Romney is going to win the Republican nomination.  I don’t care what Perry’s poll numbers are right now.  I don’t think Republican’s will vote for a guy who’s first google auto-complete term is “gay”.  (Maybe I under-estimate Republican’s tolerance, but then again, maybe I don’t.)

    Anyway, I’ve been toying with the idea of simulating the NFL season for a little while now (I did a bit of this last year, later in the season.) This year I’ve done it before any games have been played, so we’ll see how this model works out. it’s a pretty simply model and uses only data from the 2010-2011 regular season and playoffs.  Using that data, I used a logistic regression model to model the probability that one team beats another team.  Then I simulated the upcoming season 5000 times.

    Let’s begin with some pictures.  The first has nothing to do with the simulations, but it’s interesting.  It also give me a chance to quote myself.  So, here is a plot of some Chernoff faces based on the final 2010 NFL regular season team statistics.  (I posted about this before here)

    My comments from before:

    The face represents the offense and the defense is represented by hair. The size of the nose indicates sacks, the ears indicate turnovers (ear width is interceptions; ear height is forced fumbles).  The eyes indicate penalties and, finally, the size of the mouth indicates wins with a smiling face if the team made the playoffs (a really nice touch, if you ask me.)  The face at the bottom right indicates the league leader.

    Some observations on the NFL faces:  The two superbowl teams last year (Pittsburgh and Green Bay) are both located at the bottom of the graph and there faces look very, very similar.  San Diego looks similar to to both Green Bay and Pittsburgh (similar face, nose, eyes, and hair), but the big differences are the ears and, of course, the San Diego face is frowning.  Another thing that pops out at me is how similar Houston and New England look to each other.  They have very similar face shape, eyes, and hair.  The big differences are the nose and ears (sacks and turnovers).

    Here is a graph with 32 side by side boxplots representing each of the NFL teams.  Each boxplot displays the distribution of the predicted number of wins for each team.  The teams are in order of the SITW power ranking (which means it’s mostly made up).  I have also included a red W for how many wins the team had last year, a green dollar sign for the over-under betting line, and a blue P indicating whether or not the team made the playoffs last year.

    Now it’s time for my Super Bowl favorites table.  The first column lists the team, the second column lists the my predicted odds to win the 2012 Super Bowl, and column three displays my predicted probability of each team making the playoffs.  One interesting thing to note in the first couple lines of this table is that Pittsburgh is more likely to win the Super Bowl than Baltimore, but Baltimore is more likely to make it to the playoffs than Pittsburgh.  This is a result of the NFL scheduling system.  Pittsburgh and Baltimore share the exact same schedule except for two games.  Those two differing games for Pittsburgh are New England and Kansas City whereas those two games for Baltimore are the New York Jets and the San Diego Chargers.  So what is happening is that because Pittsburgh has the chance to play New England in the regular season, in the simulations, when they do make it to the playoffs, they are most often making the playoffs as a 1 seed.  Baltimore is making the playoffs more often, but they are rarely (relative to Pittsburgh) simulated to be a 1 seed.  Remember, this table is ordered by odds that a team wins the Super Bowl; it’s not ordered best to worst team.

     Team S.B. XLVI Odds  Prob(Make Playoffs)
     New England  6.8  .7018
     Pittsburgh  9.4  .6402
     Atlanta  10  .7164
     Baltimore  11  .7156
     New York Jets  14  .5352
     Chicago  16  .5842
     Green Bay  16  .5402
     Tampa Bay  21  .4636
     Philadelphia  21  .449
     New York Giants  24  .4294
     New Orleans  28  .393
     Indianapolis  37  .4328
     Miami  46  .3174
     Seattle  60  .3512
     San Diego  60  .3254
     Kansas City  63  .3596
     Minnesota  64  .2902
     Detroit  69  .3082
     Jacksonville  75  .2952
     Dallas  76  .2678
     Oakland  78  .3274
    Washington  88  .2438
    Cincinnati  99  .2514
    Cleveland  103  .2538
    St. Louis  105  .2916
    Tennessee  108  .2872
    Buffalo  131  .208
    San Francisco  146  .2826
    Arizona  160  .2732
    Houston  160  .2048
    Denver  216  .1442
    Carolina  555  .1156

    Below are some over-under bets that I like.  It seems like a lot of times forget that they are betting on the NFL. Every single team can beat every other team (See Miami beating New England as 13.5 point underdogs a few years ago.) Betters over value good teams and under value bad teams.  My two favorite bets here are Green Bay and Cincinnati.  Green Bay had a great run through the playoffs last year, but they still only won 10 regulars season games in 2010.  Add that to the fact that Aaron Rodgers is a concussion waiting to happen and winning 12 games seems like a difficult task.  Cincinnati has been blessed by the scheduling gods.  Not only did they finish fourth in their division last year earning them games against Denver and Buffalo they have also drawn the NFC west division giving them games against Arizona, Seattle, San Francisco, and St. Louis.  Then add to that two games against Cleveland and it’s not to hard to see 6+ wins in their future.  And think about this, they start their season Cleveland, Denver, San Francisco, and Buffalo.  Is it that far fetched that they start 4-0?  (Yes, it is that far fetched.  Just saying is all….)

    Team Bet Odds
    Green Bay  Under 11.5  -145
     San Diego  Under 10  +115
     Minnesota  Under 10  +110
     Houston  Under 9  +145
     Philadelphia Under 10.5  +120
     Dallas Under 9  -120
     Cincinnati Over 5.5  +135
     Carolina Over 4.5  even
     Seattle Over 6  +125
     Buffalo Over 5.5  -135
     Oakland Over 6.5  +110
    New England Under 11.5 -110

    Other bets that intrigue me.

    Team Bet Odds
     Seattle  Win Division +900
     Oakland  Win Division +700
     Chicago  Win Division +600
     Washington  Win Division +2000
     Minnesota  Win Division +1200
     Kansas City  Win Division +500
     Baltimore AFC Champs +900
     Atlanta NFC Champs +600
     Tampa Bay NFC Champs +1500
     Seattle NFC Champs +4500
     Chicago NFC Champs +2000
     Washington NFC Champs +4500
     Atlanta Super Bowl Champs +1200
     Seattle Super Bowl Champs +8000
     Tampa Bay Super Bowl Champs +3000
     Baltimore Super Bowl Champs +2000
     Chicago Super Bowl Champs +4000

    And finally, let’s make some predictions that will ultimately prove to be way off.  But it is fun to try here is what the playoffs will look like.

    The AFC.

    Team Seed Mean wins
     New England  1  10.116
     Pittsburgh  2  9.88
     Indianapolis  3  8.1634
     Kansas City  4  7.7488
     Baltimore  5  9.648
     New York Jets  6  9.3434

    I know, I know.  It’s boring and it’s exactly the same as last years AFC playoff teams down to the seeds.  But wait until you see my NFC picks!

    NFC

    Team Seed Mean wins
     Atlanta  1  9.4892
     Chicago  2  8.9044
     Philadelphia  3  8.5046
     Seattle  4  7.6006
     Green Bay  5  8.8728
     Tampa Bay  6  8.7434

    Ok.  Those weren’t that exciting either.  At least I made a stand with Tampa Bay, right?

    And now for my Super Bowl prediction.  Based solely on the numbers I am taking new England over Atlanta.  That’s wicked boring though.  So my gut is taking Tampa Bay over Baltimore 21-20.  And I’m still picking Mitt Romney.

    Cheers.

  • So, I’ve got a lot of blog posts that I meant to publish last week, but I never got around to it.  Here is a graph I made using the the auto-complete terms from Google, Yahoo, and Bing for republican presidential candidates.  I looked at the five top auto-completes from each site and scored each word 5 points if it was the first auto-complete, 4 points for second auto-complete, etc.  I did a search for each candidate twice on each site.  First using just the candidates name and a space, then the candidates name followed by the word “is” and then a space. (For example, “Mitt Romney ” and “Mitt Romney is “).   I then weighted the search engines based on their market share (about 75%, 15%, and 10% respectively).  This gives me a data set with 8 observations (8 candidates) and several dozen variables (one variable for each word).  I then used mutli-dimensional scaling to reduce the distances between the vectors down to, in this case, three dimensions.  The size of each circle is proportional to the polling percentage from RealClearPolitics on August 29, 2011 (the same day as the auto-completes were done.)  The word appearing in or next to each circle, is the word with the highest score for each candidate.

    Also, one of Michele Bachmann’s auto-complete terms on Google is “slavery”.  I couldn’t imagine what she had done to warrant this as an auto-complete term, but then I found this article by Andrew Gelman (of the blog Statistical Modeling, Causal Inference, and Social Science).  Yikes.

    Cheers.

     

  • Here are two interesting articles related to statistics that were featured on Slate.com two Mondays ago:

    The first article, by Kevin Gold, is called “The Leaky Nature of Online Privacy: Network analysis can uncover your personal details even if you choose to hide them.”  This led me to LaTanya Sweeney’s webpage (of k-anonymity fame), which I then spent quite a bit of time reading.  (I found the work on face de-identification to be very interesting.)

    On that same day on Slate, everyone’s favorite former governor of New York, Eliot Spitzer (If you haven’t seen “Client 9” yet, stop what you are doing and watch it) had an article called “World Defeats U.S. in Four Sets: How the decline of American men’s tennis can explain global economics.”  In the article, Spitzer discusses the difference between correlation and causation as it relates to tennis and the economy.

    Cheers.

     

  • Auto-complete for “Rick Perry is” on the three big search sites on 8/29/2011.

    Google Yahoo Bing
    gay an idiot an idiot
    an idiot crazy good
    a rino a scumbag a crook
    evil not a conservative bad
    not a conservative a republican a scumbag
    evil running for president
    awesome right about education
    hot
    horrible
    a joke
    
    

    Cheers.

  • Recently, I posted (“Multidimensional Scaling, Republican Presidential Candidates, and ‘a douchebag” and “Tracking the Republican candidates via google auto-complete“) about Google auto-complete and potential Republican presidential candidates.  Slate.com posted a good piece called “Google’s GOP Search Suggestion” (a day after my original post, I should note) where they look at the auto-complete for candidates names using a Google image search rather than a straight Google search.

    Cheers.
  • The tables below are for the Google and Yahoo search “Michele Bachmann “(including a space after the last name) for the various dates indicated in the table. Each column has the date of the search and the top five Google or Yahoo auto-complete terms for the search.

    Michele Bachmann – Google

    8-17-2011 8-22-2011 8-23-2011 8-29-2011 8-31-2011 – present
    quotes quotes quotes quotes quotes
    corn dog husband husband husband husband
    husband bio bio bio wiki
    elvis corn dog slavery slavery husband gay
    bio slavery husband gay husband gay hot

    Michele Bachmann – Yahoo

    8-24-2011 8-29-2011 9-1-2011 9-2-2011 9-7-2011
    hot hurricane hurricane new hair campaign manager
    for president sarasota irene margaret thatcher hurricane irene
    minnesota hot hot hot hot
    bio for president for president for president for president
    feet minnesota minnesota minnesota minnesota

    What does all this mean?  I have no idea, but I suspect it will be difficult to win a Republican party nomination and then a general election with terms like “slavery” and “husband gay” attached to your name.

    Another thought: I wonder if the political affiliations of users are constant across the three major search sites or are there a greater percentage of liberals on Bing than on Google, for instance.  Could you use auto-complete terms to gain any insight into this?  Or is this type of information perhaps already available?

    Cheers.

  • My friend Scot recently sent me a g-chat about a new search engine DuckDuckGo.  According to their website “DuckDuckGo is a general purpose search engine like Google or Bing.”  They then offer four bullet points:

    • Get way more instant answers
    • Less spam and clutter
    • Lots and lots of goodies
    • Real Privacy

    Those first three sound interesting, but what really piqued my interest was the fourth bullet: Real Privacy.  DuckDuckGo will not collect any of your browsing information, which, in turn, could have been used to identify you and potentialyl reveal what you are searching for.  Many people might not have a problem with this, but DuckDuckGo offers a very nice illustrated example of why this is potentially a problem.  They go on to say in their privacy statement:

    “It’s sort of creepy that people at search engines can see all this info about you, but that is not the main concern. The main concern is when they either a) release it to the public or b) give it to law enforcement. ”

    “Why would they release it to the public? AOL famously released supposedly anonymous search terms for research purposes, except they didn’t do a good job of making them completely anonymous, and they were ultimately sued over it. In fact, almost every attempt to anonymize data has similarly been later found out to be much less anonymous than initially thought.”

    That last line is particularly interesting.  Two examples of this that immediately come to mind are the GIC insurance example and the Netflix prize example.  All of these, the GIC, AOL, and Netflix examples, all released data to the public for research purposes.  And in all of these examples, the releasing organization realized that they could not simply release the data to the public because of privacy concerns.  They needed to do something to anonymize the data, so they did something (I’ve talked about this doing something before).  But in all of these cases, the supposedly anonymous data all turned out to be, to varying degrees, less anonymous that originally thought.  Simple ad hoc procedures like deleting information simply don’t work in protecting privacy unless you live under a rock and have no access to auxilliary information.  The only way to be completely safe is to release no data at all. However, releasing nothing to the public prevents valuable research from being done.  While GIC, AOL, and Netflix all released data that ended up being less than anonymous, you have to applaud their effort to allow researchers to do what they do: research.  The Netflix prize produced plenty of valuable research (Lesson from the Netflix Prize Challenge) and the GIC data had the potential to produce valuable, potentially life saving public health research.  Like most things in life, some balance must be found somewhere between the extremes, and the potential benefits of any research must be weighed against the potential costs of a privacy breach.

    I see it like this: If you have something of value in your house, you wouldn’t leave the door wide open; you’d lock the door.  But no matter what, if someone really wants to, they could break into your house with enough effort.  Either way it’s still illegal/unethical.  It’s the job of statisticians to put as many locks on the house as possible while still being able to reasonably use the house; it’s the lawyers job to prosecute people who break into the house, whether or not the door is well secured.

    Cheers.

  • If you don’t want to read this whole thing, just check out the graph: Multidimensional Scaling: Republican Candidates – 8/16/2011

    I was having a conversation with some friends today and someone mentioned that Rick Perry might have problems in the election because there were rumors he was gay.  So I went to google and typed in “Rick Perry is” and google kindly offered me the following auto-complete options: “gay”, “an idiot”, “a rino“, “evil”, “not a conservative”.  This got me thinking how this compared with the other candidates google auto-completes.  For instance, if you google “Mitt Romney is” you get suggestions like “a mormon” and ” an idiot” as well as three other suggestions.  I did this for all of the major candidates (sorry Thaddeus) and recorded the five google auto-complete suggestions.

    Then I created a vector for each candidate based on the google auto-complete words.  Each candidate was an observation and each word was a variable.  The candidate would get a 5 if the word was first on their list, a 4 if it was second, and so on with a 0 if the word was not mentioned in their auto-complete.

    I then used multidimensional scaling (the cmdscale function in R) to allow me to visually display the relative positions of the candidates to each other.  This all led to this graphic: Multidimensional Scaling: Republican Candidates – 8/16/2011.  The location of the circles is based on multidimensional scaling, the size of the circle is relative to their standings in a national poll taken from fivethirtyeight.com, and the top five google auto-completes are displayed in or near the appropriate circle.

    Some thoughts:

    • Every single candidate has the term “an idiot” in either the first or second auto-complete term
    • 3 candidates were listed as “hot” (Palin. Bachmann, and Romney)
    • “stupid” was only used to describe women
    • Perry and Santorum (who has a much bigger google problem that anything I’ve listed here) had “gay” listed in their autocpmpletes and Pawlenty had “definitely not gay”
    • Bachman and Palins circles are nearly identical in size (11.7% ad 11.4%, respectively) and words (they share “an idiot”, “hot”, and “stupid”)
    • “a douchebag” appears in auto-completes for Santorum, Gingrich, and Pawlenty.  I imagine it will be hard to win with this word attached to your name. (John Kerry couldn’t do it.)
    • The only overwhelmingly positive google auto-complete was for Herman Cain whose fifth auto-complete option was “awesome”
    It can’t be good for Perry that he is so close to Pawlenty and Santorum, but he does have a significant amount of support at this point.  I’ll be interested to see how these Google auto-completes changes over time and with the polls.
    For information on how Google auto-complete works, click here.
    Cheers.