• Here is a graphical representation of box office receipts for the movies from 1986 through 2008.

    This is a great display of data. We can see how much money individual moves made, as well as the length of time they were in theater. We can also easily see the overall trend of the movie industry over time and seasonally.

    Cheers.

  • I’ve recently become fascinated by politics of the Census, and, with a Census coming up in 2010, I think it’s a perfect topic for StatsInTheWild.

    Politics and the census have been an issue for decades. Here is a piece from the New York Times from August 1909. In it, they cite a letter from President Taft to the Secretary of Commerce and Labor, requesting that politics be removed from the process of the Census. However, it looks as if Taft’s pleas for a non-partisian Census are not being heeded by todays politicians.

    Recently, I posted an article about Obama nominating a new head of the Census who is an expert in sampling which has ruffled Republican feathers. Why would republicans be against sampling, a process that makes the Census more accurate and, ultimately, less expensive? Well, here is a very good article from 1998 explaining the Republican opposition to using sampling techniques for the Census.

    The basic gist of the article is that government officials from the Clinton administration wanted to use sampling methods to account for the traditional undercount of minorities. Republicans likely blocked this from happening because the people who the Census tends to miss are people who tend to vote Democrat. Republicans (led by former speaker Newt Gingrich and Bobb Barr) argued that sampling was unconstituational and was a violation of the Census Act. Their argument held up in court, and no sampling was allowed to be used for the allocation of federal funds or redistricting.

    Here is a good quote from the article if you don’t want to read the whole thing:
    “The Bureau wants to do better. By using sophisticated data-gathering and statistical-sampling techniques to augment the direct count, it believes it can reduce the total and differential undercounts and save money to boot. The National Academy of Sciences agrees, as does about every statistician worth his or her salt. In 1990, the Census Bureau thought such sampling were the way to go, but Republican officials in the Bush Administration overruled the Bureau’s experts. The courts refused to intervene. ”

    Cheers.

  • I only got one of the Final Four teams this year, but that team is UNC and they look like they have a good shot to win it. It’s never a total loss when your champion pick is still alive.

    My bracket might stink, but my “bets” have been good. For the tournament I am 8-5 with a 19.82% return per bet. Here are 4 more.

    MSU +160
    Villanova +280

    MSU +650 to win championship
    Villavona +700 to win Championship

    **********
    After the championship game, I finished 9-8 with a 9.05% return per bet. As pointed out in the comments this is a very small sample size. I agree. So I want to make it clear that I am not using this as evidence that I can “beat” the sports book. I am merely reporting my results.

    I think I can beat the sports books, but I’d need to be profitable over several hundred bets to begin to approach significant evidence that I am winning consistantly.

    Thanks for the comment.

    Cheers.

  • Here is an interesting article from the UK:
    “Numbers up: The truth about statistics”

    And in the spirit of correlation of the week, here is a good excerpt from that article:

    “Toothless post-menopausal women are three times as prone to hypertension as those with teeth”

    This news, reported in the respected journal Hypertension, might have lead to queues of denture-wearing women of a certain age at GPs’ surgeries. A study by Japanese researchers from Hiroshima University, published in 2004, suggested that tooth loss in post-menopausal women was directly linked to high blood pressure, which can increase the risk of heart disease or strokes.

    But a look past the headlines revealed a problem: the scientists based the conclusion on a study of just 98 post-menopausal women – 67 with missing teeth, and 31 with their gnashers intact. In statistical terms, that is an almost insignificant sample size.

    The problem is that the apparent cause of a link can sometimes be pure chance. The smaller the sample, the more likely this becomes. One statistician famously managed to find a statistically significant correlation (in a small enough sample) between birth rates in various European countries and the stork population, suggesting the birds therefore really do deliver babies.

    McConway’s verdict: “There’s no standard minimum group size for statistical studies – it depends what you’re measuring. If it’s something that doesn’t vary much – say, blood pressure in elite athletes – you could get away with a smaller group. But for something like this, you need a much larger sample.”

    Cheers.

  • I was watching the show “Predator X” on the History Channel tonight. Apparently, they discovered this fossil of an enormous aquatic predator. It’s pretty awesome.

    pliosaur-vs-plesiosaur

    Here is a description of the bite force of this predator: (from here)
    “At St Augustine Alligator Farm and Zoological Park in Florida , Dr. Hurum assisted evolutionary biologist Dr. Greg Erickson from Florida State University in calculating the bite force of this colossal creature. The jaws held in place a set of trihedral teeth, each measuring 12 inches, which clamped down on prey with an estimated 33,000lbs of bite force. The calculation is one of the largest bite forces ever calculated for any creature. Predator X would have had a bite a bite force was more than ten times the bite force of any animal alive today and four times the bite force of a T- Rex.”

    (Here is a link of average bite forces for humans and a few selected animals.)

    At this point you might be saying, “That IS awesome. But what does it have to do with this blog?” An astute observation. Well…

    They estimated that the bite force of the predator was 33,000 pounds. The way they estimated this was by taking measurements of the bit force of different sized crocodiles (or alligators, I can never tell the difference.) Then they plotted the data in a scatter plot, weight of crocodile versus bite force. There was a clear postive relationship between bite force and size of the animal. Then they fit a simple regression line through the data and extrapolated how much bite force a 50 foot long animal that weighed an estimated 45 tons could pack in its bite force. That’s how I believe they came up with there estimated bite force. I’ll give them the benefit of the doubt and assume they did more than that to come up with the estimate but they didn’t want to show the details on the History channel. (If you have details on how they estimated the bite force, please send them my way.)

    Let’s assume that all they did was extrapolate this simple regression line. What would be the problem with that? The problem is that they are extrapolating the linear trend outside of their domain. There is no guarantee that the bite force trend remains linear as the weight approaches the estimated 45 tons. They collected their data on animals with weight of crocodiles which can be up to 1.5 tons. It seems naive to assume that the linear trend will continue as you increase the weight of an animal.

    Here is a simple example of why extrapolating outside of your domain is a bad idea.
    Say you collected data of children’s heights and weights and you fit a regression line through the data. You’ll surely observe a postive relationship between age and height. As children get older, their heights generally increase. This increase can be approximated by a roughly linear trend say between the ages of 10-18. Also, say that we find that children grow on average of 1 inch per year between 10-18. If I were then to predict the height of a person by extrapolating out this trend I would assume that a 48 year old would be, on average, 30 inches taller than an 18 year old. Clearly, this is not true.

    So just because a trend is linear over a certain domain does not mean that that linear trend continues outside of the tested domain.

    Cheers.

  • After re-evaluating, here are my picks after the first two round of the NCAA tournament:
    Elite 8: Uconn, Pitt, Memphis, Villanova, Louisville, Oklahoma, Michigan State, UNC

    My final four is the same:Memphis, Louisville, Pitt, and UNC

    Finals: Memphis versus Pitt

    Winner: Memphis

    Good picks for the sweet 16 games: (Winner are in Bold, losers are in italics.)
    Villanova +120 (I have them as a favorite in this game).
    Michigan State -1.5 at -110
    Gonzaga +350 (They’ll probably lose, but this price is fantastic.)
    Pitt -320 (Added 3/26/2009 11:28 am)
    Results: 3-1. A bet of “100” on each of these games yields a profit of 142.16 for a 35.54% return.
    For the tournament I am 8-5 with a 19.82% return per bet.

    Check out my results from the first round.

    NIT predictions:
    San Diego State beats St. Mary’s Tonight.
    Notre Dame Beats Kentucky.

    Then, San Diego State beats Baylor to go to the Finals, and Penn State beats Notre Dame for their spot in the finals.

    And I’m switching to Penn State as my pick for the NIT champion.

    Good luck.

    And.

    Cheers.

  • Check out page 35 of the old issue of Chance magazine which has an article talking about the best way to pick NCAA basketball teams to win your office pool. Of course this probably would have been more of a help if I had posted it five days ago, but it’s still interesting. And you can use the advice next year.

    Cheer

  • Of course, we’re not gambling with American dollars. That would be illegal, so we use “standard betting units” here. (1 “standard betting unit” is approximately 1 American dollar. But it is definitely NOT an American dollar.)

    From my previous post before the first round started. Winners are in bold, losers are in italics:

    Good first round bets:
    Washington -220
    FSU -145
    Utah -110
    Illinois -200
    Arizona St -200
    Michigan +190 (This price is fantastic)
    Texas A and M +115
    Oklahoma St +115 (Seriously? They’re an underdog here?)
    Ohio St -160

    If you bet 100 “standard betting units” per game, you would have went 5-4 and be up 115.45 “standard betting units” at this point. That is a 12.8% return per bet. Not bad.

    Cheers.

  • So I’m back from ENAR and back from spring break. I’ve ben greeted back to grad school by a midtern on Friday night from 6-8pm. What a fun time for a midterm!

    Let me first start by saying that San Antonio is awesome. The river walk is great. I at at two great Mexican restaurants for lunch two days in a row and they were both incredible. And I drank Shiner Bock the whole time, which I highly recommend.

    Anyway, on Sunday night I presented a poster at ENAR (they served Shiner Bock during the poster presentations) about a paper that we wrote (and recently published) about synthetic data with binary variables. I met some very interesting people who stopped by my poster. One guy who stopped by informed me that he coined the term predictive mean matching (which i referenced on my poster). So, I asked him who he was, and he told me he was Rod Little (He’s kind of a big deal). He wrote the book on multiple imuptation: Little, R.J.A. & Rubin, D.B. (1987). Statistical Analysis with Missing Data. New York:
    John Wiley. So that was kind of neat. (I just visited Rod Little’s website and apparently this is the “most useful of all links”. (Here is another interesting article called “Calibrated Bayes: A Bayes/Frequentist Roadmap”.)

    The next day I spoke with some people from SAS and STATA, as well as, some recruiters from Smith-Hanley (who I got my first job through) and Cambridge Group. The SAS people told me about a product call JMP, which I was very impressed by it. The STATA people told me that I could buy a student STATA license for like $55 dollars and then use it commerically after I graduate. (As opposed to several thousand dollars for a SAS license that only lasts a year.) And I could use it for as long as I wanted to. The only thing I would have to pay fo rwould be upgrades. So STATA has that going for them. I am definately gonig to try it out.

    Cheers.