Showing posts with label forecasting. Show all posts
Showing posts with label forecasting. Show all posts

21 December 2007

When asteroids attack

Asteroid May Hit Mars in Next Month (AP).

Scientists tracking the asteroid, currently halfway between Earth and Mars, initially put the odds of impact at 1 in 350 but increased the chances this week. Scientists expect the odds to diminish again early next month after getting new observations of the asteroid's orbit, Chesley said.


If they expect the odds to diminish, why haven't they set them lower in the first place? I think this may be another issue of mean-median confusion -- say, there's a 1 in 3 chance that the odds will go up to 1 in 25 next month and a 2 in 3 chance it'll go away completely. But the statement seems kind of silly. Apparently efficient market theory doesn't apply to asteroids.

10 August 2007

the traffic.com "Jam Factor"

Traffic.com has something they refer to as the "Jam Factor" to tell you how much traffic there is on various major roads in your area. It's an apparently arbitrary number from 0 to 10.

Now, if I see that the jam factor is 3.1 right now, what does that mean? How long is it going to take me to get where I'm going?

They claim:

The Traffic.com Jam Factor is like a "Richter scale" for traffic. It's an overall measure of the traffic conditions on a roadway, or on a section of a roadway. Because the Jam Factor calculation uses real-time speed and travel time measurements from our sensors and those of our partners, as well as our detailed accident, construction and congestion information, it's a comprehensive measure of the state of traffic on any roadway.

The Jam Factor is measured on a scale of 0-10, with 10 being the worst traffic conditions. It is designed to give you a quick, at-a-glance picture of conditions on the roadways or personal MyTraffic Drives you care about, whether you're on our web site, looking at an emailed Traffic Report, or listening to a Traffic Report call to your phone. If you see (or hear) a high Jam Factor, you can then delve into the detailed information in the Traffic Report or on the Traffic.com site to find out more.


If you click on the name of any segment of road, it'll tell you how long it takes to get by that piece of road without traffic, and how long it'll take Right Now -- that seems like more meaningful data to me.

If I had to sum up traffic conditions in a single number, it would be the ratio of the two numbers in the previous paragraph. I'd like to know that it'll take me 50% longer to get where I'm going than it would if the roads were clear. It wouldn't surprise me to learn that the "jam factor" is just this number disguised somehow.

I prefer what, say, Google Maps does, just overlaying colors (red/yellow/green) on the road; a number, especially one with two significant digits (the Jam Factor is reported to the nearest tenth), implies some sort of precision. I'd rather have no number than a number which has extra decimal places tacked onto it to seem "scientific". (Would I be less annoyed by this if the Jam Factor was reported on a 0-100 scale, to the nearest unit? Who knows?)

Also, if they have sensors of some sort -- why am I limited to knowing how long it'll take to get from a preset point A to a preset point B? Why can't I input the exit I actually get on and off at and have it tell me how long I should expect to spend on the highway? This is a simple matter; they have the speed that the road is moving at at any given moment, so just integrate the reciprocal of speed over the length of road in question. (I realize that it's not quite this trivial, because just because a segment of road is moving at 30 mph right now doesn't mean it'll be moving at that same speed when I get there. My method is sort of like how life expectancies are calculated, which doesn't actually tell YOU how long you're going to live. But forecasting how traffic on any given day will evolve, or how medical care will evolve, is a lot harder than just observing it.)

The Richter scale itself is logarithmic, as many of you probably know -- a difference of one unit on the Richter scale corresponds to a factor of ten in the amplitude of the seismic waves (not a factor of ten in the energy released, as is widely reported; although this isn't my field, it looks like the corresponding increase in energy released is a factor of 103/2, or about 32). Other logarithmic scales that you see fairly frequently are decibels (for measuring sound) and Google PageRank (on the 0 to 10 scale); a decibel corresponds to an increase in volume of 100.1, and I've read various things but one unit of PageRank seems to correspond to a factor of about five. (No one outside of Google knows for sure.) All of these situations have one thing in common -- most earthquakes, sounds, and web pages are relatively small, and some are much more prominent, so there's a good reason to spread out the small end of the scale and compress the large end. But in the case of the Jam Factor, this isn't necessary; if it takes me an hour on average to drive somewhere, on a really lucky day it might take forty minutes, and on a really unlucky day it might take three hours, but it's never going to take even as much as a hundred hours.

(Incidentally, as you may h ave guessed by now, my preferred mode of transportation is walking. A large part of this is that even though it's slow, I always know exactly how long it's going to take me. There is essentially no situation that slows down or speeds up my walking speed. And even if there were, it never takes me twice as long to walk somewhere as it ordinarily would, whereas driving times that are twice the average are routine. Unfortunately, it's not totally clear whether I save fossil fuel by walking, because I have to eat more food in order to walk.)

02 August 2007

you can't predict the climate, either

Julie Rehmeyer, of MathTrek, writes about how predicting the weather -- or climate -- is basically impossible. The article begins:
Climate models may never produce predictions that agree with one another, even with dramatic improvements in their ability to imitate the physics and chemistry of the atmosphere and oceans. That's the conclusion of a report by James McWilliams, an applied mathematician and earth scientist at the University of California, Los Angeles. The mathematics of complex models guarantees that they will differ from one another, he argues. Therefore, says McWilliams, climate modelers need to change their approach to making predictions.

I had been under the impression that this was already known. It's known as "sensitive dependence on initial conditions"; if we know the weather to within a certain precision ε right now, then after one day we know the weather to within kε, after two days we know it to within k2ε, and so on, where k is some constant larger than 1 . More technically, the Lyapunov exponent of the weather is positive.

I've probably read several dozen versions of the story that is usually told about Edward Lorenz's toy model of the weather. (It would be interesting to see a web page that gives the various ways this particular story has been told, something like this page which gives over a hundred versions of the story of the young Gauss summing 1 + 2 + ... + 100.) The story, if you're not familiar with it, goes like this: Lorenz had a toy model of the weather in his computer, a system of differential equations. (I want to say it was a system of three equations, but I might be confusing it with the Lorenz attractor. Then again, they may actually be the same system.) It was the sixties, so computers were slow. Lorenz had his computer print out the position of the system in phase space at time 0, 1, 2, ...; one day he was looking at one of these printouts and saw a pattern he wanted to investigate. He fired up the computer again and typed in a line from the printout and told it to evolve the system from that point. The system evolved differently in the second run than the first; Lorenz thought it was a mistake, but eventually realized that the figures on the printout were rounded versions of the actual numbers in the computer, so he was introducing a small error by doing this, which was quickly amplified.

The MathTrek article is about climate, not weather, though, and it addresses this point (even anticipating my complaint about Lorenz!). Still, my instinct would have been -- even before reading this -- that the climate is a complex system just like the weather. The Lyapunov time -- the reciprocal of the Lyapunov exponent -- is much larger for climate than for weather. (I am quite confident saying that the average high temperature in Philadelphia next August will be about eighty-three degrees, and I am confident enough in this that when I take my air conditioner down when the summer ends, I will store it in my closet, instead of selling it. But I have a much worse idea what the weather will be on August 2, 2008.) Roughly speaking, climate is the average of weather, and averages change much less quickly than the things being averaged. I am confident that the Phillies will win about 29 of their remaining 55 games and just miss the playoffs, which is something I've gotten quite used to. (I'd be pleasantly surprised if they prove me wrong.) I have no idea whether they'll win tomorrow. (I would have said "I have no idea whether they'll win today," but they're up by four runs right now. It's only the fourth inning, though, so they have time to fall apart.)

The actual study is available here. Apparently the state of the art in climate and weather forecasting is to run a variety of different models on the same initial data; if they end up giving similar results then you can be fairly confident in the correctness of the forecast, while if they vary widely you know the forecast isn't so good. This is an experimental way of determining how sensitive the forecast is to the assumptions of the model. Although I'm not a meteorologist, it would be kind of interesting to see this on weather forecasts. I'm not sure how useful it would be for temperature; would I take a forecast high of "94, plus or minus 3" any differently than a forecast high of just "94"? Probably not. But for, say, snowfall estimates it could be incredibly useful. I don't care so much if they say there will be "two inches of snow". What I really want to know is if there's a chance of having some amount of snow that will seriously inconvenience me (say, over six inches). But I doubt you'll hear this on the TV news, because "we don't really know what the weather is going to be" kills the ratings -- even though "everyone knows" that the TV weather people don't really know what the weather is going to be. I've been known to actually use something like this "ensemble forecasting" myself -- I go to a bunch of different weather forecasts and see what they say. I'm not sure if it actually helps me, but it makes me feel better, usually because when I'm checking multiple weather web sites it means I'm procrastinating.

24 July 2007

predicting baseball attendance

Back in July I wondered where baseball attendance numbers come from, and pointed to J. C. Bradbury's post about the Harry Potter effect on baseball attendance; he came up with a quick-and-dirty estimate of the number of people who didn't go to baseball games because they were reading the new Harry Potter book. This got me thinking -- how could we know how many people we expect to go to any given baseball game?

This paper by Brian Stoner gives the results of a regression analysis which forecasts the average attendance of a baseball team, over the course of a season, as a function of average ticket price, number of seats in the team's stadium, the previous year's attendance, the team payroll in the current year and number of wins in the previous year, the team's television market size, and whether or not the team has a new stadium; these turn out to explain most (98%) of the variation over the sample period. For the rest of this post, let's assume that we can predict a team's average attendance; we know it'll be, say, 32,000 per game. (Equivalently, we can predict the attendance for the entire season, since the average is just this figure divided by 81.)

But what I was wondering about is not average attendance, but attendance at a specific game, or equivalently how much the attendance is expected to be above or below the average. If you look at the attendance figures for a given team, their deviation from the average attendance depends on other figures which vary from day to day. First, there are factors that probably affect any outdoor activity, not just baseball:

  • day of week -- weekend baseball is more popular than weekday baseball

  • time of day -- day games and night games are different, although which is more popular probably depends on day of week and the season

  • weather. Pre-purchased tickets probably depend on the average weather; for example, my family has a long-standing policy of not going to day games in June through August, because there's too much risk of the weather being unbearable. I suspect fans in colder-weather markets than Philadelphia avoid games in early April. (Philly can be cold in early April; I wore a winter coat to the Phillies' second game of the season this year.) After the many games canceled due to poor weather in April (including that Indians-Mariners series in Cleveland that got snowed out!),Gregory Goodrich, a meteorologist, examined the frequency of such "miserable baseball days". Walkup tickets, by contrast, probably depend quite heavily on the actual weather. Also, weather matters a lot less in a domed stadium; only in particularly miserable weather people won't even want to leave their homes.


I don't think these three variables could be treated independently; I'd probably prefer a day game in April (the nights can be very cold!) but a night game in July. And people probably care less about weekday versus weekend when school isn't in session.

There are then the factors that are unique to sports:

  • quality of the visiting team. In general I suspect people want to see the home team play a good opponent more than a bad opponent. But at the same time people like to see the home team win. Perhaps people most want to see a visiting team with the same record as the home team, which maximizes the chances of a competitive game. (Note that I'm assuming that the quality of the home team is factored into the average attendance.)

  • identity of the visiting team. When two teams in the same division play each other, I expect the attendance is higher than usual. Similarly, when teams have a long-standing rivalry more people come to games. This is distinct from the "quality of the visiting team" -- Phillies fans are more likely to come down to the ballpark to see the Mets than the Brewers or Dodgers, who have similar records as of this writing. Giants-Dodgers, Red Sox-Yankees, Cubs-Cardinals, etc. games probably sell more tickets than other intradivisional games (Giants-Diamondbacks, Red Sox-Orioles, Cubs-Astros, etc.)

  • time of season. This is distinct from "weather" because while one might expect the same weather in May and September, the two are very different in the context of a baseball season; a game near the end of the season is usually seen to be "more exciting", at least if at least one of the two teams involved has the potential of making the playoffs.

  • playoff importance. Games that are important in determining who ends up in the playoffs should have more attendance; for example, games where the two teams are first and second in their division or in a wild-card race.

I don't know if anyone's done this analysis (and if they have, it wouldn't surprise me if it's a baseball-team front office that isn't talking!) but I'd like to see the results.