Showing posts with label Google. Show all posts
Showing posts with label Google. Show all posts

08 October 2009

Evidence that mathematicians have a big Internet presence

If you google genealogy, the first hit is The mathematics genealogy project. In some sense, according to "the Internet", mathematical genealogy is more important than real genealogy! (I have a feeling that the biological parents of mathematicians would be offended by this, so I won't tell my parents.)

If you google AMS, the first hit is the American Mathematical Society. (Societies of meteorologists, musicologists, Montessori schools, etc. show up further down the list.) I have a musicologist friend that I joke with this about, claiming that the mathematical society is the real AMS.

I suspect this is because mathematicians found the Internet early, and seem to be more likely to have personal web pages than even most other academics.

Edit, 4:18 pm: In a comment by Boris, I'm reminded that Google personalizes search results, so what I've said is not true.

12 December 2008

Where are the mathematicians?

Why are people in Iowa interested in combinatorics? Combinatorics is more popular in Iowa than in any state but Massachusetts.

Google now has a feature called "Google Insights"; you can type in a search term and see where people are searching for it, how frequency of searches varies with time, etc. In states where there's a lot of volume it's possible to zoom in; in Massachusetts it's possible, for example, and most of the interest is in Cambridge. Given that there is a Big Important University and a liberal arts school that has a well-known mathematics department in Cambridge, that's not surprising. But I can't zoom in on Iowa.

(It's possible to get results by country, too, but these results seem ridiculously skewed; I suspect that Google may be normalizing by the number of Internet users in a given area, and the user pool is different in different places.)

Another interesting one: "probability" is popular in Maryland, and among cities in that state it's most popular in College Park and Laurel. College Park is where the University of Maryland is. Laurel is where the NSA is. You can see similar things in other states; for example, in New York, "probability" is most common in Stony Brook, Troy (RPI), and Ithaca (Cornell). In Pennsylvania, it's University Park (Penn State), Bethlehem (Lehigh), and State College (Penn State again). The general pattern seems to be first a few college towns, then the big cities -- the places with the fourth and fifth highest numbers for "probability" in Pennsylvania are Pittsburgh and Philadelphia.

Most mathematical search terms I could think of are highly seasonal -- they're less common in the (Northern Hemisphere) summer, when schools aren't in session. That seems to imply that lots of the people doing the searching are students. I couldn't find a mathematics-related search term that didn't show this seasonality; I don't know if it can be done, because only search terms that receive a reasonably large amount of traffic are reported on the site at all, and things which are important enough to get lots of traffic are probably studied in schools.

07 September 2008

First number not in Google?

Walking Randomly asks: what's the smallest positive integer that gives no Google hits?

I don't know exactly, but it's in the high eight digits. Why, you ask?

Well, I searched a series of numbers, starting with 1 and roughly doubling. This gave the following data:




Search stringNumber of hits
13033319 158
26066638 37
52133277 17
104266555 12
208533110 4

(Each number in the left column is either twice the previous one, or twice the previous one plus one.)

So here's what I'm thinking; numbers around 26 million, have, on average, 37 google hits. Since containing a given number is a Rare Event, I'm guessing that we can treat the number of hits of numbers around that magnitude as Poisson with parameter 37. The probability that a Poisson(37) random variable is equal to zero is exp(-37), or about one in 1016. So the first integer with no Google hits is probably larger than 26 million.

But by the same argument, numbers around 52 million have probability exp(-17), or about one in 25 million, of having no Google hits. So we expect one between, say, 52 million and 77 million.

And by the time we get up near 100 million, we should be seeing these numbers at a frequency of one in exp(12), or six in a million; they should become commonplace.

To get a better model, I'd take more samples. The number of Google hits for a number seems to follow a power law; the number of hits for n is a constant times n for some exponent α, somewhere between 1 and 1.5. (There are issues around saying that things follow power laws, though; it's easy to see them even when they're not there.) And there are various complications -- for example, powers of ten and powers of two are more common than the numbers around them. And how do we know that the occurrence of a number on various web pages is actually independent? To be honest, we don't; if a number exists on a web page, it's there for a reason, and if one person has something to say about it, why not someone else?

11 August 2008

Google ads disappoint again...

Google ad, via gmail: "Quantum Corporate Funding - Quantumfunding.com - Factoring For Contractors Direct Contractor Funding Source".

I was hoping it would be about quantum computers to do integer factoring, say using Shor's algorithm, but something seemed a bit off about the text; it had "quantum" and "factoring" in it, but also words like "funding" and "contractor".

"Factoring" in this context is apparently a banking service that involves selling your accounts receivable (i. e. money people owe your business) in exchange for a (slightly smaller) amount of cash, which you get now. Wikipedia has more. I'm not quite sure why it has that name.

But it wouldn't have surprised me to see somebody saying that they have invented a quantum computer large enough to do non-trivial factorization via a Google advertisement. I'm not saying they'd be right, but I've seen other mathematics/physics crackpottery delivered through that channel.

08 August 2008

A story of Google, Wikipedia, and languages I only sort of read

The Viquipèdia article "Mathemàtiques" has a bunch of amusing pictures that are meant to be icons of different types of mathematics: a Rubik's cube for abstract algebra, a Koch snowflake for fractal geometry, the Lorenz attractor for chaos theory, dice for probability, an elliptic curve for number theory, and so on.

Some areas don't translate into pictures as well: for category theory they have a commutative diagram, for combinatorics the six permutations of [3], etc.

Also, here are Representacions matemàtiques de diversos camps. (The English version of the article does not have this picture currently; they have a picture of Euclid.)

Much of the article seems to be a straight translation of the English version, but I find myself focusing more on the pictures when reading the Catalan version, because I don't actually read Catalan. But I read French and, to a lesser extent, Spanish, so I can figure things out.

(As to why I'm looking at the Catalan wikipedia -- well, I ended up there because Kowalski said he googled 1.70521 when it appeared in some of his work, and I wanted to see the results.))

Such a search finds Wikipedia articles, usually tables of mathematical constants, in Serbian, English, Esperanto, Catalan, Japanese, Thai, Czech, Turkish, Japanese, Serbo-Croatian, and Bosnian. This is both an illustration of the universality of mathematics and the extent to which Wikipedia is an international enterprise. (My apologies if I misidentified any of these languages!)

But the first hit upon Googling 1.70521 (upon this writing) was Kowalski's post, though it's twenty minutes old. That shows you how fast Google is at indexing. (By the time you read this, who knows? This post might be the first hit.)

31 July 2008

A linguistic oddity

Necessary but not sufficient: 596,000 google hits.

Sufficient but not necessary: 134,000 google hits.

Exercise for readers: explain the vast gap here. It seems like the two should be equally common, but when a friend of mine used "sufficient but not necessary" that sounded strange to me; that's what led me to Google, which shows that indeed this phrase is much less common than the reverse.

27 August 2007

scalene triangles on Google

Would you believe that the #1 "hot trend" on Google Trends yesterday was scalene triangle?

I am not making this up.

The Google Trends page for yesterday is actually feeding me a fair bit of traffic right now, to this page in which I solved an unrelated puzzle about scalene triangles.

This would be mystifying to me, but there was a question on "Are You Smarter Than A 5th Grader" last night asking how many angles in a scalene triangle are the same. (Incidentally, I'm not entirely sure whether the answer is zero or one. All the angles are different. There was another question on that show that had two possible answers -- they asked for the world's longest river, whether the world's longest river is the Nile or the Amazon depends on how you measure.)

The #2 "hot trend" yesterday was "mars moons", and #8 was "feet in a mile"; "proper noun" is #21, and "worlds longest river" [sic] is #38. These also related to questions on that show.

The peak time for this search was 4 PM, which seems mystifying until you realize that Google uses Pacific time; this is 7 PM Eastern, which is the time the show aired in the East. Indeed, the four hours when "scalene triangle" was most searched were 4 PM, 7 PM, 5 PM, and 6 PM Pacific (in that order); these are, of course, 7 PM Eastern, Pacific, Central, and Mountain, respectively, which I suspect is the ranking of U. S. time zones from most to least populated. I suspect that it's possible to distinguish between trends that occur because of something on TV and trends that occur because of something that happens in the "real world" (i. e. a breaking news story) by looking at this; breaking news stories should show a single peak, while TV-inspired searches should show a large peak and then a small peak three hours later.

From what I can gather from about Google Trends, the quantity they're ranking is the number of searches done on a particular query in the day in question divided by the number of searches done on a "typical" day. My suspicion is that the various questions on the show were equally likely to be Googled, but more people Google for river lengths or grammatical terminology than scalene triangles on a normal day.