Monday, April 11, 2011

How to write a grant proposal for industry

I've recently had the pleasure of reviewing proposals for Google's Research Awards program. This is a huge program that gives away millions of dollars a year to a large number of university research projects ranging from machine vision to human-computer interaction to mobile systems. After spending eight years in academia struggling to get funding for my own research, it is quite nice to be on the other side of the table and be the one helping to give away the money, rather than begging for it.

First of all, this is nothing like reviewing proposals for, say, the NSF, where you have 15-20 page proposals and project sizes ranging from 3-5 years and anywhere from one to ten PIs. The Google proposals are (thankfully) short -- only 3 pages -- and generally ask for funding for a couple of grad students for a year, plus some funding for a summer month of a PI and maybe some equipment. So the individual grants are small, which in some sense is frustrating since it's hard to propose anything big and groundbreaking when it's just for a year. Essentially this means that most PIs are only asking for money for things they are already working on, rather than spearheading work in a new direction. (Arguably Google should be giving away more money to fewer schools for longer periods of time, but this is my personal opinion.) Also, NSF proposals are reviewed by a panel of (mostly) academics drawn from all over the place, who are supposed to be impartial. At Google, the reviewers are both researchers and software engineers who may or may not be working in areas related to the proposal itself.

Most of the proposals I saw left something to be desired. Far too many of them were asking for money for things that we already know how to do (like write an Android app for your pet project) or which have been done to death (like develop yet another mobile ad hoc routing protocol). On the other hand there were the rare proposals that had an exciting new idea and proposed to do something that Google was not about to go do on its own.

A few tips if you ever apply to this program (and I certainly encourage you to do so).

First, think about who the reviewers are. They are mostly not people like me. Most of the engineers at Google don't have a lot of experience reading and reviewing research proposals (let alone writing them), and many of them are not going to be in your immediate research area. Try to reach out and explain why your work is important at a broader level; the reviewers in this case are typically not your peers.

Second, think about why Google should fund this research. The key question I asked myself was not whether Google would benefit from the research, but rather why should Google fund a university to do this project, rather than just do it ourselves. If Google can hire a couple of engineers to solve a problem, I don't see any reason for us to fund a university to do it instead. On the other hand, if the university PIs are going to do something hard, or groundbreaking, or risky that Google would not have the time or resources to do, we should fund it.

There's also the related question of why should Google fund a research effort rather than another funding agency, such as the NSF. This one is a lot easier to answer: I know from experience that it's damned hard to get money from the NSF for many kinds of projects, and Google can help seed new research efforts that would be difficult to get off the ground otherwise. But if a project seems like it can and should be funded through another agency, that makes it less attractive from my perspective.

Third, try to get the exciting ideas up front. Most Googlers are extremely busy and probably won't spend as much time reading the proposals as you'd like them to. If you bury the lead it will be much harder for the reviewers to see the big idea and get excited about your work. It also helps to establish your credentials in the proposal itself -- not just your CV, but a paragraph or two in the main text saying who you are and why you are the right person to do this research is incredibly helpful (especially in the case when the reviewer is outside your area).

Finally, it always helps to have a champion at Google. If there is someone within the company that you know personally, who can vouch for your work and wants to work with you on a project, this helps tremendously. Having a grad student spend time at Google as an intern is a great way to make those connections.

It's just a guess, but I would not be surprised if other companies' research grants worked in much the same way. While I was at Harvard, I got a lot of funding through places like Microsoft Research, Intel, IBM, and Sun, all of which have fantastic university research programs. (Well, except for Sun, which no longer exists.) Of course, keep in mind that I don't speak for the rest of the Google Research Awards committee, and other reviewers very likely use different criteria than I do.

This is my personal blog. The views expressed here are mine alone and not those of my employer. 

Sunday, April 3, 2011

The death of Intel Labs and what it means for industrial research

Intel recently announced that it is closing down its three "lablets" in Berkeley, Seattle, and Pittsburgh. I know a lot of people who work at the Intel Labs and in fact spent a year at the Berkeley lab before joining Harvard in 2003.  (I should be clear that not all of Intel Research is closing down -- just the lablets.) All of the researchers have been told to find new jobs, though some of them are getting picked up by Intel-sponsored research centers at the nearby Universities.

The Intel Labs were a fantastic experiment to rethink how industrial research should be done. They first started in 2001 under the model that full-time Intel researchers would work side-by-side with faculty and students from the nearby universities. All of the research was done under an open intellectual property model where results were co-owned by the university and Intel. In fact the labs were not inside of the Intel corporate network and operated largely autonomously from the rest of Intel. This allowed projects to be done seamlessly across the Intel/academic barrier and for students to come and go without restrictions on the IP.

Some fantastic work came out of the Labs. The Berkeley lab drove most of the early work on sensor networks and TinyOS, especially while David Culler and his various students were there. The Seattle lab developed PlaceLab (the precursor to WiFi based localization found in every cell phone platform today); WISP (the first computing platform powered by passive RFID); and lots of great work on security of wireless networks. The Pittsburgh lab did work on camera-based sensor networks, cloud computing, and robotics. All of these projects have benefitted tremendously from the close ties that the Labs had with the university.

Before the Labs opened, Intel Research was consistently ranked one of the lowest amongst all major technology companies in terms of research stature and output. I feel that the Labs really put Intel Research on the map by involving world-class academics and doing high-profile projects. They have attracted some of the top PhDs and offered a much more academic alternative to a place like, say, IBM Research.

I have no idea why Intel decided to close the labs. The official press release is devoid of any rationale, and obviously tries to spin the positive angle (the establishment of the university-based research centers which will replace the Labs). I've spoken with a number of the researchers there since the announcement, and have formed my own theories about why Intel is shutting them down. The most obvious possibility is that the Labs are incredibly expensive to run, and it's hard to link the work they do to Intel's bottom line. After all, very little of the work done at the Labs is picked up by Intel's product groups. The Labs' mission has always been to inform the five-to-ten-year roadmap for the company. It's unclear to me whether they have been successful in this, though at least they have inspired some entertaining commercials.

Personally, I'm worried about what this means for industrial computer science research. Here is one of the world's largest and most wealthy tech companies, closing down a set of labs that employs some of the top minds in the field, which by all measures has been really successful in producing novel and high-impact research. If Intel can't figure out how to leverage that amazing talent pool, it does not bode well for the rest of the industry.

Maybe this suggests is that the conventional industrial research model is simply broken. The only (important) places left that use this model are Microsoft, IBM, and HP.  These companies can afford to set up big labs with lots of PhDs and pay them to do whatever the hell they want with little accountability, but maybe this model is no longer sustainable. As I've written before, Google takes a very different approach, one in which there is no division between "research" and "engineering." The advantage is that it's always clear how the research activities relate to the company's priorities, although it does mean that researchers are not doing purely "academic" work, the main output of which is more papers.

One closing thought. Perhaps Intel realizes it can have far more impact by setting up large, high-impact research programs within universities rather than run its own labs. In some ways I can appreciate this point of view: help the universities do what they do best. But the way this is being done is unlikely to be successful. The first such Intel center on visual computing involves something like 25 PIs spread across eight universities. Each PI is only getting enough to fund work they were already doing, so this is an example of doing something that looks good on paper but is unlikely to move the needle at all for these research groups. This seems like a missed opportunity for Intel.

Obligatory disclaimer: This is my personal blog. The views expressed here are mine alone and not those of my employer.

Wednesday, March 23, 2011

Carriers are not ready for tablets

This week I spent at least two hours on the phone trying to convince both AT&T and Verizon to give me online access to accounts I set up for tablets that I am testing -- the Samsung Galaxy Tab (AT&T) and Motorola Xoom (Verizon). They are both great devices; I like the Galaxy Tab's form factor (like a paperback book) and the Xoom is incredibly fast. But it is clear that the wireless carriers have no idea how to incorporate these devices into their billing and customer service ecosystem. It was such a painful and frustrating experience that I wonder how the cellular carriers expect to leverage these devices as more tablets come onto the market.

First, my story with AT&T. I bought the Galaxy Tab a few months ago which came with an AT&T SIM card pre-installed. When you boot the device for the first time, there's a widget which takes you to a registration page, which I filled out to activate the tablet on AT&T's network. Since then I have not received a bill (to my knowledge) for the usage, and couldn't remember whatever password I might have used to set up the device. Nothing on AT&T's website seemed to offer any help, as it is completely oriented towards phones.

Finally, I gave up and called AT&T to re-register the Tab manually. The first customer service rep had no idea how to do this. They kept asking for the phone number, which the device does not have (at least when you enter the Settings menu it lists the phone number as "unknown"). Given that I had not received any bills I suspected the Tab was never registered, so we had to look it up by IMEI number. The rep could not pull up any account information. She ended up transferring me to technical support at Samsung, of all places. The Samsung rep was very friendly but couldn't help with this problem, either -- it seemed to be an AT&T issue (and I agree). I ended up having to call AT&T back, and went through the same painful process of explaining what my problem was. This rep ended up transferring me over to a different sales rep who tried to help me set up the account from scratch.

This is when things started to go downhill. All I wanted as an unlimited (or as close as possible) data plan for the Galaxy Tab. I could see online that AT&T has a 2GB/month data plan for tablets but the rep kept telling me that "their internal systems don't necessarily show what is on the website." (First warning sign there.) Eventually he managed to pull up the right plan but couldn't seem to figure out how to add a Galaxy Tab -- the device wasn't showing up in his menus. It sounded like he had never activated a tablet before. After around 20 minutes on hold he managed to figure it out, so I think I finally have the Galaxy Tab set up for data access. I was promised that I would get an email from AT&T confirming the new account, but it never arrived. So I guess I am going to have to call them back. I am dreading this.

Verizon was almost as bad. Like the Tab, I had set up the Xoom using the registration app on the tablet itself. I had made a note of the username used to set up the account, but not the password. Verizon's website offers no ability to request a new password except via SMS to the device -- and the Xoom doesn't receive SMS messages, since it's not a phone. The only way to request a new password is to spend around 45 minutes on the phone with various Verizon reps with the result being that a new password is being sent to me by postal mail in five business days. (What is this, the nineteenth century?) Of course, since I moved recently, my mailing address on file with Verizon was incorrect. Fortunately, the service rep allowed me to change the postal address over the phone -- meaning that they trust me enough to let me change my mailing address, not not enough to reset my online account password. This makes absolutely no sense and seems designed to drive users away.

The lesson here is that the wireless carriers have no clue how to incorporate tablets. They are treating them like phones, which they aren't.

Wednesday, March 16, 2011

Thinking back on 8 years in Boston

Tomorrow I will be packing up and moving from Boston to Seattle with my family. I thought now would be a good time to reflect on living in Boston as a city and recall some of my best memories here.

I've never lived in Seattle, though have been there many times -- it seems like a wonderful city, full of funky crazy people and absolutely beautiful geography. I'm not terribly excited about the rainy weather, though something tells me it can't be any worse than the Boston winters, when I always feel cooped up. I really miss getting out to go hiking with the dog or mountain biking during the winter months in New England -- and now that I have a kid it's especially hard to get out when it's well below freezing outside (he has a lot lower tolerance for the winter weather than I do). Rain I can deal with; negative 20 wind chills and a foot and a half of snow are something different altogether.

On finishing grad school at Berkeley in 2002, I had a few faculty job offers, and my wife was looking for residency programs in psychiatry. Our decision came down to two cities: Boston, where I had an offer at Harvard, and Pittsburgh, where I had an offer at CMU. I had lived in Boston for a while during college, so I knew I liked the city. But CMU was a very tempting offer, being a much more highly-ranked CS program than Harvard. My wife and I visited Pittsburgh a couple of times and actually liked it a lot: it was a very friendly place, and the CMU folks went all out to show us a good time. At one point we actually made the decision to move to Pittsburgh but decided to sleep on it. The next day we had to ourselves, without anyone showing us around. The only Mexican place that served "fish tacos" appeared to be Van De Kamp's fish sticks on a tortilla. I could not find a music store that allowed you to browse the CDs without an attendant unlocking a glass case to let you inside. We tried to find a decent shopping mall, hoping they would have a good record store, but found ourselves in the mall where the movie Dawn of the Dead was filmed -- I did not make it more than three paces beyond the door before we realized it was a terrible mistake. I'm sorry, Pittsburgh might be a wonderful city for some people, but it was not for us.

Monroeville Mall on a typical Sunday afternoon.
So we moved to Boston in 2003. We drove across the country in my little beat-up two-door Ford, stopped at places like Moab and St. Louis and Nashville, a little like On the Road in reverse. We arrived on a hot, sticky summer day in a thunderstorm and moved into an apartment on Dana Street in Cambridge, not far from Harvard Square. Almost immediately we felt like outsiders. Boston is a very old city, and it shows -- the old brick buildings around Harvard, the beat-up sidewalks, everything dripping with history and significance. Paul Revere. George Washington. Boston Common. Old Granary Burial Ground. The feeling was so different than the newness that is so pervasive in California. It took some getting used to.

These gravestones have been here for a while. (From http://www.flickr.com/photos/harvardavenue/66097831/)
That first summer was one of the most exciting in Red Sox history. I had never paid attention to baseball before, but it soon became clear that if I wasn't conversant in Curt Schilling or Manny Ramirez I was going to be left out of a lot of conversations. Like a lot of people in Boston, I got caught up in the excitement of the 2003 postseason when the Sox were narrowly beat by the Yankees for the ALCS title. The next year was even more exciting, in that the Sox went on to win the World Series for the first time in 86 years. The whole city went totally apeshit. For so many people living here, it was like a moment of rapture that they never believed would actually happen, the culmination of a lifetime of inferiority to cities like New York. I feel sad for some long-time Bostonians who now have no excuse to be cynical about their station in life.

People were just a little excited. (From http://www.flickr.com/photos/eandjsfilmcrew/412103968/)
Not long after we moved to Boston they finally banned smoking in bars and restaurants (which we'd been used to in California) and actually allowed alcohol sales on Sundays. It was as if the city were modernizing before our eyes. It would take another six years for Cambridge to allow outdoor dining at restaurants, which is still encumbered with backwards vestiges of the Puritans (you can only have a drink while sitting outside if you also order food).

Going out in Boston was a bit more a dressy, formal affair than we were used to in Berkeley. At all but the very highest end restaurants in San Francisco, jeans and a t-shirt were acceptable attire; not so in Boston. On the positive side, Boston has a great foodie scene. At first we were totally lost trying to find good places to eat; Zagat's is useless and Yelp simply reflects the lowest common denominator. At first we were convinced that people in Boston had no idea what real, good Mexican food was -- those greasy platters of "enchiladas" covered in melted cheese that are so popular at hellholes like Casa Romero are not it. Then we discovered Chowhound, and got turned on to a whole world of hole-in-the-wall places serving authentic Mexican and Salvadorean and Sichuan and Cambodian. Unlike Berkeley, where you can swing a dead cat and hit three burrito joints and a place with out-of-this-world sushi, in Boston it takes a bit more digging, but there is great food here.

Some of my best memories of living in Boston...

Walking my dog, Juneau, to work at Harvard every day, stopping at the coffee shop on the way, and taking her to the dog park on the way home for an hour or so of play time with the other dogs.

Juneau would occasionally help me reviewing papers, too.

Sitting outside on a warm summer evening, firing up the grill, having friends over for dinner and drinks until late.

Every single fall, feeling the first day of cold air and getting excited for the leaves to start changing.

This image has not been enhanced.

Riding my bike along the banks of the Charles River, whizzing by rollerbladers and clueless tourists walking four abreast in the middle of a bike path.


The morning after a big snowstorm, seeing the world transformed and noticing how quiet everything was with the snow on the ground.

Shoveling is always fun too.


Late nights out in Boston Chinatown with Gu, drinking beer and eating Korean food.


Watching the sun rise out of the window of Mount Auburn hospital on a hot July day a couple of hours before my son was born.

I was a dad not long after this picture was taken.

Friday, March 11, 2011

Running a successful program committee

Yesterday we held the program committee meeting for the 13th Workshop on Hot Topics in Operating Systems (HotOS), for which I am serving as the program chair. This is the premier workshop in the OS community and focuses on short (five page) position papers meant to bring out exciting new research directions for the field. In some years it has been more exciting than others. What tends to happen is that people send five-page versions of an SOSP submission they are working on, which (in my opinion) is not the best use of this venue. When HotOS becomes an SOSP preview I think it misses an important opportunity to discuss new and crazy ideas that would not make it into a regular conference.

We accepted 33 out of 133 submissions. The number of accepted papers is a bit higher in previous years because I wanted to be more inclusive, but also recognized that we could fit more presentations in at the workshop when you don't have 25-minute talk slots. There is no reason a 5-page paper needs a 25-minute talk, and I think it goes against the idea of the workshop to turn it into a conventional conference.

This was, by far, the best program committee experience I ever had, and I was reflecting on some of the things that made it so successful.

I've been on a lot of program committees, and sometimes they can be a very painful experience. Imagine a dozen (or more) people crammed into a room, piles of paper and empty coffee cups, staring at laptops, arguing about papers for 9 or 10 hours. Whether a PC meeting goes well seems to be the result of many factors...

Pick the best people. We had a stellar program committee and I knew going in that everyone was going to take the job very seriously. Everyone did a fantastic job and wrote wonderful and thoughtful reviews. These folks were invested in HotOS as a venue, were the kind of people who often submit papers to the workshop, and care deeply about the systems community as a whole. The discussion at the PC meeting itself was great, nobody seemed to get cranky, and even after 8+ hours of discussing papers there was still a lot of energy in the room. This is helped a lot by the content -- HotOS papers tend to be more "fun" and since they are so short, you can't nitpick every little detail about them.

Set expectations. I tried to be very organized and made sure that the PC knew what was expected of them, in terms of getting reviews done on time, coming to the PC meeting in person, what my philosophy was for choosing papers, and how we were going to run the discussion at the meeting. I think laying out the "rules" up front helps a lot since it keeps things running smoothly. I've blogged about this before but I think establishing some ground rules for the meeting is really useful.

Get everyone in the room. Having a face-to-face PC meeting is absolutely key to success. Everyone came to the PC meeting in person, except for one person whose family fell ill at the last minute and had to cancel, but even he phoned in for the entire meeting (I can't imagine being on the phone for more than eight hours!). I made sure the PC knew they were expected to come in person, and nailed down the meeting date very early, so everyone was able to commit. Letting some people phone in is a slippery slope. I can't count how many PC meetings I've been to that have been hampered by painful broken conference call or Skype sessions.

Use technology. We used HotCRP for managing submissions and reviews, which is by far the best conference management system out there. During the PC meeting itself, I shared a Google spreadsheet with the TPC which had the paper titles, topic area, accept/reject decision, and a one-line summary of the discussion. The summary was really helpful for remembering what we thought about a paper when revisiting it later in the day. Below is a snippet (with the paper numbers and titles blurred out). The "order" column below is the order in which the paper was discussed. This way, everyone in the PC could see the edits being made in real time and there was rarely confusion about which paper we were discussing next.


Pre-reject and pre-accept. I rejected around half of the submissions before the PC meeting and gave the PC a chance to revive any such paper for discussion (none were). I also "pre-accepted" about 10 papers that were noncontroversial; we saved discussion of these for the end of the day, since they were easy cases. We ended up discussing a total of 69 papers at the meeting, which meant we had to go at a pretty good clip.

Be definitive.  With very few exceptions, we tried to reach a clear accept/reject decision on each paper as we discussed it, and did not table any papers for later discussion. There was one case where we were hung on what to do with a paper and decided to push the discussion until the end of the day. In cases where there was disagreement, I would mark a paper as "presumed reject" or "presumed accept" and put down the name of the person who wanted to argue for the opposite outcome later. That gave us a chance to move on when there was an insurmountable debate, and it was clear that the champion (or anti-champion) of a paper would have a chance to have their say.


Take everyone out to a nice dinner afterwards. As far as I'm concerned, this was the best part of hosting the PC meeting.

Thursday, February 24, 2011

What life was like before the Web

It's kind of shocking that the Web has only been around since the mid-1990s, but a lot of younger people that I work with have no idea what it was like to use the Internet before the Web. I'm not that much of an old-timer, but I thought it would be amusing to talk about the pre-HTTP Internet. A few reminisces of the good old days...

  • Everything was text-based;
  • There was (almost) no spam;
  • There was no such thing as a search engine;
  • You did everything using these crappy UNIX text-based command-line tools.

I first started using the Internet back in high school, around 1990, on an IBM RT UNIX system connected via dialup from school. Back then, there were three main uses of the Internet: email, USENET, and FTP. Email was not very common but some universities had it, and by the time I started college in 1992, you automatically got an email address as an undergraduate.

USENET

USENET was a huge distributed newsgroup system. It was based on peer-to-peer file exchange well before we called it that. There were thousands of newsgroups on topics ranging from the C programming language to the rock band Rush. Yes, it still exists; I think it's largely overrun with spam and warez these days, and I haven't looked at it in years. At the time, spam was almost unheard of so newsgroup discussions tended to stay on-topic. I was the moderator for comp.os.linux.announce for a while, which meant that every announcement to the Linux community (like a new kernel release or software package port) came to my inbox for approval before I posted it to the world.

At some point I want to write a book about USENET culture circa 1992. It was a very interesting place. Groups like talk.bizarre were frequented by the likes of Roger David Carrasso (who pulled off some of the best trolls imaginable); Kibo (founder of alt.religion.kibology); and who can forget the utterly brilliant and bizarre stories by RICHH?

I was a member of the "inner circle" for a group called alt.fan.warlord, which was centered on making fun of ridiculous signatures at the end of USENET posts, like this:


     Paul Tomblin, Head             _            _   ____
     Automated Test Tools Team     | |          | | |  __|   ___._`.*.'_._
    _______________   ______   ____| |________  | |_| |__   +  * .o   u.* `
   /  ________  _  \ |  __  | /  ________  _  \ |  ______|  . ' ' |\^/|  `.
   | |  | |  / / | | | |  | | | | |  |  / / | | | | | |            \V/
   | |__| | / /__| |_| |  | | | |_|  | / /__| |_| | | |            /_\
   \ _____/ \__________|  |_|  \___|_| \__________| |_|       === _/ \_ ===
   //
   \\____   Phone: (613) 723-6500x8018      Mail: Gandalf Data Limited
   /  _  \  Fax:   Don't know it yet              130 Colonnade Road South
   | |_| |  Email: ptomblin@gandalf.ca            Nepean, Ontario
   \_____/   or    ab401@freenet.carleton.ca      K2E 7J5 CANADA



There was a group called alt.hackers that was a moderated group with no moderator. In order to post you needed to figure out how to circumvent the moderation mechanism (which was simply a matter of adding an extra header line to your post).

FTP

USENET was all about discussions, although there were newsgroups where you could post binary files -- typically encoded in an ASCII format like UUENCODE, and broken up into a dozen or more individual posts that you would have to manually stitch back together and decode. This became a popular way to post low-resolution porn GIFs, but was pretty much useless for anything larger than a few megabytes. A better way to download files was to use FTP, which allowed you to download (and upload) files to a remote FTP server. Of course, FTP had this totally unusable command-line interface which required you to type a bunch of commands just to get one file. This site at Colorado State helpfully explains how to use FTP, like anybody still needs to know how. A typical FTP session looked like this:


%   ftp cs.colorado.edu  
Connected to cs.colorado.edu.  
220 bruno FTP server (SunOS 4.1) ready.  
Name (cs.colorado.edu:yourlogin): anonymous  
331 Guest login ok, send ident as password.  
Password:  
230-This server is courtesy of Sun Microsystems, Inc.  
230-  
230-The data on this FTP server can be searched and accessed via WAIS, using  
230-our Essence semantic indexing system.  Users can pick up a copy of the  
230-WAIS ".src" file for accessing this service by anonymous FTP from  
230-ftp.cs.colorado.edu, in pub/cs/distribs/essence/aftp-cs-colorado-edu.src  
230-This file also describes where to get the prototype source code and a  
230-paper about this system.  
230-  
230-  
230 Guest login ok, access restrictions apply.  
ftp> cd /pub/HPSC  
250 CWD command successful.  
ftp> ls  
200 PORT command successful.  
150 ASCII data connection for /bin/ls (128.138.242.10,3133) (0 bytes).  
ElementsofAVS.ps.Z  
   . . .
execsumm_tr.ps.Z  
viShortRef.ps.Z  
226 ASCII Transfer complete.  
418 bytes received in 0.043 seconds (9.5 Kbytes/s)  
ftp> get README  
200 PORT command successful.  
150 ASCII data connection for README (128.138.242.10,3134) (2881 bytes).  
226 ASCII Transfer complete.  
local: README remote: README  
2939 bytes received in 0.066 seconds (43 Kbytes/s)  
ftp> bye  
221 Goodbye.  


All this just to download a single README file from the site.

In order to use FTP, you needed to know what FTP sites were out there and manually poke around each one of them to see what files they seemed to host, using "cd" and "ls" commands in the crappy command-line client. There was no such thing as a search engine. So, people in the FTP community made this giant "master list" of every FTP site out there and a short (one-line) summary of what the side had, e.g., "Linux, UNIX utils, GIFs." This master list was mirrored on a bunch of sites but of course that was a manual process, and in effect the mirrors were often conflicting and out-of-date. These days you can find a giant list of FTP sites on the Web, of course.

The Birth of the Web

Back in 1994 or so I was doing research as an undergraduate at Cornell. At the time I was a huge USENET junkie, and from time to time I would see these funny things with "http://" in people's signatures. I had no idea what they were, but at some point saw a USENET post that explained that if you wanted to browse those funny "URL" things you needed to download something called Mosaic from an FTP site at NCSA. When you launched the Mosaic Web browser it brought up this page which was the home page for the entire World Wide Web. At some point I managed to get the original NCSA Web server running and put Cornell's CS department on the Web. Here's an archive of the Cornell Robotics and Vision Lab web page that I made back around 1995.

Not long after this a startup from Stanford called Yahoo created a (manual) index of every web page -- still not quite searchable, but at least you could find things. I remember using the original Google when it was hosted at Stanford and had a whopping "25 million pages" indexed.

Saturday, January 22, 2011

Does Google do "research"?

I've been asked a lot by folks recently about whether the work I'm doing now at Google is "research" and whether one can really have a "research career" at Google. This has also led to a lot of interesting discussions about what the role of research is in an industrial setting. TL;DR -- yes, Google does research, but not like any other company I know.

Here's my personal take on what "research" means at Google. (Don't take this as any official statement -- and I'm sure not everyone at Google would agree with this!)

They don't give us lab coats like this, though I wish they did.
The conventional model for industrial research is to set up a lab populated entirely by PhDs, whose job is mostly to write papers, and (in the best case) inform the five-to-ten year roadmap for the company. Usually the "research lab" is a separate entity from the product side of the company, and may even be physically remote.

Under this model, it can be difficult to get anything you build into production. I have a lot of friends and colleagues at places like Microsoft Research and Intel Labs, and they readily admit that "technology transfer" is not always easy. Of course, it's not their job to build real systems -- it's primarily to build prototypes, write papers about those prototypes, and then move on to the next big thing. Sometimes, rarely, a research project will mature to the point where it gets picked up by the product side of the company, but this is the exception rather than the norm. It's like throwing ping pong balls at a mountain -- it takes a long time to make a dent.

But these labs aren't supposed to be writing code that gets picked up directly by products -- it's about informing the long-term strategic direction. And a lot of great things can come out of that model. I'm not knocking it -- I spent a year at Intel Research Berkeley before joining Harvard, so I have some experience with this style of industrial research.

From what I can tell, Google takes a very different approach to research. We don't have a separate "research lab." Instead, research is distributed throughout the many engineering efforts within the company. Most of the PhDs at Google (myself included) have the job title "software engineer," and there's generally no special distinction between the kinds of work done by people with PhDs versus those without. Rather than forking advanced projects off as a separate “research” activity, much of the research happens in the course of building Google’s core systems. Because of the sheer scales at which Google operates, a lot of what we do involves research even if we don't always call it that.

There is also an entity called Google Research, which is not a separate physical lab, but rather a distributed set of teams working in areas such as machine learning, information retrieval, natural language processing, algorithms, and so forth. It's my understanding that even Google Research builds and deploys real systems, like Google’s automatic language translation and voice recognition platforms.

(Update 23-Jan-2011: Someone pointed out that Google also has a "Quantitative Analyst" job role. These folks work closely with teams in engineering and research to analyze massive data sets, build models, and so forth -- a lot of this work results in research publications as well.)

I like the Google model a lot, since it keeps research and engineering tightly integrated, and keeps us honest. But there are some tradeoffs. Some of the most common questions I've fielded lately include:

Can you publish papers at Google? Sure. Google publishes hundreds of research papers a year. (Some more details here.)You can even sit on program committees, give talks, attend conferences, all that. But this is not your main job, so it's important to make sure that the research outreach isn't interfering with your ability to do get "real" work done. It's also true that Google teams are sometimes too busy to spend much time pushing out papers, even when the work is eminently publishable.

Can you do long-term crazy beard-scratching pie-in-the-sky research at Google? Maybe. Google does some crazy stuff, like developing self-driving cars. If you wanted to come to Google and start an effort to, say, reinvent the Internet, you'd have to work pretty hard to convince people that it could be done and makes sense for the company. Fortunately, in my area of systems and networking, I don't need to look that far out to find really juicy problems to work on.

Do you have to -- gulp -- maintain your code? And write unit tests? And documentation? And fix bugs? Oh yes. All of that and more. And I love it. Nothing gets me going more than adding a feature or fixing a bug in my code when I know that it will affect millions of people. Yes, there is overhead involved in building real production systems. But knowing that the systems I build will have immediate impact is a huge motivator. So, it's a tradeoff.

But doesn't it bother you that you don't have a fancy title like "distinguished scientist" and get your own office? I thought it would bug me, but I'm actually quite proud to be a lowly software engineer. I love the open desk seating, and I'm way more productive in that setting. It's also been quite humbling to work side by side with these hotshot developers who are only a couple of years out of college and know way more than I do about programming.

I will be frank that Google doesn't always do the best job reaching out to folks with PhDs or coming from an academic background. When I interviewed (both in 2002 and in 2010), I didn't get a good sense of what I could contribute at Google. The software engineering interview can be fairly brutal: I was asked questions about things I haven't seen since I was a sophomore in college. And a lot of people you talk to will tell you (incorrectly) that "Google doesn't do research." Since I've been at Google for a few months, I have a much better picture and one of my goals is to get the company to do a better job at this. I'll try to use this blog to give some of the insider view as well.

Obligatory disclaimer: This is my personal blog. The views expressed here are mine alone and not those of my employer.

Startup Life: Three Months In

I've posted a story to Medium on what it's been like to work at a startup, after years at Google. Check it out here.