Monday, April 14, 2008

Crawling the Deep Web

Our work on crawling the Deep Web has received some attention over the last few days. It started with a post on Google's Webmaster blog. Judging by the number of in-links to the blog (see the bottom the page) and the several news articles that picked it up, there were quite a few reactions on the blogosphere and beyond.

Matt Cutts, Google's main interface to web masters gives a nice explanation of why this work is useful to site owners. Anand Rajaraman details some of the history behind the technology that led to this work.

In summary, a nice example of research on data management having impact on the Web.

Wednesday, April 2, 2008

Bar-coding in Costa Rica

I spent last weekend in the Area de Conservacion Guanacaste (ACG) in Costa Rica with a few of my colleagues. We were hosted by Dan Janzen and Winnie Hallwachs (his wife), an incredibly inspiring pair of biologists. Among many other awards, Dan is also a recipient of the Kyoto Prize in 1987. For the past 30 years, Dan and Winnie have spent half of every year in Costa Rica, creating the ACG, while spending the other half professing at the University of Pennsylvania. I'll have to skip the details of how we got there, but do ask me in person when you see me (and if you need to juggle my memory, use the phrase "party in the sky").

You can find the pictures from the trip here.

So you're probably wondering what was a guy like me, with questionable credentials in Biology, is doing in such a biologically intense area?

Imagine that every living species and plant had a barcode, just like products in a supermarket. Furthermore, imagine that you had a device, the size of a cell phone, such that when you found a specimen in the forest, you can put the specimen into the device and it would tell you all the known information about it. In addition to being a useful device to take on hikes, such a device can have major impact on agriculture and controlling the spread of disease.

The International Bar-Code of Life Project (iBol) is trying to do exactly that, based on genomic techniques. Specifically, it turns out that with over 98% accuracy, the CO1 gene uniquely determines the species. In contrast, the traditional approach to determining species is based on morphological features. By sequencing the CO1, Janzen and many others have been able to uncover several mysteries, showing that species that look very similar are actually different, and vice versa. Janzen runs the biggest specimen collection operation (Costa Rica happens to have a huge number of different species, hence Janzen's conservation goal). Currently he sends them to the University of Guelph in Canada for sequencing (in a lab run by Paul Hebert who was also there), but they envision that in a decade, we'll be able to build the small device.

We spent the weekend in numerous and intense discussions on Biology, walking through the forest seeing it first hand, and actually participating in the process of collecting specimens and preparing them to be sent for sequencing.

In the discussions we tried to understand the challenges involved in this project (including arguments by its critics). It actually turns out that determining species can often be very subjective, for two reasons. First, the determination typically needs to be done with only partial information about the set of specimens available and unless you can find other evidence, morphology is typically the deciding factor. Second, and somewhat more surprising to me, not all biologists completely agree on what the concept of species even means. The most accepted definition is based on the ability to mate and create viable offsprings, but there are other opinions as well (e.g., it's the morphology stupid). In fact, when it's not even clear (to me, at least) that classification into species is as important as it's traditionally been considered, since many of the questions we're asking about animals or plants depend on other genetic and environmental traits.

And yes, there are huge data management challenges here. Many scientists are collecting data and each putting it into their own format. They would like to share their data but also maintain control of their own. They'd like to publish the data on the web and make it accessible to the masses. They need to manage uncertainty and provenance. Ironically, one of the closest systems I know that is considering some of these issues is Orchestra, built by Zack Ives at the... University of Pennsylvania (i.e., a few buildings away from Janzen's office).

Then there was the flight back, but I can't talk about that either. Overall, an incredible experience! Many thanks to Dan and Winnie (and their crew) for hosting us and sharing their incredible knowledge and passion!

Sunday, March 16, 2008

Two Books by Geraldine Brooks

I just finished reading two books by Geraldine Brooks and highly recommend them. Both books are fiction inspired by true historical events. In both cases, Brooks manages to vividly recreate the periods in which the plot is taking place and bring them back to life. The research that goes into her books is really impressive.

The first book is Year of Wonders. It is based on the story of the little village of Eyam in Derbyshire, England in 1666. That was the year plague swept through the village devastating it. The story tells of how the plague entered the village and its effects on its inhabitants through the story of Anna Frith, a housemade at the village's rectory.

The second book, People of the Book, was published just this year. It is based on the story of the Haggadah of Sarajevo. The Haggadah is the story of the exodus of the Israelites from Egypt, and is read on the eve of Passover (with a lot of food and a minimum of 4 glasses of wine weaved in). The history of this Haggadah goes back to 15th century Spain and has an amazing story of survival through Venice, Vienna and Sarajevo (at least). One of the most interesting aspects of its story is that the acts (often of heroism) to save the Haggadah were typically carried out by non-Jews -- Muslims or Christians, who appreciated the value of the book. The book portrays vividly several periods in history some of which had Christians, Muslims and Jews were living in peace together (Spain, before the inquisition). It's really a great read (regardless of one's religion). And yes, there was even a Halevy involved in this book's history!

Monday, January 28, 2008

A "Web Moment"

I'm sure each and every one of you has had at least one "web moment", where the power of the web simply jumped out at you. It may have been after a web-search yielded an amazing result, or realizing that you're driving to a dinner meeting at http://... using directions you got from online maps, and checking traffic conditions on the web from your car. (I do realize, however, that there is an entire generation out there who has no idea what the heck I'm talking about, and equates the pre-web world roughly with the 19th century).

I (or rather, my dad) had such a moment yesterday, when he found a document on the web, signed by his father in 1956, that he had no idea existed.

Yad Vashem, the Holocaust Museum in Jerusalem, collected a database of people who were killed in the holocaust. They asked anyone who knew holocaust victims to report their details and contribute them to this database. I found the web site for searching that database (thanks to Yair Kurzion), and told my dad about it.

His task was not easy. Searching for a Levy in database of Jews is like searching for Smith in a phone book of a big American city, and even restricting the search with the first name and the city of origin did not help a lot. But magically, he pulled up the first result and found a document signed by my grandfather (who passed away 40 years ago). My grandfather had reported the death of his brother and sister, who were both deported from Thessaloniki, Greece to the gas chambers in Poland in 1942, along with the vast majority of the vibrant Jewish community of that city. Fortunately, my grandfather, who was a Zionist at heart, left Thessaloniki in 1933 for Tel-Aviv to later be part of the creation of the State of Israel.

Friday, December 28, 2007

My Dad is 80!





My dad turned 80 this month, and we celebrated the event with a workshop and reception in his honor at the Weizmann Institute of Science, in Rehovot, Israel. The full set of pictures from the event can be seen here.

My dad is a professor of Chemistry at the Weizmann Institute. After fighting in the Israeli War of Independence, he was finally able to focus on his studies. He completed his Ph.D in a little less than 2 years(!!) in 1955 at Syracuse University in New York (fortunately, because that's where he met my mom). When he's asked how he did that, he simply shows the picture below.




He still doesn't understand why it took me an entire 5 years to do my Ph.D, and worse, in a field that uses the term 'science' in a questionable fashion. (When I got promoted to full professor he finally figured I might be doing something right).

My dad's main claim to fame is a 1-page article he wrote during his post-doc. The following makes the point better than I can - it's a quote from Krzysztof Matyjaszewski (a CMU professor) and Axel Muller (professor at the U. of Beyreuth, Germany) in their foreword to the December 2006 special issue of the Journal of Progress in Polymer Science on "50 Years of Living Polymerization":


On June 5, 1956, Michael Szwarc, together with Moshe Levy and Ralph Milkovich pubished an article entitled "Polymerization initiated by electron transfer to monomer - A new method for formation of block copolymers", J Am Chem Soc (1956), 2656.

In this article the term "living polymer" appeared for the first time. It caused a revolution in polymer science.


Michael Szwarc (who was also my dad's Ph.D adviser) received the Kyoto Prize for this work in 1991.

He has worked in many areas over the years, but since the mid-80's my dad has been one of the pioneers in solar energy research, studying methods for chemical storage of solar energy so it can be used any time and transported to less sunny locations. In fact, he published two papers on using solar energy for chemical transformations this year! As a befitting token or recognition, he received an awesome Google solar t-shirt...

And he definitely needs the t-shirt. He still gets up every morning at 6am to either play tennis, or go for a run & workout, which includes running up 15 flights of stairs in the solar tower at the institute!


It was a great event, and a wonderful chance to see many of my dad's colleagues throughout his career, some of whom I had not seen since I was a kid. It was also the first full gathering of all the family's grandchildren.

Sunday, November 11, 2007

Dataspaces for Veterans

In the U.S. we are marking Veteran's Day tomorrow. There are many ways in which we should be thanking our veterans and making their lives better. I'd like to report a rather unique way.

I recently had the opportunity to visit the Veteran's Administration Hospital in Washington DC and learn first-hand about their patient-record system. I was pleased to see the principles of dataspaces in action, clearly enabling better healthcare services.

The VA provides services to veterans of the American military and has around 150 hospitals, 800 clinics and 200 nursing centers scattered around the country. To support these services, the VA maintains electronic records for all their patients, a system that has won them many accolades in recent years. The system stores the patients' prescriptions, doctor visits, lab tests and other data about each patient. As their patients often move around and receive treatment in various locations, when a doctor views the data about a patient, it needs to be integrated from multiple VA locations. Each of these locations is running their own system. In addition, data about their patients may reside in systems of the Department of Defense (and their healthcare providers) and various drugstore chains.

Clearly, this is an incredible data integration problem. Today they are aware of at least 130 different "implementations" of their electronic record system, i.e., different schemas. Also, given the different local needs of hospitals and clinics, imposing a single schema on all the VA centers would not work. Using a data integration solution at this scale and in such a dynamic environment would be extremely difficult.

Instead, what the VA did is standardize on a very small subset of patients' attributes, namely attributes describing patients' vital signs. Outside of this set of attributes, hospitals are free to develop their own local data organizations. However, the system lets the healthcare providers see all the data even if it's not completely integrated. So for example, if a doctor wants to see what happened to a patient while they were at a remote location, then the remote data may appear as plain text, and therefore the doctor would have to work a little harder to digest it, and won't be able to pose the queries she could pose on the local data. But being able to see the data in some form is infinitely better than not seeing it at all, and the doctors are extremely happy with the system's capabilities.


The VA also demonstrated two examples of the pay-as-you-go principle that is at the foundation of dataspaces. The first was the fact that they decided that vital signs are critical, so their data sources are aligned on the attributes relating to those (effectively, creating semantic mappings involving the attributes of vital signs), and they plan to continue agreeing on terminology as they see fit. Second, they had a culture that allowed for local innovation, class-3 applications, that represented needs at the local level. When these needs were perceived to be important throughout the organization, they promoted them to class-1 applications, and required all their systems to support them.

Just to make it clear, when I walked in the door they did not greet me and say: "Pleased to see you Dr. Halevy; we'd love to show you our dataspace system". What I'm describing is a post-rationalization of a system that was developed over more than a decade. I believe that their loose integration was the key to their success.

Wednesday, October 24, 2007

A Murder Mystery with a Twist

I just finished reading dot.dead, a Silicon Valley murder mystery by Keith Raffel. Yes, a murder happening in Palo Alto at the home of Ian Michaels, a high-tech executive; searching for clues on Stanford campus and running the dish to think deep thoughts and unravel the mystery.


Will Ian Michaels be indicted and spend the rest of his life in jail? Or perhaps he will be promoted to that COO job he's been eying for a while, or even leave his company and start his own? And in the process, how many eligible (or non-eligible) women will try to seduce him?

Read the book and find out. Not bad for an author who used to be a high-tech guy himself.