Showing posts with label Information Retrieval. Show all posts
Showing posts with label Information Retrieval. Show all posts

Sunday, May 11, 2008

Powerset Natural Language Search


Powerset, a company we only remember seeing as conference sponsors, now actually has something working. After receiving an email from them, I tried out several queries. At best, it seems to answer most Wh-questions and certain whole-part relations.

Try out the same query on Google.

Thursday, August 2, 2007

Recommending scientific papers

I noticed a new feature in Citeseer which tries suggest an "alternate document" for a paper.
Clearly it does not do what it implies to do and it doesn't show up for all papers. (Experimental?) So, an interesting question is how does one recommend scientific papers? Something more than mere document similarity is required. If I am reading a CRF paper then there is no point in listing all papers containing similar words. Just listing nodes connected to inward and outward links of the paper in the citation graph wont suffice either. Ideal recommendations for a paper would depend on the role the user is playing. When I am reading a paper about some new topic, I would like to get pointed to original papers on the topic, some recent papers on the topic, and may be some survey papers or books. On the other hand when I am writing a paper, I would like to be pointed to all papers related to the topic (recall important than precision here to avoid reviewer comments on "missing reference") in some magical order that puts papers more relevant to your work above. Also these papers might not be related in directly through citations. If there is a recent related work in the Annals of Statistics, for instance, then it should show up when I am working on, say, approximate inference methods for graphical models. (Possible to deduce this from my previous queries?)

In spite of more information being present in a scientific paper than its text, recommending or ranking papers appears to be quite challenging.

Tuesday, July 24, 2007

Readings from SIGIR 2007

SIGIR 2007 is happening now at Amsterdam!

Latent Concept Expansion Using Markov Random Fields, Donald Metzler, Bruce Croft

Random Walks on the Click Graph, Nick Craswell, Martin Szummer

Towards Automatic Extraction of Event and Place Semantics from Flickr Tags, Tye Rattenbury, Nathaniel Good, Mor Naaman

Clustering of Documents with Local and Global Regularization, Fei Wang, Changshui Zhang, Tao Li

Detecting, Categorizing and Clustering Entity Mentions in Chinese Text, Wenjie Li, Donglei Qian, Chunfa Yuan, Qin Lu

Principles of Hash-based Text Retrieval, Benno Stein

DiffusionRank: A Possible Penicillin for Web Spamming, Haixuan Yang, Irwin King, Michael R. Lyu

Context Sensitive Stemming for Web Search, Fuchun Peng, Nawaaz Ahmed, Xin Li, Yumao Lu

Combining Content and Link for Classification using Matrix Factorization, Shenghuo Zhu, Kai Yu, Yun Chi, Yihong Gong

ARSA: A Sentiment-Aware Model for Predicting Sales Performance Using Blogs, Yang Liu, Jimmy Huang, Aijun An, Xiaohui Yu

Heavy-Tailed Distributions and Multi-Keyword Queries, Arnd Konig, Surajit Chaudhuri, Liying Sui, Kenneth Church

Tuesday, December 26, 2006

Book Review: Geometry and Meaning

Title: Geometry and Meaning
Author: Dominic Widdows
URL : http://infomap.stanford.edu/book/


This is an excellent book for even high school students to learn about IR. However, if you have read a few papers in this field then reading this book is a waste of time, except for the jewels in boxes throughout the book.

Monday, December 18, 2006

New IR book

Just discovered this book by Chris Manning, Prabhakar Raghavan, and Hinrich Schütze.
This is going to be in my reading list for the IR Course, next spring.