Wired has an article about spock.com, a people search engine that combines crawled and user added content. From the few searches I did, looks like this is good for celebrity names than a regular person with web content. For instance, searching a name like "David Smith" produces these results. Of the top 10 results, only 3 of them actually have the name "David Smith" or something closer and the first result is not one of them. Compare this with a general purpose search engine like Google. Among a dozen random NLP/ML academic names (professors) I tried, it only got Jason Eisner and Tom Mitchell correct. One reason for this poor recall is probably they don't get content from user home pages.
(Some sites where this data is derived from include MySpace, Friendster, IMDB, Wikipedia, ratemyprofessors.com, etc.)
Nevertheless, this website is a representative of interesting KDD-style problems that one could do with people names. It is also interesting as people names that we look for fall in the "long tail" without sufficient data to support calling for clever machine learning techniques.
Wednesday, August 15, 2007
People Search on the Web
-
Delip Rao
at
4:16 PM
0
comments
Principal Components: data mining, IR, KDD, NLP, search
Sunday, August 12, 2007
Digital Reasoning awarded contextual similarity patent?
I was lead to this article on Forbes via Damien's post. The article is about a company Digital Reasoning getting patent on what sounded to me as contextual similarity. Their "white paper" makes reference to a patent number 7249117 (via USPTO). Unlike research papers, reading the patent document was so difficult. Will get to it sometime later but here is an extract from their press release about what their technology can do.
* Learn the meanings of words, classes of words, and other symbols based on how they are used in context in natural language
* Create and manipulate models of this "meaning" - i.e. the mathematical patterns of usage - including the detection of groups or similar categories of words or development of hierarchies or creation of relationships between words
* Improve the models based on human feedback or using other structured information after model construction
* The representation or sharing of this model or learning in an ontology, graph structure, or programming languages
Anyone from the ACL/ML/AI community can immediately recognize this and start citing their favorite papers on these topics starting from at least a decade ago. A promotional video from the company on YouTube can be found here. Excerpt from the video: "... We treat the text representation of human language as a signal ... ".
I think everyone should stop taking patents seriously. Wishful thinking?
-
Delip Rao
at
12:30 PM
1 comments
Principal Components: "machine learning", data mining, NLP, patents
Wednesday, July 18, 2007
Reading List from KDD 2007
KDD 2007 will be on Aug 12-15 in the neighborhood at San Jose. Here is my selection:
"Extracting Semantic Relations from Query Logs", Ricardo Baeza-Yates and Alessandro Tiberi
"Efficient Incremental Clustering with Constraints", Ian Davidson, S.S. Ravi, and Martin Ester
"A Probabilistic Framework for Relational Clustering", Bo Long, Zhongfei Zhang, and Philip S. Yu
"Tracking Multiple Topics for Finding Interesting Articles", Raymond Pon, Alfonso Cardenas, David Buttler, and Terence Critchlow
"Feature Selection Methods for Text Classification", Anirban Dasgupta, Petros Drineas, Boulos Harb, Vanja Josifovski, and Michael Mahoney
"Hierarchical Mixture Models: a Probabilistic Analysis", Mark Sandler
"Information distance from a question to an answer", Xian Zhang, Yu Hao, Xiaoyan Zhu, and Ming Li
"Statistical Change Detection for Multi-Dimensional Data", Xiuyao Song, Mingxi Wu, Chris Jermaine, and Sanjay Ranka
"Constraint-Driven Clustering", Rong Ge, Martin Ester, Wen Jin, and Ian Davidson
"Enhancing Semi-Supervised Clustering: A Feature Projection Perspective", Wei Tang, Hui Xiong, Shi Zhong, and Jie Wu
-
Delip Rao
at
12:57 PM
0
comments
Principal Components: data mining, KDD, learning, reading list, research