As a grad student working on NLP how do you explain what you are working on, to friends and family? I inevitably end up referring to the Google search engine even though what I do is quite far from IR. Actually, thats not true. These days IR seems to consume everything but thats another story.
This reminds me of a funny conversation at CLSP recently:
Sanjeev is telling us about an incident where a concerned parent of a young child with a speaking disability is asking him for his opinion. Apparently, she is confused about "Language and Speech Processing" in CLSP.
Keith butts in: "Run a few more iterations of EM and he'll be fine."
Sunday, February 24, 2008
What do you do?
-
Delip Rao
at
5:50 PM
1 comments
Principal Components: Geek Humor, NLP
Thursday, February 14, 2008
A song on parsing
We all know Jason's love for parsing from his work but it takes a different level of dedication to write a Valentine's Day song about parsing.
As Jason says, "Parsers just want to be appreciated, like everyone else."
-
Delip Rao
at
1:23 AM
0
comments
Principal Components: Geek Humor, NLP, Parsing
Wednesday, October 17, 2007
Funny bone
The frequentist exclaimed, "All your Bayes are belong to us!" to which the Bayesian responded, "Well, it depends."
-
Delip Rao
at
5:07 PM
0
comments
Principal Components: Geek Humor, Humor, math, statistics
Thursday, September 20, 2007
NIPS papers are out
For a full list see here. Some papers I want to read based on my current interests:
Random Projections for Manifold Learning
Chinmay Hegde, Michael Wakin, Richard Baraniuk
The Distribution Family of Similarity Distances
Gertjan Burghouts, Arnold Smeulders, Jan-Mark Geusebroek
Manifold Sculpting
Michael Gashler, Dan Ventura, Tony Martinez
A learning framework for nearest neighbor search
Lawrence Cayton, Sanjoy Dasgupta
Learning Bounds for Domain Adaptation
John Blitzer, Koby Crammer, Alex Kulesza, Fernando Pereira, Jennifer Wortman
Convex Relaxations of EM
Yuhong Guo, Dale Schuurmans
A Randomized Algorithm for Large Scale Support Vector Learning
Krishnan Kumar, Chiru Bhattacharya, Ramesh Hariharan
Bundle Methods for Machine Learning
Alex Smola, S V N Vishwanathan, Quoc Le
Regularized Boost for Semi-Supervised Learning
Ke Chen, Shihai Wang
Learning the structure of manifolds using random projections
Yoav Freund, Sanjoy Dasgupta, Mayank Kabra, Nakul Verma
A complexity measure for intuitive theories
Charles Kemp, Noah Goodman, Joshua Tenenbaum
-
Delip Rao
at
10:58 PM
0
comments
Principal Components: learning, machine learning, ML, NIPS
Saturday, August 18, 2007
NLP and Global Warming
Those of us who were at EMNLP-CONLL 2007 remember the "NLP and Global Warming" exchange between James Clarke, Jason Eisner, and Dan Bikel at the Q/A session of the Clarke and Lapata paper. The transcript of this funny conversation is now online, thanks to Jason.
I really liked Hal's ending remark.
-
Delip Rao
at
12:40 AM
0
comments
Wednesday, August 15, 2007
People Search on the Web
Wired has an article about spock.com, a people search engine that combines crawled and user added content. From the few searches I did, looks like this is good for celebrity names than a regular person with web content. For instance, searching a name like "David Smith" produces these results. Of the top 10 results, only 3 of them actually have the name "David Smith" or something closer and the first result is not one of them. Compare this with a general purpose search engine like Google. Among a dozen random NLP/ML academic names (professors) I tried, it only got Jason Eisner and Tom Mitchell correct. One reason for this poor recall is probably they don't get content from user home pages.
(Some sites where this data is derived from include MySpace, Friendster, IMDB, Wikipedia, ratemyprofessors.com, etc.)
Nevertheless, this website is a representative of interesting KDD-style problems that one could do with people names. It is also interesting as people names that we look for fall in the "long tail" without sufficient data to support calling for clever machine learning techniques.
-
Delip Rao
at
4:16 PM
0
comments
Principal Components: data mining, IR, KDD, NLP, search
Sunday, August 12, 2007
Digital Reasoning awarded contextual similarity patent?
I was lead to this article on Forbes via Damien's post. The article is about a company Digital Reasoning getting patent on what sounded to me as contextual similarity. Their "white paper" makes reference to a patent number 7249117 (via USPTO). Unlike research papers, reading the patent document was so difficult. Will get to it sometime later but here is an extract from their press release about what their technology can do.
* Learn the meanings of words, classes of words, and other symbols based on how they are used in context in natural language
* Create and manipulate models of this "meaning" - i.e. the mathematical patterns of usage - including the detection of groups or similar categories of words or development of hierarchies or creation of relationships between words
* Improve the models based on human feedback or using other structured information after model construction
* The representation or sharing of this model or learning in an ontology, graph structure, or programming languages
Anyone from the ACL/ML/AI community can immediately recognize this and start citing their favorite papers on these topics starting from at least a decade ago. A promotional video from the company on YouTube can be found here. Excerpt from the video: "... We treat the text representation of human language as a signal ... ".
I think everyone should stop taking patents seriously. Wishful thinking?
-
Delip Rao
at
12:30 PM
1 comments
Principal Components: "machine learning", data mining, NLP, patents