Tuesday, 1 November 2011

ISMIR 2011: data, data and more data

This year's ISMIR conference saw some powerful critiques of current evaluation datasets for MIR tasks, but also some great new releases of data that should help us start to do a better job.

First the criticisms. Julian Urbano pointed out a widening gap in sophistication between the annual MIREX algorithm bake-off and more established equivalents such as TREC for document and image search. For anyone not completely persuaded by his arguments, Fabien Gouyon and his team demonstrated convincingly that current music autotagging algorithms fail to generalise from one dataset to another i.e. at the moment there is no hard evidence that they are really learning anything at all: the main cause is probably that the available reference datasets are simply too small.

And now the data:

Wow that's a lot of new data. Time to get down to some algorithm development!

Wednesday, 26 October 2011

ISMIR 2011: crowd sourcing

One of the themes coming out of today's sessions at ISMIR is a growing interest in crowd sourcing groundtruth and evaluation data for MIR tasks. Masataka Goto introduced his amazing Songle application, which uses an impressively-crafted Flash application to display automatically extracted chord symbols, melody line, and more, for any mp3 on the web, and also allows users to correct the automatic annotations by hand. Meanwhile Dan Stowell showed a delightfully simple web interface designed for use in the school classroom, which displays automatically extracted harmony information alongside any embedded YouTube music video: truly MIR for the masses.

Finally here are the slides from my own talk about tempo estimation:

Friday, 26 August 2011

Talks...

I'll be presenting a paper on crowd-sourcing data to improve musical tempo estimation at this year's ISMIR conference in Miami in October. This reports some research around a little interactive demo which we released a while ago on Last.fm's labs website. If you haven't tried it yet then give it a go, it's quite fun... and we'll get a bit more data.

I'm also giving another talk on Hadoop at the Stack Overflow Dev Days in London in November.  The nice Stack Overflow folks have also given me some $100 discount codes to hand out, drop me a line if you'd like  one, they are also valid for the upcoming Dev Days in other cities. Update: Dev Days has been cancelled, but it turns out that I'll be speaking on Hadoop at UCL on 16 November.

Friday, 15 April 2011

Algorithms on Hadoop

I gave a talk last night on some fun things we do with our Hadoop cluster at Last.fm including
  • topic modelling with LDA
  • graph-based recommendation with Label Propagation
  • audio analysis
Here are the slides.

Tuesday, 5 October 2010

RecSys 2010: music

How do teenagers learn about new music?  Audrey LaPlante has spent some time actually asking them, and presented some of their answers at the Workshop on Music Recommendation and Discovery.


Some of her findings were extremely interesting.  All of the teenagers she spoke to said that their tastes had changed substantially over time, and that the changes were due to changes in their social network.  Most had music geek friends whom they actively consulted for music recommendations, even though they were not influential people in other respects.  Although close contacts were most likely to be sources of new music, those chosen to play that role were "almost always those whose social network were more different from theirs, mostly those who were going to a different school".

I'll be interested to follow Audrey's research and see how we can learn to make online social networks an equally great place for young people to discover new music.

RecSys 2010: YouTube

The YouTube team had a poster at RecSys descibing their recommender in some detail.  The design is intentionally simple, and apparently entirely implemented as a series of MapReduce jobs.

They first compute fixed-length lists of related videos, based in principle on simple co-occurrence counts in a short period i.e. given a video i, the count for a candidate related video j could be the number of users who viewed both i and j within the last 24 hours.  The counts are normalised to take into account the relative popularity of different videos, and no doubt massaged in numerous other ways to remove noise and bias.  As the paper says, "this is a simplified description".

At recommendation time they build a set of seed videos representing the query user, based on the user's views, favourites, playlists, etc.  They then assemble a candidate pool containing the related videos for all of the seeds.  If the pool is too small, they expand it by adding the related videos of all the videos already in the pool, though always keeping track of the original seed video for messaging purposes.  The candidates in the pool are reranked based on a linear combination of values expressing the popularity of a given candidate video, the importance of its seed to the user, the overall popularity of the candidate and its freshness.  Finally the recommended videos are diversified, using simple constraints on the number of recommended videos that can be associated with any one seed, or have been uploaded by any one user.  Diversification is particularly important as related videos are typically very tightly associated with their seed.

Precomputed recommendations are cached and served up a few at a time to a user each time they visit the site.  Each recommendation is easily associated with an explanation based on its seed video: "recommended because you favorited abc".  While this system isn't going to win any best paper prizes it is certainly effective: 60% of all video clicks from the YouTube homepage are for recommendations.

RecSys 2010: social recommendations

Mohsen Jamali won the best paper award at RecSys for his presentation on SocialMF, a model-based recommender designed to improve ratings-based recommendations for users who have made few ratings but who have friends, or friends of friends, who have provided plenty.  Mohsen's previous solution to the same problem was TrustWalker.  TrustWalker predicts ratings from a set of item-item similarities and a social graph, where nodes are users and edges represent trust or friendship relationships.  The rating for item i for user u is predicted by taking a short random walk on the graph, stopping at some friend-of-a-friend v and returning v's rating for item j, where j is the most similar item to i which v has rated.  Closed form expressions for these predictions don't scale at all well, so to make a prediction TrustWalker actually executes the random walk a few times and returns the average rating.  On a test set of ratings by 50k users for 100k product reviews from the Epinions website, TrustWalker does indeed show a significant benefit in both coverage and prediction accuracy for cold start users over baseline methods that don't leverage the social graph.


SocialMF is a Bayesian model-based solution to the same problem: latent factors for all users and items are learned jointly from ratings and the social graph.  Ratings for cold start users can then be predicted from their learned factors.  When tested on the epinions dataset, and a new one of ratings by 1M users for 50k movies crawled from Flixster, SocialMF again does indeed improve the accuracy of predicted ratings for cold start users over latent factor models that don't take the social graph into account.

The model-based approach is elegant and perhaps even scalable: learning time is linear in the number of users, and the paper reports a runtime of 5.5 hours for 1M users.  But it lacks the powerful explanations of the simpler system: "recommended for you because your friend xyz likes it".