Wednesday, August 08, 2012

The Data is Free and Computing is Cheap, but Imagination is Dear

Recently published research, What Makes Paris Look like Paris?, attempts to classify images of street scenes according to their city of origin.  This is a fairly typical supervised machine learning project, but the source of the data is of interest.  The authors obtained a large number of Google Street View images, along with the names of the cities they came from.  Increasingly, large volumes of interesting data are being made available via the Internet, free of charge or at little cost.  Indeed, I published an article about classifying individual pixels within images as "foliage" or "not foliage", using information I obtained using on-line searches for things like "grass", "leaves", "forest" and so forth.

A bewildering array of data have been put on the Internet.  Much of this data is what you'd expect: financial quotes, government statistics, weather measurements and the like- large tables of numeric information.  However, there is a great deal of other information: 24/7 Web cam feeds which are live for years, news reports, social media spew and so on.  Additionally, much of the data for which people once charged serious bucks is now free or rather inexpensive.  Already, many firms augment the data they've paid for with free databases on the Web.  An enormous opportunity is opening up for creative data miners to consume and profit from large, often non-traditional, non-numeric data which are freely available to all, but (so far) creatively analyzed by few.


Jeremy Dalletezze said...

I am just learning the tricks of utilizing free web data, but never thought of the image route. With so much seo emphasis of images, i could definitely see some useful classifying research based on images & alt tags/titles/etc...
Thanks for the idea and cool post. Would definitely appreciate a follow-up post on that foliage project giving some tips/advice. Kind Regards,

Will Dwinnell said...

The publication to which I referred was the Jan-26-2007 posting, Pixel Classification Project, to my Data Mining in MATLAB log, at

Sadly, the figures accompanying that posting are no longer available.

Note that further details are provided in a later (Feb-02-2007) posting,

As I said, this is a very straightforward supervised learning project, and the derived features would be familiar to any image processing novice.

Regardless, this is an excellent illustration of what can be made of free data from the Internet.