Recently, I have been reading two books which may be of interest to data miners, Statistical Rules of Thumb by Gerald Van Belle (ISBN-13: 978-0471402275) and Common Errors in Statistics (and How to Avoid Them), by by Phillip I. Good and James W. Hardin (ISBN-13: 978-0471794318). Both impart practical advice based on extensive experience and statistical rigor, yet avoid becoming hung up on academic issues.
While both are written from the point of view of traditional statisticians, they do suggest the use of some less traditional techniques, such as the bootstrap and robust regression. A wide range of topics is covered, such as sample size determination, hypothesis testing and treatment of missing values. Both books also include some material written for audiences working in specific fields, such as environmental science and epidemiology. Material in these two books will vary in applicability to data mining, given the traditional statistical focus on smaller data sets and parametric modeling.
I highly recommend both of them. Tables of contents can easily be found on-line, and an entire chapter of Statistical Rules of Thumb is available at: Chapter 2: Sample Size.
Showing posts with label books. Show all posts
Showing posts with label books. Show all posts
Thursday, October 09, 2008
Wednesday, December 27, 2006
Two Book Recommendations
In my data mining courses, there are two books I always recommend to course attendees who are new to data mining. The first is Data Preparation for Data Mining by Dorian Pyle. I like this book because data preparation is usually the most time-consuming step in the data mining process, and there is only one book I know of that is written entirely for the purpose of data preparation (the second hit in the amazon list I linked is a data prep for SAS book, but that one is SAS-specific).
The second book I recommend is for the analyst who is not a statistician is Data Mining: Practical Machine Learning Tools and Techniques by Witten and Frank. They do a great job of describing algorithms and techniques in data mining in an intuitive way; there are few equations and derivations to cloud the issues for non-mathematicians. The biggest critique I have is that there is no description of neural networks, one of the key algorithms in data mining software packages. But that doesn't dampen my enthusiasm for the book. (If you would like a good, free description of neurla networks, go to the SAS Neural Network FAQ.)
The second book I recommend is for the analyst who is not a statistician is Data Mining: Practical Machine Learning Tools and Techniques by Witten and Frank. They do a great job of describing algorithms and techniques in data mining in an intuitive way; there are few equations and derivations to cloud the issues for non-mathematicians. The biggest critique I have is that there is no description of neural networks, one of the key algorithms in data mining software packages. But that doesn't dampen my enthusiasm for the book. (If you would like a good, free description of neurla networks, go to the SAS Neural Network FAQ.)
Labels:
books,
data mining
Tuesday, December 19, 2006
The Best Data Mining Book of 2005
A bit late, but better late than never! Actually, I just heard Stephen Levitt speak at SPSS Directions in November and was reminded, of course, of his book Freakonomics. In 2005, I recommended the book to my data mining course attendees as my favorite data mining book of the year, despite the term "data mining" never appearing (to the best of my knowlege) in the book at all. I think a quote in the preface summarizes why I liked it:
What interested Levitt were the stuff and riddles of everyday life... 'He (Levitt) is an intuitionist. He sifts through a pile of data to find a story no one else has found. He figures a way to measure an effect that veteran economists had declared unmeasurable.It was the idea of "sifting", a prominent term in the Gartner Group definition (and one that I like in particular) that struck me. And all the examples Levitt gives in his book are examples of uncovering patterns in data that are not the most obvious answers, but rather are ones that fit the data better (in his opinion). I like the book because he approaches data with a forensic mindset.
Labels:
books,
data mining,
economics
Subscribe to:
Posts (Atom)