Mining Website Logs For Hidden Robots.
We worked on a project recently, where we had found that a client’s entire site was screen scraped and put up at a different location without their permission. This of course led to a drop in search engine traffic as the new site was considered duplicate content.
The client receives a ton of traffic each day and none of the traditional visitor analytic tools (Google Analytics/Piwik) were providing any real insight into what and who was crawling their site in general.
We decided to bring out our old tool, Perl, to help data mine for the robots.
First, let me state, that by no means will this ever stop or detect every robot that is nefariously crawling your website. Most really crafty developers would mask their robot in their code, by stating that it is actually a Chrome browser running on Windows 7. It would take its time and slowly/carefully/randomly download pages to grab the content.
The purpose of our exercise below was to see what other reported robots where mining our site for links and/or possible competitive link ranking.
Apparently there was some sort of global attack against a variety of hosting providers last night (4/11/2013) targetting users of WordPress. The massive botnet attack targeted WordPress accounts and went directly after the login screen.