> Webbots, Spiders, and Screen Scrapers is for programmers and businesspeople
Ugh so which one is it, if it tries to satisfy both groups I reckon it will be an unsatisfying read for either one.
Hate to tell you, but there is more than enough free info out there to learn this book without buying it. All you have to do is read.
Hate to tell you, but there is more than enough free info out there to learn this book without buying it. All you have to do is read.



Hate to tell you, but there is more than enough free info out there to learn this book without buying it. All you have to do is read.
If you know what you are doing sure, but I had know idea where to begin I found the book very informative.384 pages on writing bots using PHP and cURL? I'm going to take a wild stab in the dark here, and say 364 pages of that is worthless shit.
*insert concise and fast python scraper here*
This is a great talk on the subject
PyCon 2009: Scrape the Web: Strategies for programming websites that don't expect it (Part 1 of 3) - Python Miro Community - All Python Video, All the Time
I love using Python for web scraping, what are the tools / libraries you prefer for doing it?
Would python be a better tool for scraping? I started learning it a few months ago when I was having trouble using php to parse pdf files. What a nightmare that was but I was able to do it with python.
I like it more, although it's obviously just an opinion, if you want help with python scraping, hit me up.
Nice video, thanks!
I'm mostly using my own urllib2 wrapper, https://github.com/mattseh/python-web for non-screenscraping, here's an example I just wrote today using it:
https://gist.github.com/1517279
Ditto but for Ruby. Check out Nokogiri (parser) - includes mechanize on the backend for the actual grabs.