I'll be ordering this for sure in 2012

CaviB

Google Hammered 24/7
Jul 11, 2009
520
3
0
Canada
[ame=http://www.amazon.com/Webbots-Spiders-Screen-Scrapers-Automate/dp/1593273975/ref=sr_1_8?s=books&ie=UTF8&qid=1324723305&sr=1-8]Amazon.com: Webbots, Spiders, and Screen Scrapers: Scrape, Parse, Automate, and Own the Internet (9781593273972): Michael Schrenk: Books[/ame]
51rMMDI5YML._SL500_AA300_.jpg
 
  • Like
Reactions: guerilla


Hate to tell you, but there is more than enough free info out there to learn this book without buying it. All you have to do is read.
 
> Webbots, Spiders, and Screen Scrapers is for programmers and businesspeople

Ugh so which one is it, if it tries to satisfy both groups I reckon it will be an unsatisfying read for either one.
 
> Webbots, Spiders, and Screen Scrapers is for programmers and businesspeople

Ugh so which one is it, if it tries to satisfy both groups I reckon it will be an unsatisfying read for either one.

Not really, douche bags call webbots or spiders, screen scrapers, even though they are the same fucking thing generally. A true screen scraper to me would be utilizing a browser to automate things (like Chrome). However that's not the term most people use.
 
Hate to tell you, but there is more than enough free info out there to learn this book without buying it. All you have to do is read.

Maybe your right I do have the first one just finished it recently, for some reason I'm always on the look out for more and more information on the same subject in books I rarely get any ideas from scripts posted online.
 
Hate to tell you, but there is more than enough free info out there to learn this book without buying it. All you have to do is read.

Having actual paper book gives you +20% skill, maybe even +30% if you read it in public.
 
Actually I am catching myself Screwing up again!

All this time I have read these books gained the knowledge and have failed to put this into action. I need to start using it before it rots away. I mean you should never stop learning but it's time to put this shit to action as you can see I've read a few books on these subjects.
books1sm.jpg

books2sm.jpg

books3sm.jpg
 
384 pages on writing bots using PHP and cURL? I'm going to take a wild stab in the dark here, and say 364 pages of that is worthless shit.

It's not difficult to make a bot with PHP/cURL, and it's not like there's a bunch of different methodologies you can choose from. HTTP protocol is HTTP protocol.
 
Hate to tell you, but there is more than enough free info out there to learn this book without buying it. All you have to do is read.

Don't tell Rage9 he's right, because it'll go to his head, but this is a very true fact. Hell, my blog post usually comes up when you search for advanced data scraping and Smaxor's comes up when you search for PHP web scraper or php screen scraper (which is what I learned concepts on years ago).

Books are nice, but PHP topics have been beaten to death with online tutorials. You can find anything if you know how to do a proper Google search. Oh, and let's not forget the PHP Warchest here on WF
 
384 pages on writing bots using PHP and cURL? I'm going to take a wild stab in the dark here, and say 364 pages of that is worthless shit.
If you know what you are doing sure, but I had know idea where to begin I found the book very informative.
 
right now I'm working on a project for a webbot that can go after comments, download them and categorize them, finally putting them into a database to be used on a site I am making.
 
Would python be a better tool for scraping? I started learning it a few months ago when I was having trouble using php to parse pdf files. What a nightmare that was but I was able to do it with python.
 
This is a great talk on the subject

PyCon 2009: Scrape the Web: Strategies for programming websites that don't expect it (Part 1 of 3) - Python Miro Community - All Python Video, All the Time

I love using Python for web scraping, what are the tools / libraries you prefer for doing it?

Nice video, thanks!

I'm mostly using my own urllib2 wrapper, https://github.com/mattseh/python-web for non-screenscraping, here's an example I just wrote today using it:

https://gist.github.com/1517279
 
Would python be a better tool for scraping? I started learning it a few months ago when I was having trouble using php to parse pdf files. What a nightmare that was but I was able to do it with python.

I like it more, although it's obviously just an opinion, if you want help with python scraping, hit me up.
 
Ditto but for Ruby. Check out Nokogiri (parser) - includes mechanize on the backend for the actual grabs.

I love mechanize, using spoofed user-agents, proxies, eventlet (green threads) with cookie management and easy form filling without any need for UI is super nice.