Facebook Scraping



It looks like the fan lists are AJAX/JSON, so wouldn't you have to have a way to sniff the packets to pull the data out? Can it do that?

Disclaimer: I don't know shit about this stuff, so I might be over complicating. Could be an easier way.
 
Here we go, found someone doing exactly that:

My strategy to set this data free was to sniff the network traffic with the Wireshark tool, then replay the HTTP calls with a ruby script. The script below will iterate over the page’s fans, save the pages as JSON in plain text files, then load the text files and convert them to CSV files in the format we used above for groups. Note that if you run this you will need to substitute the value of your cookies and the form values in the HTTP post body. This insures you are authenticated as yourself when you connect to Facebook.

Code:
#!/bin/ruby
require 'net/http'
require 'rubygems'
require 'active_support'

url = URI.parse("http://www.facebook.com/ajax/social_graph/fetch.php")

5.times do |page|
  req = Net::HTTP::Post.new(url.path, {
    'Content-Type' => 'application/x-www-form-urlencoded; charset=UTF-8',
    'User-Agent' => 'Mozilla/5.0 (Macintosh; U; Intel Mac OS X 10_6_1; en-us) AppleWebKit/531.9 (KHTML, like Gecko) Version/4.0.3 Safari/531.9',
    'Origin' => 'http://www.facebook.com',
    'X-Svn-Rev' => '189912',
    'Accept-Language' => 'en-us',    
    'Cookie' => 'lotsamagicstuff',
    })

  req.body = "page=#{page}&edge_type=fan&limit=100&class=FanManager&node_id=47337974490&post_form_id=magic&fb_dtsg=magic&post_form_id_source=AsyncRequest&__a=1"

  response = Net::HTTP.start(url.host, url.port) do |http| 
    http.request(req)
  end

  puts "Got Page #{page}"
  File.open("/tmp/fb_fanjack_page#{page}.js", 'w').write(response.body)
end

# this will be the header for the CSV file
puts 'facebook_id,Name,City-state,facebook_pic_big'

5.times do |page|
  js = File.read("/tmp/fb_fanjack_page#{page}.js")
  js = js.gsub(/for \(;;\);/, '')
  data = ActiveSupport::JSON.decode(js)
  users = data["payload"]["user_info"]
  users.each do |user|
    puts '"' + [user[0], user[1]['title'], user[1]['subtitle'], user[1]['pic']].join('","') + '"'
  end
end

Don't know shit about Ruby, but hopefully I can get this working. Solves 50% of the problem. Once I have this set, it should be fairly trivial to go out and get friend counts. Should be able to do that w/ imacros (I think).
 
Visual Basic...

create a new webBrowser control after including the proper DLL in your project
navigate it to the url you want to go to... open a file, put the entire contents of the first HTML tag in there.

WebBrowser1.Navigate2 "your URL"
open "filename" for output as #1
print #1, "<html>" & WebBrowser1.Document.All(1).innerhtml & "</html>"
close #1

Why does this work? because the current active ajax stuff that you see in the browser has become part of the DOM or Document Object Model, so everything in the DOM exists in the innerhtml property of the first element on the page which is usually the <HTML> element. IT might not be the first element, cause of document type tags and whatnot, but just look at the source, count how many tags down the HTML element is, then that's your number.
 
You need to install ruby, rubygems and rails (he requires activesupport which I believe part of Rails). There are packages that does install the entire thing in one click.

I did use RfaceBook Gem which is a ruby gem that has functions that do a lot on facebook. Here is the URL: Using Ruby on Rails with Facebook Platform - Facebook Developer Wiki
and you can look at facebooker (Booker Tutorial on Facebook). You could use one of those to make doing what you wana easier...


Best,


Here we go, found someone doing exactly that:



Code:
#!/bin/ruby
require 'net/http'
require 'rubygems'
require 'active_support'

url = URI.parse("http://www.facebook.com/ajax/social_graph/fetch.php")

5.times do |page|
  req = Net::HTTP::Post.new(url.path, {
    'Content-Type' => 'application/x-www-form-urlencoded; charset=UTF-8',
    'User-Agent' => 'Mozilla/5.0 (Macintosh; U; Intel Mac OS X 10_6_1; en-us) AppleWebKit/531.9 (KHTML, like Gecko) Version/4.0.3 Safari/531.9',
    'Origin' => 'http://www.facebook.com',
    'X-Svn-Rev' => '189912',
    'Accept-Language' => 'en-us',    
    'Cookie' => 'lotsamagicstuff',
    })

  req.body = "page=#{page}&edge_type=fan&limit=100&class=FanManager&node_id=47337974490&post_form_id=magic&fb_dtsg=magic&post_form_id_source=AsyncRequest&__a=1"

  response = Net::HTTP.start(url.host, url.port) do |http| 
    http.request(req)
  end

  puts "Got Page #{page}"
  File.open("/tmp/fb_fanjack_page#{page}.js", 'w').write(response.body)
end

# this will be the header for the CSV file
puts 'facebook_id,Name,City-state,facebook_pic_big'

5.times do |page|
  js = File.read("/tmp/fb_fanjack_page#{page}.js")
  js = js.gsub(/for \(;;\);/, '')
  data = ActiveSupport::JSON.decode(js)
  users = data["payload"]["user_info"]
  users.each do |user|
    puts '"' + [user[0], user[1]['title'], user[1]['subtitle'], user[1]['pic']].join('","') + '"'
  end
end
Don't know shit about Ruby, but hopefully I can get this working. Solves 50% of the problem. Once I have this set, it should be fairly trivial to go out and get friend counts. Should be able to do that w/ imacros (I think).
 
Visual Basic...

create a new webBrowser control after including the proper DLL in your project
navigate it to the url you want to go to... open a file, put the entire contents of the first HTML tag in there.

WebBrowser1.Navigate2 "your URL"
open "filename" for output as #1
print #1, "<html>" & WebBrowser1.Document.All(1).innerhtml & "</html>"
close #1

Ye olde skewl.