One of my friends owns a large site very similar to YouTube. Last month and months preceding the googlebot slurped his site for 50gb+ of data. Yahoo/MSN crawlers don't even compare(less than 1gb on average). We all know google can't extrapolate a lot of data from mediums other than text.
The basic conspiracy is: YouTube analyzes his site for unique content and automatically uploads it to YouTube.
How the theory came up: He ripped a series of very old French cartoons off VHS tapes, within 24 hours they were on YouTube (subsequent uploads to other tube sites like metacafe etc.. took much much longer, and looked more 'natural').
He's trying to devise some more advanced tracking at the moment and I suggested he tried this to see if he's right,
1. Upload 10 absolutely unique videos on his site, unique description, identifiers etc..
2. Upload 10 videos ripped off YouTube using a same quality rip, descriptions etc..
3. Look at tracking to see if the 10 unique videos were scraped/uploaded, and the 10 ripped off ones not touched.
I'm posting this here to see if anyone else owns some tube sites and had similar experiences? We know he can simply block the crawler but that's besides the point, I just want to know if this technique is being used. I really can't think of any other reason why Google would crawl his site for that amount of data every month.
Maybe this is old news, but I never heard about it and tried some basic searches on google and found no mention of this technique being used by Google.
The basic conspiracy is: YouTube analyzes his site for unique content and automatically uploads it to YouTube.
How the theory came up: He ripped a series of very old French cartoons off VHS tapes, within 24 hours they were on YouTube (subsequent uploads to other tube sites like metacafe etc.. took much much longer, and looked more 'natural').
He's trying to devise some more advanced tracking at the moment and I suggested he tried this to see if he's right,
1. Upload 10 absolutely unique videos on his site, unique description, identifiers etc..
2. Upload 10 videos ripped off YouTube using a same quality rip, descriptions etc..
3. Look at tracking to see if the 10 unique videos were scraped/uploaded, and the 10 ripped off ones not touched.
I'm posting this here to see if anyone else owns some tube sites and had similar experiences? We know he can simply block the crawler but that's besides the point, I just want to know if this technique is being used. I really can't think of any other reason why Google would crawl his site for that amount of data every month.
Maybe this is old news, but I never heard about it and tried some basic searches on google and found no mention of this technique being used by Google.