scrapy框架(https://github.com/scrapy/scrapy)提供了一个库,供登录需要身份验证的网站时使用,https://github.com/scrapy/loginform.
我已经浏览了这两个程序的文档,但是我似乎无法弄清楚如何让 scrapy 在运行之前调用登录表单。只需登录表单即可正常登录。
Thanks
loginform
只是一个库,与 Scrapy 完全解耦。
您必须编写代码以将其插入您想要的蜘蛛中,可能是在回调方法中。
以下是执行此操作的结构示例:
import scrapy
from loginform import fill_login_form
class MySpiderWithLogin(scrapy.Spider):
name = 'my-spider'
start_urls = [
'http://somewebsite.com/some-login-protected-page',
'http://somewebsite.com/another-protected-page',
]
login_url = 'http://somewebsite.com/login-page'
login_user = 'your-username'
login_password = 'secret-password-here'
def start_requests(self):
# let's start by sending a first request to login page
yield scrapy.Request(self.login_url, self.parse_login)
def parse_login(self, response):
# got the login page, let's fill the login form...
data, url, method = fill_login_form(response.url, response.body,
self.login_user, self.login_password)
# ... and send a request with our login data
return scrapy.FormRequest(url, formdata=dict(data),
method=method, callback=self.start_crawl)
def start_crawl(self, response):
# OK, we're in, let's start crawling the protected pages
for url in self.start_urls:
yield scrapy.Request(url)
def parse(self, response):
# do stuff with the logged in response
本文内容由网友自发贡献,版权归原作者所有,本站不承担相应法律责任。如您发现有涉嫌抄袭侵权的内容,请联系:hwhale#tublm.com(使用前将#替换为@)