我试图从allabolag.se中提取一些数据。我想要关注例如http://www.allabolag.se/5565794400/befattningar但scrapy不能正确地获取链接。它在URL中的“%2”后面缺少“52”。Scrapy - 由于编码无法关注链接
但scrapy到达下面的链接:https://www.owasp.org/index.php/Double_Encoding
我在这个网站,它得到的东西做的编码读我如何解决这个问题?
我的代码如下:
# -*- coding: utf-8 -*-
from scrapy.spiders import CrawlSpider, Rule
from scrapy.linkextractors import LinkExtractor
from scrapy.selector import HtmlXPathSelector
from scrapy.http import Request
from allabolag.items import AllabolagItem
from scrapy.loader.processors import Join
class allabolagspider(CrawlSpider):
name="allabolagspider"
# allowed_domains = ["byralistan.se"]
start_urls = [
"http://www.allabolag.se/5565794400/befattningar"
]
rules = (
Rule(LinkExtractor(allow = "http://www.allabolag.se", restrict_xpaths=('//*[@id="printContent"]//a[1]')), callback='parse_link'),
)
def parse_link(self, response):
for sel in response.xpath('//*[@id="printContent"]'):
item = AllabolagItem()
item['Byra'] = sel.xpath('/div[2]/table/tbody/tr[3]/td/h1').extract()
item['Namn'] = sel.xpath('/div[2]/table/tbody/tr[3]/td/h1').extract()
item['Gender'] = sel.xpath('/div[2]/table/tbody/tr[3]/td/h1').extract()
item['Alder'] = sel.xpath('/div[2]/table/tbody/tr[3]/td/h1').extract()
yield item
抓取时是否出现错误? – Rahul