如何用六步教会你使用python爬虫爬取数据

　　import os

　　import requests

　　from bs4 import BeautifulSoup

　　#爬虫头数据

　　cookies = {

　　'SINAGLOBAL': '6797875236621.702.1603159218040',

　　'SUB': '_2AkMXbqMSf8NxqwJRmfkTzmnhboh1ygvEieKhMlLJJRMxHRl-yT9jqmg8tRB6PO6N_Rc_2FhPeZF2iThYO9DfkLUGpv4V',

　　'SUBP': '0033WrSXqPxfM72-Ws9jqgMF55529P9D9Wh-nU-QNDs1Fu27p6nmwwiJ',

　　'_s_tentry': 'www.baidu.com',

　　'UOR': 'www.hfut.edu.cn,widget.weibo.com,www.baidu.com',

　　'Apache': '7782025452543.054.1635925669528',

　　'ULV': '1635925669554:15:1:1:7782025452543.054.1635925669528:1627316870256',

　　}

　　headers = {

　　'Connection': 'keep-alive',

　　'Cache-Control': 'max-age=0',

　　'Upgrade-Insecure-Requests': '1',

　　'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/84.0.4147.89 Safari/537.36 SLBrowser/7.0.0.6241 SLBChan/25',

　　'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9',

　　'Sec-Fetch-Site': 'cross-site',

　　'Sec-Fetch-Mode': 'navigate',

　　'Sec-Fetch-User': '?1',

　　'Sec-Fetch-Dest': 'document',

　　'Accept-Language': 'zh-CN,zh;q=0.9',

　　}

　　params = (

　　('cate', 'realtimehot'),

　　)

　　#数据存储

　　fo = open("http://www.jb51.net/article/微博热搜.txt",'a',encoding="utf-8")

　　#获取网页

　　response = requests.get('https://s.weibo.com/top/summary', headers=headers, params=params, cookies=cookies)

　　#解析网页

　　response.encoding='utf-8'

　　soup = BeautifulSoup(response.text, 'html.parser')

　　#爬取内容

　　content="#pl_top_realtimehot > table > tbody > tr > td.td-02 > a"

　　#清洗数据

　　a=soup.select(content)

　　for i in range(0,len(a)):

　　a[i] = a[i].text

　　fo.write(a[i]+'

　　fo.close()

您可能感兴趣的文章:

如何用六步教会你使用python爬虫爬取数据

如何用六步教会你使用python爬虫爬取数据

相关文章

大家感兴趣的内容

最近更新的内容