我试图解析数据从网站使用beautifulsoap在python和最后我拉数据从网站所以我想保存数据在json文件但它保存数据如下根据代码我写的
json文件
[
{
"collocation": "\nabove average",
"meaning": "more than average, esp. in amount, age, height, weight etc. "
},
{
"collocation": "\nabsolutely necessary",
"meaning": "totally or completely necessary"
},
{
"collocation": "\nabuse drugs",
"meaning": "to use drugs in a way that's harmful to yourself or others"
},
{
"collocation": "\nabuse of power",
"meaning": "the harmful or unethical use of power"
},
{
"collocation": "\naccept (a) defeat",
"meaning": "to accept the fact that you didn't win a game, match, contest, election, etc."
},
我的代码:
import requests
from bs4 import BeautifulSoup
from selenium import webdriver
import pandas as pd
import json
url = "https://www.englishclub.com/ref/Collocations/"
mylist = [
"A",
"B",
"C",
"D",
"E",
"F",
"G",
"H",
"I",
"J",
"K",
"L",
"M",
"N",
"O",
"P",
"Q",
"R",
"S",
"T",
"U",
"V",
"W"
]
list = []
for i in range(23):
result = requests.get(url+mylist[i]+"/", headers=headers)
doc = BeautifulSoup(result.text, "html.parser")
collocations = doc.find_all(class_="linklisting")
for tag in collocations:
case = {
"collocation": tag.a.string,
"meaning": tag.div.string
}
list.append(case)
with open('data.json', 'w', encoding='utf-8') as f:
json.dump(list, f, ensure_ascii=False, indent=4)
但是,比如,我想为每个字母都准备一个列表,比如,一个列表代表A,另一个列表代表B,这样我就可以很容易地找到哪个字母开头的列表,然后使用它。我该怎么做呢?你可以在json文件中看到,在搭配的开头总是有\
,我该怎么删除它呢?
2条答案
按热度按时间m528fe3b1#
我已经为修改的部分写了一个注解。我们为字典中的每个字母创建了一个数组。所以在将来的使用中,你可以只使用键来获得它们,而不用担心索引
然而这是输出
cngwdvgl2#
在循环中,定义
doc
后,尝试执行以下操作:例如,对于字母B,它应该输出:
您可以将输出元素分配给CSV、 Dataframe 或其他任何类型。