简短回答: 您要查找的玩家数据是NOT在该网址中。
那么你可能想问why? 我见过他们在那个页面中,为什么他们不在那里?
所以我会尝试解释一下当您浏览该网址时会发生什么使用 Chrome 等现代浏览器。
You:输入网址并按 Enter 键。
Chrome:明白了。我会尽快为您找到该页面,请稍候。 (从该网址获取内容),现在我拥有它了!但等一下让我
在我向你展示之前先阅读/解析它,(阅读里面的内容
内容),哦,糟糕,这个 javascript 告诉我要获取额外的内容
来自另一个网址的信息,好的,我会这样做;哦等等,这是另一个
告诉我在标题中加载广告,我不喜欢它,但是
我只会做我被告知的事情;等一下,这些 css 告诉我
以粗体显示玩家姓名,还不错;哦,这是另一张照片
url xxx 我需要加载,没问题...天哪,有多少东西
给我处理?我对这个网站不满意...(正在开发
一堆其他东西...)终于一切准备就绪!现在检查一下!
You:xxx玩家其实还不错,我去看看。 (点击玩家xxx)
Chrome:: ......
正如您每次浏览网页时所看到的那样,浏览器会执行许多“幕后”操作来向用户显示网页。所以基本上:输入的 url >> 获取的 url 内容 >> 解析的内容 >> 获取的其他内容 >> 呈现的所有内容 >> 显示的页面(一个或多个步骤可以同时完成)
并且使用您的代码,它只是“从 url 获取的内容”,也您想要的那些统计数据恰好是“附加内容”,必须从其他地方加载,所以这就是为什么你一无所获。
那么我如何获得这些统计数据呢?一旦您知道负责加载这些统计信息的网址,只需跟踪它们即可。我如何找到这些网址?好吧,你总是可以阅读 JavaScript...如果你足够耐心的话...
得到你想要的最简单的方法就是分析该页面加载时的流量,并找出所有幕后流量。我会推荐fiddler,但您可以使用任何您认为合适的工具。
Now let's see what happens when you load that page:
实际上,有数百个请求来完全呈现您访问的页面,您所需要做的就是找出哪一个提供“实际”或“真实”统计数据。即使其中包含“StatisticsFeed”,也有一个 url,它可能就是这个吗?让我们来看看:
{
"playerTableStats": [{
"name": "Conor Hourihane",
"firstName": "Conor",
"lastName": "Hourihane",
"playerId": 134172,
"height": 181,
"weight": 62,
"age": 25,
"isManOfTheMatch": false,
"isActive": true,
"isOpta": true,
"playedPositions": "-MC-",
"positionText": "Midfielder",
"playedPositionsShort": "M(C)",
"teamId": 142,
"teamName": "Barnsley",
"seasonId": 6365,
"seasonName": "2016/2017",
"tournamentId": 7,
"tournamentRegionId": 252,
"tournamentRegionCode": "gb-eng",
"regionCode": "ie",
"tournamentName": "Championship",
"tournamentShortName": "EC",
"rating": 7.8705882352941181,
"ranking": 1,
"apps": 17,
"subOn": 0,
"minsPlayed": 1530,
"manOfTheMatch": 4,
"yellowCard": 5.0,
"redCard": 0.0,
"goal": 3,
"assistTotal": 8,
"shotsPerGame": 2.2352941176470589,
"aerialWonPerGame": 0.6470588235294118,
"passSuccess": 81.370449678800867
},
{
"name": "Anthony Knockaert",
"firstName": "Anthony",
"lastName": "Knockaert",
"playerId": 86794,
"height": 172,
"weight": 69,
"age": 25,
"isManOfTheMatch": false,
"isActive": true,
"isOpta": true,
"playedPositions": "-AML-AMR-",
"positionText": "Midfielder",
"playedPositionsShort": "AM(LR)",
"teamId": 211,
"teamName": "Brighton",
"seasonId": 6365,
"seasonName": "2016/2017",
"tournamentId": 7,
"tournamentRegionId": 252,
"tournamentRegionCode": "gb-eng",
"regionCode": "fr",
"tournamentName": "Championship",
"tournamentShortName": "EC",
"rating": 7.6722222222222216,
"ranking": 2,
"apps": 18,
"subOn": 1,
"minsPlayed": 1471,
"manOfTheMatch": 5,
"yellowCard": 4.0,
"redCard": 0.0,
"goal": 6,
"assistTotal": 0,
"shotsPerGame": 2.3888888888888888,
"aerialWonPerGame": 0.22222222222222221,
"passSuccess": 83.420593368237348
},
{
"name": "Lewis Dunk",
"firstName": "Lewis",
"lastName": "Dunk",
"playerId": 86441,
"height": 192,
"weight": 88,
"age": 25,
"isManOfTheMatch": false,
"isActive": true,
"isOpta": true,
"playedPositions": "-DC-",
"positionText": "Defender",
"playedPositionsShort": "D(C)",
"teamId": 211,
"teamName": "Brighton",
"seasonId": 6365,
"seasonName": "2016/2017",
"tournamentId": 7,
"tournamentRegionId": 252,
"tournamentRegionCode": "gb-eng",
"regionCode": "gb-eng",
"tournamentName": "Championship",
"tournamentShortName": "EC",
"rating": 7.660000000000001,
"ranking": 3,
"apps": 18,
"subOn": 0,
"minsPlayed": 1620,
"manOfTheMatch": 3,
"yellowCard": 8.0,
"redCard": 0.0,
"goal": 1,
"assistTotal": 1,
"shotsPerGame": 0.61111111111111116,
"aerialWonPerGame": 3.5,
"passSuccess": 79.72251867662753
},
{
"name": "Tom Clarke",
"firstName": "Tom",
"lastName": "Clarke",
"playerId": 133974,
"height": 180,
"weight": 77,
"age": 28,
"isManOfTheMatch": false,
"isActive": true,
"isOpta": true,
"playedPositions": "-DC-",
"positionText": "Defender",
"playedPositionsShort": "D(C)",
"teamId": 181,
"teamName": "Preston",
"seasonId": 6365,
"seasonName": "2016/2017",
"tournamentId": 7,
"tournamentRegionId": 252,
"tournamentRegionCode": "gb-eng",
"regionCode": "gb-eng",
"tournamentName": "Championship",
"tournamentShortName": "EC",
"rating": 7.6126315789473677,
"ranking": 4,
"apps": 19,
"subOn": 0,
"minsPlayed": 1692,
"manOfTheMatch": 4,
"yellowCard": 0.0,
"redCard": 0.0,
"goal": 2,
"assistTotal": 0,
"shotsPerGame": 0.89473684210526316,
"aerialWonPerGame": 5.4736842105263159,
"passSuccess": 66.666666666666657
},
{
"name": "Pontus Jansson",
"firstName": "Pontus",
"lastName": "Jansson",
"playerId": 121123,
"height": 194,
"weight": 89,
"age": 25,
"isManOfTheMatch": false,
"isActive": true,
"isOpta": true,
"playedPositions": "-DC-",
"positionText": "Defender",
"playedPositionsShort": "D(C)",
"teamId": 19,
"teamName": "Leeds",
"seasonId": 6365,
"seasonName": "2016/2017",
"tournamentId": 7,
"tournamentRegionId": 252,
"tournamentRegionCode": "gb-eng",
"regionCode": "se",
"tournamentName": "Championship",
"tournamentShortName": "EC",
"rating": 7.5976923076923066,
"ranking": 5,
"apps": 13,
"subOn": 0,
"minsPlayed": 1126,
"manOfTheMatch": 1,
"yellowCard": 6.0,
"redCard": 0.0,
"goal": 1,
"assistTotal": 0,
"shotsPerGame": 0.53846153846153844,
"aerialWonPerGame": 3.5384615384615383,
"passSuccess": 86.336633663366342
},
{
"name": "Angus MacDonald",
"firstName": "Angus",
"lastName": "MacDonald",
"playerId": 110825,
"height": 184,
"weight": 70,
"age": 24,
"isManOfTheMatch": false,
"isActive": true,
"isOpta": true,
"playedPositions": "-DC-",
"positionText": "Defender",
"playedPositionsShort": "D(C)",
"teamId": 142,
"teamName": "Barnsley",
"seasonId": 6365,
"seasonName": "2016/2017",
"tournamentId": 7,
"tournamentRegionId": 252,
"tournamentRegionCode": "gb-eng",
"regionCode": "gb-eng",
"tournamentName": "Championship",
"tournamentShortName": "EC",
"rating": 7.5066666666666677,
"ranking": 6,
"apps": 12,
"subOn": 0,
"minsPlayed": 1080,
"manOfTheMatch": 0,
"yellowCard": 3.0,
"redCard": 0.0,
"goal": 0,
"assistTotal": 0,
"shotsPerGame": 0.33333333333333331,
"aerialWonPerGame": 4.833333333333333,
"passSuccess": 72.147651006711413
},
{
"name": "Marc Roberts",
"firstName": "Marc",
"lastName": "Roberts",
"playerId": 138949,
"height": 183,
"weight": 81,
"age": 26,
"isManOfTheMatch": false,
"isActive": true,
"isOpta": true,
"playedPositions": "-DC-",
"positionText": "Defender",
"playedPositionsShort": "D(C)",
"teamId": 142,
"teamName": "Barnsley",
"seasonId": 6365,
"seasonName": "2016/2017",
"tournamentId": 7,
"tournamentRegionId": 252,
"tournamentRegionCode": "gb-eng",
"regionCode": "gb-eng",
"tournamentName": "Championship",
"tournamentShortName": "EC",
"rating": 7.503125,
"ranking": 7,
"apps": 16,
"subOn": 0,
"minsPlayed": 1440,
"manOfTheMatch": 1,
"yellowCard": 3.0,
"redCard": 0.0,
"goal": 2,
"assistTotal": 2,
"shotsPerGame": 0.625,
"aerialWonPerGame": 7.0625,
"passSuccess": 61.595547309833023
},
{
"name": "Bradley Johnson",
"firstName": "Bradley",
"lastName": "Johnson",
"playerId": 12490,
"height": 178,
"weight": 68,
"age": 29,
"isManOfTheMatch": false,
"isActive": true,
"isOpta": true,
"playedPositions": "-MC-ML-",
"positionText": "Midfielder",
"playedPositionsShort": "M(CL)",
"teamId": 20,
"teamName": "Derby",
"seasonId": 6365,
"seasonName": "2016/2017",
"tournamentId": 7,
"tournamentRegionId": 252,
"tournamentRegionCode": "gb-eng",
"regionCode": "gb-eng",
"tournamentName": "Championship",
"tournamentShortName": "EC",
"rating": 7.4954545454545443,
"ranking": 8,
"apps": 11,
"subOn": 0,
"minsPlayed": 952,
"manOfTheMatch": 1,
"yellowCard": 4.0,
"redCard": 0.0,
"goal": 2,
"assistTotal": 1,
"shotsPerGame": 1.3636363636363635,
"aerialWonPerGame": 4.0909090909090908,
"passSuccess": 71.908127208480565
},
{
"name": "Christophe Berra",
"firstName": "Christophe",
"lastName": "Berra",
"playerId": 8287,
"height": 186,
"weight": 81,
"age": 31,
"isManOfTheMatch": false,
"isActive": true,
"isOpta": true,
"playedPositions": "-DC-",
"positionText": "Defender",
"playedPositionsShort": "D(C)",
"teamId": 165,
"teamName": "Ipswich",
"seasonId": 6365,
"seasonName": "2016/2017",
"tournamentId": 7,
"tournamentRegionId": 252,
"tournamentRegionCode": "gb-eng",
"regionCode": "gb-sct",
"tournamentName": "Championship",
"tournamentShortName": "EC",
"rating": 7.4789473684210526,
"ranking": 9,
"apps": 19,
"subOn": 0,
"minsPlayed": 1710,
"manOfTheMatch": 3,
"yellowCard": 4.0,
"redCard": 0.0,
"goal": 0,
"assistTotal": 1,
"shotsPerGame": 0.94736842105263153,
"aerialWonPerGame": 6.2105263157894735,
"passSuccess": 58.636363636363633
},
{
"name": "Adam Webster",
"firstName": "Adam",
"lastName": "Webster",
"playerId": 109922,
"height": 191,
"weight": 0,
"age": 21,
"isManOfTheMatch": false,
"isActive": true,
"isOpta": true,
"playedPositions": "-DC-",
"positionText": "Defender",
"playedPositionsShort": "D(C)",
"teamId": 165,
"teamName": "Ipswich",
"seasonId": 6365,
"seasonName": "2016/2017",
"tournamentId": 7,
"tournamentRegionId": 252,
"tournamentRegionCode": "gb-eng",
"regionCode": "gb-eng",
"tournamentName": "Championship",
"tournamentShortName": "EC",
"rating": 7.4780000000000006,
"ranking": 10,
"apps": 15,
"subOn": 1,
"minsPlayed": 1227,
"manOfTheMatch": 2,
"yellowCard": 1.0,
"redCard": 0.0,
"goal": 0,
"assistTotal": 0,
"shotsPerGame": 0.2,
"aerialWonPerGame": 5.0666666666666664,
"passSuccess": 58.256029684601117
}],
"paging": {
"currentPage": 1,
"totalPages": 34,
"resultsPerPage": 10,
"totalResults": 338,
"firstRecordIndex": 1,
"lastRecordIndex": 10
},
"statColumns": ["apps",
"subOn",
"minsPlayed",
"goal",
"assistTotal",
"yellowCard",
"redCard",
"shotsPerGame",
"passSuccess",
"aerialWonPerGame",
"manOfTheMatch"]
}
确切地!那么现在怎么办?模拟这个请求并解析内容,因为它已经是 JSON 格式,内置模块json
会很容易地完成这项工作,你甚至不必使用BeautifulSoup
你可能会问,为什么我直接浏览这个链接什么也没有呢?这是因为他们在服务器上设置了限制,以便只有具有有效标头的请求才会获得提要。那么我该如何绕过它呢?模拟“生动”使用正确的参数(主要是标题),以便他们相信你。