首页
学习
活动
专区
圈层
工具
发布
社区首页 >问答首页 >如何在简单的网页抓取中停止302url重定向?

如何在简单的网页抓取中停止302url重定向?
EN

Stack Overflow用户
提问于 2016-12-28 16:56:14
回答 0查看 2.2K关注 0票数 0

我正在尝试使用Python中的Requests库抓取网站,当我尝试时:

代码语言:javascript
复制
r = requests.get('http://www.cell.com/cell-stem-cell/home', allow_redirects = False)
>>> r.status_code
302
>>> r.text
'The URL has moved <a href="https://secure.jbs.elsevierhealth.com/action/getSharedSiteSession?redirect=http%3A%2F%2Fwww.cell.com%2Fcell-stem-cell%2Fhome&rc=0&code=cell-site">here</a>\n'

当我尝试的时候:

代码语言:javascript
复制
>>> r = requests.get("https://secure.jbs.elsevierhealth.com/action/getSharedSiteSession?redirect=http%3A%2F%2Fwww.cell.com%2Fcell-stem-cell%2Fhome&rc=0&code=cell-site")
>>>
>>> r.text
'\n\n\n\n\n<style type="text/css">\n    .hidden {\n        display: none;\n        visibility: hidden;\n    }\n</style>\n\n<!-- hidden iFrame for each of the SSO URLs -->\n<div class="hidden">\n    \n        <iframe src="//acw.secure.jbs.elsevierhealth.com/SSOCore/update?utt=81c120bb854495181ef4ef3f679b12261e956c5-JKh">Your browser doesn\'t support iFrames!</iframe>\n    \n        <iframe src="//acw.sciencedirect.com/SSOCore/update?utt=81c120bb854495181ef4ef3f679b12261e956c5-JKh">Your browser doesn\'t support iFrames!</iframe>\n    \n        <iframe src="//acw.scopus.com/SSOCore/update?utt=81c120bb854495181ef4ef3f679b12261e956c5-JKh">Your browser doesn\'t support iFrames!</iframe>\n    \n        <iframe src="//acw.sciverse.com/SSOCore/update?utt=81c120bb854495181ef4ef3f679b12261e956c5-JKh">Your browser doesn\'t support iFrames!</iframe>\n    \n        <iframe src="//acw.mendeley.com/SSOCore/update?utt=81c120bb854495181ef4ef3f679b12261e956c5-JKh">Your browser doesn\'t support iFrames!</iframe>\n    \n        <iframe src="//acw.elsevier.com/SSOCore/update?utt=81c120bb854495181ef4ef3f679b12261e956c5-JKh">Your browser doesn\'t support iFrames!</iframe>\n    \n</div>\n\n\n\n<noscript>\n    <a href="CANT POST LINK BECAUSE OF LACK OF REPUTATION POINTS OF STACK OVERFLOW">Redirect</a>\n</noscript>\n\n<!-- redirect to the product page after all iFrames are rendered -->\n<script>\n    setTimeout(redirectFun,2000);\n    var iFramesList = document.getElementsByTagName("iframe");\n    var renderedIFramesCount = 0;\n    var numberOfIFrames = iFramesList.length;\n    for (var i = 0; i < iFramesList.length; i++) {\n        var iFrame = iFramesList[i];\n        bindEvent(iFrame, \'load\', function(){\n            renderedIFramesCount = renderedIFramesCount + 1;\n            if (renderedIFramesCount >= numberOfIFrames)\n            {\n                redirectFun();\n            }\n        });\n    }\n    var doRedirect = true;\n    function redirectFun() {\n        if (doRedirect)\n            window.location.href = "CANT POST THIS WEBSITE BECAUSE OF MY REPUTATION POINTS ON STACKOVERFLOW";\n        doRedirect = false;\n    }\n\n    function bindEvent(el, eventName, eventHandler) {\n        if (el.addEventListener){\n            el.addEventListener(eventName, eventHandler, false);\n        } else if (el.attachEvent){\n            el.attachEvent(eventName, eventHandler);\n        }\n    }\n</script>\n\n'

我只想得到原始网站的HTML。

EN

回答

页面原文内容由Stack Overflow提供。腾讯云小微IT领域专用引擎提供翻译支持
原文链接:

https://stackoverflow.com/questions/41358519

复制
相关文章

相似问题

领券
问题归档专栏文章快讯文章归档关键词归档开发者手册归档开发者手册 Section 归档