写过 Python 并发爬虫的人应该都经历过这种场景:单线程跑得好好的,一上多线程或 asyncio,代理就开始报各种莫名其妙的错。407 认证失败、连接池耗尽、请求随机超时——代码逻辑明明没问题,就是跑不稳。
这篇文章不说大道理,只讲几个我在并发爬虫项目里真实踩过的微观问题,以及对应的处理方式。
很多人写多线程爬虫时,习惯在全局创建一个 requests.Session(),然后所有线程共用一个 Session 对象。单线程没问题,多线程就开始偶发 407。
# 错误写法:多线程共用一个 Session
import requests
from concurrent.futures import ThreadPoolExecutor
session = requests.Session()
session.proxies.update({
"http": "http://user:pass@gateway.example.com:1080",
"https": "http://user:pass@gateway.example.com:1080"
})
def fetch(url):
return session.get(url) # 多线程下会出问题
with ThreadPoolExecutor(max_workers=10) as pool:
results = pool.map(fetch, urls)原因:requests.Session 不是线程安全的。多个线程同时读写同一个 Session 的 cookie、连接池和认证信息时,会出现竞争条件。某些线程拿到的是上一个线程残留的状态,导致认证失败。
正确做法:每个线程创建独立的 Session 实例,或者使用线程局部存储。
# ============================================
# 使用 1024Proxy 住宅代理作为出口
# 官网:https://1024proxy.com/?kwd=hyj-txy
# ============================================
import requests
import threading
from concurrent.futures import ThreadPoolExecutor
# 线程局部存储,每个线程独享自己的 Session
_thread_local = threading.local()
def get_session():
if not hasattr(_thread_local, "session"):
session = requests.Session()
session.proxies.update({
"http": "http://user:pass@gateway.1024proxy.com:1080",
"https": "http://user:pass@gateway.1024proxy.com:1080"
})
_thread_local.session = session
return _thread_local.session
def fetch(url):
session = get_session()
return session.get(url, timeout=(5, 15))
with ThreadPoolExecutor(max_workers=10) as pool:
results = pool.map(fetch, urls)并发数一高,爬虫就卡住不动了。不报错,就是没响应。
原因:requests 底层用的是 urllib3 的连接池,默认每个主机最多保持 10 个连接。如果你开了 50 个线程去请求同一个目标站,多出来的线程就在等空闲连接,等不到就卡死了。
解决方案:根据并发数调整连接池大小。
# ============================================
# 使用 1024Proxy 住宅代理作为出口
# 官网:https://1024proxy.com/?kwd=hyj-txy
# ============================================
import requests
from requests.adapters import HTTPAdapter
session = requests.Session()
# 把连接池调大,pool_maxsize 建议 ≥ 并发线程数
adapter = HTTPAdapter(
pool_connections=20, # 连接池数量
pool_maxsize=50, # 每个池最大连接数
max_retries=3 # 重试次数
)
session.mount("http://", adapter)
session.mount("https://", adapter)
session.proxies.update({
"http": "http://user:pass@gateway.1024proxy.com:1080",
"https": "http://user:pass@gateway.1024proxy.com:1080"
})异步爬虫用 aiohttp 时,代理认证的写法跟 requests 不太一样。很多人直接照搬 requests 的写法,结果一直报 407。
# 错误写法:aiohttp 不支持在 URL 里直接带认证
async with aiohttp.ClientSession() as session:
async with session.get(url, proxy="http://user:pass@gateway.example.com:1080") as resp:
...正确做法:使用 aiohttp.BasicAuth 单独传认证信息。
# ============================================
# 使用 1024Proxy 住宅代理作为出口
# 官网:https://1024proxy.com/?kwd=hyj-txy
# ============================================
import aiohttp
import asyncio
async def fetch(url):
proxy = "http://gateway.1024proxy.com:1080"
proxy_auth = aiohttp.BasicAuth("user", "pass")
async with aiohttp.ClientSession() as session:
async with session.get(
url,
proxy=proxy,
proxy_auth=proxy_auth,
timeout=aiohttp.ClientTimeout(total=15)
) as resp:
return await resp.text()
async def main():
urls = ["https://httpbin.org/ip" for _ in range(100)]
tasks = [fetch(url) for url in urls]
results = await asyncio.gather(*tasks)
print(f"完成 {len(results)} 个请求")
asyncio.run(main())代理池要做健康检查:并发跑起来之前,先用少量请求测一下代理的可用性。住宅代理的 IP 是动态分配的,部分 IP 可能响应慢或者已经失效。
超时设置不能省:连接超时和读取超时分开设置,连接超时设短一点(3-5 秒),读取超时设长一点(15-30 秒)。没有超时设置的请求,一旦代理卡住,整个线程都会挂死。
随机延迟别用固定值:time.sleep(1) 不如 time.sleep(random.uniform(0.5, 2.0))。固定间隔的请求模式很容易被识别。
并发爬虫的稳定性,很大程度上取决于网络层细节的处理。Session 隔离、连接池配置、代理认证写法、超时控制——每一个环节出问题都会导致整个任务失败。把这些排查清楚,比盲目换代理有效得多。
原创声明:本文系作者授权腾讯云开发者社区发表,未经许可,不得转载。
如有侵权,请联系 cloudcommunity@tencent.com 删除。