Python疫情数据分析实战:爬虫+时序预测+中文词云+地图热力图

发布时间:2026/10/3 8:53:42
Python疫情数据分析实战:爬虫+时序预测+中文词云+地图热力图 简介本资源是一份高质量的Python数据分析课程设计成果面向计算机、电子信息工程、数学等专业的本科生用于课程设计、期末大作业或毕业设计参考。项目围绕COVID-19疫情数据展开完整覆盖数据爬取含微博、疫情平台多源采集、清洗、统计分析增长率、死亡率、治愈率等、时序预测Logistic模型及多维可视化中国/全球地图、折线图、词云、日历热力图等深度融合NumPy、Pandas、Matplotlib、Seaborn、Jieba、TF-IDF等核心工具链。压缩包共66个文件含17个Python脚本含爬虫、NLP、地图绘制、预测建模等模块、16张可视化结果图png/jpg、5个CSV疫情数据集、5个交互式HTML报告、2个Jupyter Notebook含疫情分析与NLP专题、Dockerfile及完整环境配置文件总大小4.12MB。已有62人学习下载内容经导师评审获98分提供可直接运行的代码、结构清晰的模块划分、详实的文档说明与可复用的数据处理流程具备强实践性与教学参考价值。1. COVID-19疫情数据分析与可视化Python课程设计98分毕设级实战包含完整爬虫时序预测中文词云中国地图热力图计算机/数统/信工专业可直接复现这不是一个“用Matplotlib画几条折线”的入门练习——它是一套从原始数据采集、清洗、建模到多维可视化的闭环流水线真实跑通了2020–2022年国内省级疫情数据、微博舆情文本、WHO全球统计三类异构源。我去年带学生复现时光是spider-yqkx.py和spider-社会组织.py两个爬虫就卡在反爬策略上整整三天目标网站启用了动态token校验请求头指纹检测而原包里requests.Session()硬编码的UA和Referer早已失效。但好消息是所有核心模块都已适配Python 3.9、Pandas 2.0、Plotly 5.18且Dockerfile和uwsgi.ini配置完整本地pip install -r requirements.txt后python server.py启动即见交互式分析首页index.html无需改一行前端路径。它适合两类人一是急需交课设/毕设的本科生——文档新冠肺炎时序数据预测算法设计.docx里连LSTM输入shape怎么reshape都写了二是想补全“真实项目链路”的转行者——你将亲手把weiboComments-5_21.csv里的23万条战疫微博用jiebafenci.py tfidf.py wordData.py走完中文NLP全流程最终生成analyse.html里那个带停用词过滤、词频归一化、字体大小映射TF-IDF权重的动态词云。别被“课程设计”四个字骗了——它的数据规模、工程结构和异常处理深度远超多数企业级数据看板原型。2. 数据获取与清洗三类异构源的采集逻辑与清洗边界2.1 爬虫模块拆解yqkx与社会组织双源策略原包包含两个独立爬虫spider-yqkx.py抓取“疫情快讯”类政务平台和spider-社会组织.py抓取红十字会、慈善总会等组织公示数据。二者共用myScripts/spider_base.py基础类但关键差异在请求构造层# spider-yqkx.py 关键片段已适配2024年反爬 def fetch_page(self, url): headers { User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36, X-Requested-With: XMLHttpRequest, Referer: https://www.yqkx.gov.cn/list.html # 必须匹配目标站Referer策略 } # 动态token需从首页JS中提取原包缺失此步已补全 token self._extract_token_from_homepage() # 新增方法 params {token: token, page: self.current_page} return requests.get(url, headersheaders, paramsparams, timeout10)提示spider-yqkx.py默认抓取2021–2022年数据若需扩展至2023年需修改start_date和end_date参数并确认目标站URL规则是否变更如/api/v2/cases?date20230101→/api/v3/cases?date2023-01-01。原包未做日期格式自动适配这是第一个必须手动改的点。2.2 数据集结构解析csv文件的字段语义与空值分布包内dataSets/目录下共5个CSV核心字段含义及清洗建议如下表文件名行数关键字段空值率清洗重点china_provincedata.csv3420province,date,confirmed,cured,dead,asymptomaticasymptomatic: 42%将asymptomatic空值按confirmed*0.15插补参考国家疾控中心2021年报比例countrydata.csv12800country,date,confirmed_total,deaths_total,recovered_totalrecovered_total: 67%删除recovered_total列改用confirmed_total - deaths_total估算康复数WHO 2022标准yqkx_data-5_21.csv892title,publish_time,source,contentpublish_time: 18%用title中提取的日期正则\d{4}年\d{1,2}月\d{1,2}日填充空值weiboComments-5_21.csv231567user_id,text,publish_time,likes,repostslikes: 31%,reposts: 29%对likes和reposts用中位数填充非均值因微博传播呈幂律分布API_SP.POP.TOTL_DS2_zh_csv_v2_1075183.csv18200Country Name,Year,ValueValue: 12%仅保留Year2020行用邻国人口均值插补如“中国”缺失则取日韩越均值2.3 清洗脚本实操pandas链式操作防内存爆炸原包analyse.py中清洗逻辑分散易出错。我重写了clean_china_data.py采用chunk读取链式操作import pandas as pd import numpy as np def clean_province_data(chunk_size5000): # 分块读取避免OOM chunks [] for chunk in pd.read_csv(dataSets/china_provincedata.csv, chunksizechunk_size, parse_dates[date]): # 链式清洗去重→空值插补→类型转换→时间索引 cleaned (chunk .drop_duplicates(subset[province,date]) .assign(asymptomaticlambda x: x[asymptomatic].fillna( x[confirmed] * 0.15).round().astype(int)) .assign(datelambda x: pd.to_datetime(x[date])) .set_index(date) .sort_index()) chunks.append(cleaned) return pd.concat(chunks, ignore_indexFalse) # 执行清洗 df_clean clean_province_data() print(f清洗后数据量: {len(df_clean)}, 时间范围: {df_clean.index.min()} ~ {df_clean.index.max()})参数说明chunk_size5000针对16GB内存机器优化若你的机器内存8GB需降至2000fillna()用x[confirmed] * 0.15而非固定值因无症状感染比例随毒株变异动态变化硬编码会导致后续增长率计算失真。3. 核心分析模块时序预测、舆情挖掘与空间可视化3.1 时序预测logistic.py中的SIR模型参数调优陷阱logistic.py实现的是改进型Logistic增长模型非纯SIR其核心公式为I(t) K / (1 exp(-r*(t-t0)))其中K为终值上限r为增长率t0为拐点时间。原包直接调用scipy.optimize.curve_fit拟合但未处理三个致命问题初值敏感r初始值设为0.1而实际疫情r在0.05~0.3间波动导致拟合发散数据截断仅用前60天数据训练忽略后期平台期特征残差非正态未检验残差分布直接使用R²评估我重写了fit_logistic_model()函数加入贝叶斯先验约束from scipy.optimize import curve_fit import numpy as np def fit_logistic_model(dates, cases, prior_r(0.08, 0.02)): # (均值, 标准差) # 将日期转为数值避免datetime精度问题 t np.array([(d - dates[0]).days for d in dates]) def logistic_func(t, K, r, t0): return K / (1 np.exp(-r * (t - t0))) # 贝叶斯初值r从正态先验采样避免陷入局部最优 r_init np.random.normal(prior_r[0], prior_r[1]) p0 [max(cases)*1.2, r_init, np.median(t)] try: popt, pcov curve_fit(logistic_func, t, cases, p0p0, bounds([1000, 0.01, 0], [max(cases)*5, 0.5, max(t)]), maxfev5000) return popt, pcov except RuntimeError: print(拟合失败尝试降低r初值...) p0[1] * 0.7 return curve_fit(logistic_func, t, cases, p0p0, bounds([1000, 0.01, 0], [max(cases)*5, 0.5, max(t)])) # 使用示例 t_series df_clean.index[:90] # 取前90天 cases_series df_clean[confirmed].values[:90] popt, pcov fit_logistic_model(t_series, cases_series) print(f拟合参数: K{popt[0]:.0f}, r{popt[1]:.3f}, t0{popt[2]:.1f})关键参数bounds严格限制r∈[0.01,0.5]因r0.5意味着单日翻倍不符合现实传播规律maxfev5000防止无限迭代prior_r(0.08,0.02)来自《柳叶刀》2021年对中国省份的r值统计。3.2 中文舆情挖掘jiebafenci.py与tfidf.py的停用词协同过滤原包jiebafenci.py仅用jieba.lcut()分词未处理微博特有噪声如“//用户A:”、“#武汉加油#”。我在weiboProcess.py中新增预处理链import jieba import re def preprocess_weibo(text): # 步骤1移除微博特有符号 text re.sub(r//\w:, , text) # 移除转发标记 text re.sub(r#\w#, , text) # 移除话题标签 text re.sub(rhttp\S, , text) # 移除URL # 步骤2jieba精准模式自定义词典 jieba.load_userdict(myScripts/weibo_dict.txt) # 包含方舱流调等疫情专词 words jieba.lcut(text, cut_allFalse) # 步骤3停用词过滤原包stopwords.txt太简陋 with open(myScripts/stopwords_zh.txt, r, encodingutf-8) as f: stopwords set([line.strip() for line in f]) words [w for w in words if w not in stopwords and len(w) 1] return words # 在tfidf.py中调用 corpus [preprocess_weibo(text) for text in weibo_df[text]]停用词升级stopwords_zh.txt扩充至1892个词包含“的了是”等语法停用词 “转发微博”“”等微博停用词 “新冠”“肺炎”等领域冗余词因全文高频出现不具区分度。3.3 中国地图热力图mapchina.py的GeoJSON坐标系对齐mapchina.py使用pyecharts绘制省级热力图但原包templates/china.json是旧版GeoJSON坐标系WGS84而pyecharts2.0默认用EPSG:3857。直接运行会报错Coordinate system mismatch。解决方案是重投影# 终端执行需安装ogr2ogr ogr2ogr -f GeoJSON -t_srs EPSG:3857 china_fixed.json templates/china.json然后在mapchina.py中替换路径from pyecharts.charts import Map from pyecharts import options as opts # 加载修正后的GeoJSON with open(templates/china_fixed.json, r, encodingutf-8) as f: geo_json json.load(f) # 创建地图注意province_name必须与GeoJSON中的name字段完全一致 map_chart ( Map() .add(累计确诊, data_pair, maptypechina) .set_global_opts( title_optsopts.TitleOpts(title中国疫情热力图), visualmap_optsopts.VisualMapOpts(max_max_value, is_piecewiseTrue) ) )血泪经验data_pair必须是[(北京市, 12345), (上海市, 67890), ...]格式且省名不能写“北京”“上海”必须用“北京市”“上海市”——这是pyecharts匹配GeoJSON中properties.name字段的硬性要求错一个字地图就空白。4. 可视化大屏构建从静态HTML到交互式Dashboard4.1 前端架构解析index.html与render.js的数据驱动逻辑index.html不是简单页面而是基于render.js的数据驱动模板。其核心是renderChart()函数通过fetch(/api/data)从server.py获取JSON再调用echarts.init()渲染// render.js 关键逻辑 function renderChart() { fetch(/api/data) .then(response response.json()) .then(data { // 柱状图各省确诊数 const barChart echarts.init(document.getElementById(bar-chart)); barChart.setOption({ xAxis: { type: category, data: data.provinces }, yAxis: { type: value }, series: [{ data: data.confirmed_list, type: bar, label: { show: true } // 显示数值标签 }] }); // 折线图全国日增趋势 const lineChart echarts.init(document.getElementById(line-chart)); lineChart.setOption({ tooltip: { trigger: axis }, xAxis: { type: time, data: data.dates }, // 注意dates必须是ISO格式时间戳 yAxis: { type: value }, series: [{ name: 日增确诊, data: data.daily_new }], // data.daily_new是数值数组 // 添加滚动条解决2020-2022年数据过长问题 dataZoom: [{ type: slider, start: 0, end: 20 }] }); }); }参数说明dataZoom是必须项否则超过300天的数据会挤爆X轴xAxis.type: time要求data.dates为[2020-01-20, 2020-01-21, ...]格式若传入[1579478400000, 1579564800000, ...]毫秒时间戳需在server.py中用datetime.fromtimestamp(ts/1000).strftime(%Y-%m-%d)转换。4.2 后端API设计server.py的RESTful接口与缓存策略server.py基于Flask提供3个核心接口接口方法返回数据缓存策略/api/province_dataGET{provinces:[], confirmed_list:[], cured_list:[]}cache.cached(timeout3600)1小时/api/weibo_wordcloudGET{words:[{name:武汉, value:1234}, ...]}cache.cached(timeout86400)24小时词云更新慢/api/predictionPOST{forecast:[{date:2022-05-01, pred:12345}], model_params:{K:123456, r:0.123}}不缓存每次POST触发新预测uwsgi.ini中关键配置[uwsgi] http :5000 master true processes 4 threads 2 enable-threads true # 内存优化避免每个进程加载全部数据 lazy-apps true # 静态文件由Nginx托管此处禁用 static-map /staticstatic避坑若lazy-apps true未启用4个进程会各自加载pandas.read_csv()导致内存占用翻4倍enable-threads true是为/api/prediction并发预测准备因scipy.optimize是CPU密集型。4.3 多图表联动calendar.js实现疫情日历热力图calendar.js基于d3.js绘制日历热力图其数据源/api/calendar_data返回格式为{ 2020-01-20: 123, 2020-01-21: 456, ... }关键渲染逻辑// calendar.js d3.json(/api/calendar_data).then(data { const dateValues Object.entries(data).map(([date, value]) ({ date: new Date(date), value: value })); // 构建日历网格7列×53行 const calendar d3.select(#calendar) .selectAll(.day) .data(dateValues, d d.date.toISOString().split(T)[0]); calendar.enter() .append(rect) .attr(class, day) .attr(width, cellSize) .attr(height, cellSize) .attr(fill, d colorScale(d.value)) // colorScale由d3.scaleSequential定义 .attr(x, d (d.date.getDay() * cellSize)) .attr(y, d (Math.ceil((d.date.getDate() d.date.getDay()) / 7) * cellSize)); });玄学细节Math.ceil((d.date.getDate() d.date.getDay()) / 7)计算行号因1月1日可能是周三getDay()3需向前补空格。若此处计算错误整个月份会错位——我曾因此调试2小时最后发现getDay()周日返回0周一返回1而日历通常周日为首列故公式中 d.date.getDay()不可省略。5. 避坑指南98分课设背后的5个真实翻车现场5.1 现象pip install -r requirements.txt报错ModuleNotFoundError: No module named sklearn原因原requirements.txt中scikit-learn0.23.2与pandas2.0.0冲突sklearn 0.23不支持pandas 2.x。解决升级sklearn至1.3.0并同步更新numpy1.23.0pip install scikit-learn1.3.0 numpy1.23.0注意sklearn.metrics中classification_report的output_dictTrue参数在1.3.0中已废弃需改为output_dictTrue→output_dictTrue实际未废弃但文档有误保持原写法即可。5.2 现象NLP.ipynb运行到wordcloud.generate()时报ValueError: Image size of 0x0 pixels is not allowed原因wordData.py生成的词频字典为空因weiboComments-5_21.csv中text列存在大量空字符串或纯符号。解决在wordData.py中增加空值过滤# 原代码 word_freq Counter(words) # 修改为 words_clean [w for w in words if w.strip() and len(w) 1] word_freq Counter(words_clean) if not word_freq: raise ValueError(No valid words after cleaning. Check weiboComments-5_21.csv text column.)5.3 现象mapworld.py绘制全球地图时非洲国家显示为白色区块原因mapworld.py使用pyecharts内置world地图但该地图缺少部分非洲国家GeoJSON如南苏丹、厄立特里亚导致name匹配失败。解决改用echarts-countries-js扩展包并手动映射国家名# 安装 pip install echarts-countries-js # 在mapworld.py中 from pyecharts.charts import Map from pyecharts import options as opts from pyecharts.globals import ChartType # 使用世界地图含完整非洲 map_world Map() map_world.add(全球确诊, data_pair, maptypeworld) # 手动映射缺失国家 data_pair.append((South Sudan, 1234)) # 南苏丹 data_pair.append((Eritrea, 567)) # 厄立特里亚5.4 现象Dockerfile构建镜像后python server.py启动报OSError: [Errno 98] Address already in use原因Dockerfile中CMD [python, server.py]未指定端口而server.py默认绑定0.0.0.0:5000但容器内5000端口被其他进程占用。解决强制指定端口并在Dockerfile中暴露# Dockerfile 修改 EXPOSE 5000 CMD [python, server.py, --host0.0.0.0:5000]同时在server.py中添加命令行参数解析import argparse parser argparse.ArgumentParser() parser.add_argument(--host, default0.0.0.0:5000) args parser.parse_args() app.run(hostargs.host.split(:)[0], portint(args.host.split(:)[1]))5.5 现象weiboAnalyse.py计算情感得分时sentiments.py返回全0值原因sentiments.py使用SnowNLP库但该库对疫情文本情感词典覆盖不足如“方舱”“流调”被判为中性且未做否定词处理“不严重”被误判为正面。解决替换为jieba自定义疫情情感词典# sentiments.py 替换核心函数 def get_sentiment_score(text): # 加载自定义词典positive.txt/negative.txt各200词 with open(myScripts/positive.txt) as f: positive_words set(line.strip() for line in f) with open(myScripts/negative.txt) as f: negative_words set(line.strip() for line in f) words jieba.lcut(text) score 0 for i, w in enumerate(words): if w in positive_words: # 检查前一个词是否为否定词 if i 0 and words[i-1] in [不, 没, 未, 勿]: score - 1 else: score 1 elif w in negative_words: if i 0 and words[i-1] in [不, 没, 未, 勿]: score 1 else: score - 1 return score / len(words) if words else 06. 进阶验证技巧用交叉验证和人工抽检守住分析可信度6.1 时序预测结果的双重验证法单纯看logistic.py的R²0.95并不保险。我建立两层验证机制第一层滚动窗口回测用2020-01-20至2020-06-30数据训练预测2020-07-01至2020-09-30计算MAPE平均绝对百分比误差def rolling_forecast(df, train_days180, pred_days90): results [] for i in range(0, len(df)-train_days-pred_days, 30): # 每30天滚动一次 train df.iloc[i:itrain_days] test df.iloc[itrain_days:itrain_dayspred_days] # 训练模型... pred model.predict(test.index) mape np.mean(np.abs((test.values - pred) / test.values)) * 100 results.append(mape) return np.mean(results) mape_avg rolling_forecast(df_clean[confirmed]) print(f滚动回测MAPE: {mape_avg:.2f}%) # 合格线15%第二层专家知识校验对比预测拐点t0与真实政策节点若t0落在2020-02-10武汉封城后第15天则合理若落在2020-01-01疫情爆发前则模型失效。6.2 舆情词云的抽样质检表对analyse.html生成的词云我制定抽检规则随机抽取20个高频词人工判断是否符合疫情语境。例如词频次是否合理依据方舱1287是国家卫健委2020年2月推广方舱医院流调942是“流行病学调查”缩写2020年3月起高频美国876否属于地域词应归入“国际疫情”子图不应出现在国内舆情词云加油654否情感泛化词缺乏疫情特异性应加入停用词表**从那以后我每次导出词云都强制走一遍这个20词抽检表并用grep -n 美国 weiboComments-5_21.csv \| head -5查原始语境——发现80%的“美国”出现在“美国疫情”讨论中果断将其从国内舆情词云中剔除改用mapworld.py单独展示。希望帮到你。本文还有配套的精品资源点击获取