InternLM3-Instruct 推理与对话实战指南:Transformers、ModelScope 与 Streamlit Web Demo

发布时间:2026/10/6 16:04:09
InternLM3-Instruct 推理与对话实战指南:Transformers、ModelScope 与 Streamlit Web Demo 大模型人工智能基础模型AI Agent【免费下载链接】InternLMOfficial release of InternLM series (InternLM, InternLM2, InternLM2.5, InternLM3).项目地址https://gitcode.com/gh_mirrors/in/InternLM点击查看免费下载本指南基于 InternLM 官方仓库的 chat/README.md系统讲解 InternLM3-8B-Instruct 的三种使用方式通过 HuggingFace Transformers 加载推理、通过 ModelScope 加载推理、以及通过 Streamlit 前端 Web Demo 进行交互式对话。读完本文你将掌握模型加载时的精度选择与显存优化技巧、基于 ChatML 模板的多轮对话组装方式、量化推理8-bit/4-bit的完整配置以及 Web Demo 中「普通回答 / 深度思考Deep Thinking」双推理模式与对比功能的源码级实现原理。为什么选择这三条推理路径InternLM3-Instruct 系列当前仓库针对internlm/internlm3-8b-instruct是上海人工智能实验室开源的可商用对话模型。仓库将其推理入口整理为三条互补的路径覆盖从「脚本内联推理」到「浏览器可视化对话」再到「服务化部署」的完整场景Transformers 路径面向 Python 开发环境模型权重、分词器均可由 HuggingFace 生态的标准 API 加载适合在自有代码中集成推理逻辑ModelScope 路径面向国内网络环境与 ModelScope 生态通过snapshot_download将模型权重先落地为本地目录再加载Web Demo 路径基于 Streamlit 的前端界面开箱即用地演示多轮对话并支持推理模式切换与结果对比。仓库根目录的 requirements.txt 给出了基础依赖transformers4.38、streamlit、sentencepiece而 Web Demo 章节要求transformers4.48实际使用时请以更高版本为准。通过 Transformers 加载 InternLM3-8B-Instructchat/README.md提供的最小可用代码如下我们先逐段拆解其关键点import torch from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer AutoTokenizer.from_pretrained(internlm/internlm3-8b-instruct, trust_remote_codeTrue) # Set torch_dtypetorch.float16 to load model in float16, otherwise it will be loaded as float32 and might cause OOM Error. model AutoModelForCausalLM.from_pretrained(internlm/internlm3-8b-instruct, trust_remote_codeTrue, torch_dtypetorch.float16) # (Optional) If on low resource devices, you can load model in 4-bit or 8-bit to further save GPU memory via bitsandbytes. # InternLM3 8B in 4bit will cost nearly 8GB GPU memory. # pip install -U bitsandbytes # 8-bit: model AutoModelForCausalLM.from_pretrained(internlm/internlm3-8b-instruct, device_mapauto, trust_remote_codeTrue, load_in_8bitTrue) # 4-bit: model AutoModelForCausalLM.from_pretrained(internlm/internlm3-8b-instruct, device_mapauto, trust_remote_codeTrue, load_in_4bitTrue) model model.eval() messages [ {role: system, content: You are an AI assistant whose name is InternLM.}, {role: user, content: Please tell me five scenic spots in Shanghai}, ] tokenized_chat tokenizer.apply_chat_template(messages, tokenizeTrue, add_generation_promptTrue, return_tensorspt) generated_ids model.generate(tokenized_chat, max_new_tokens512) generated_ids [ output_ids[len(input_ids):] for input_ids, output_ids in zip(tokenized_chat, generated_ids) ] response tokenizer.batch_decode(generated_ids)[0]参数要点与显存控制trust_remote_codeTrue是必需的InternLM3 的建模代码modeling_internlm3.py与分词器逻辑随权重仓库发布必须显式授权加载远程代码否则from_pretrained会直接报错。torch_dtypetorch.float16用于控制加载精度不指定时 Transformers 默认按 float32 加载8B 参数模型会成倍占用显存极易触发 OOM显存充足也可改用torch.bfloat16。仓库测试用例 tests/test_hf_model.py 中对 InternLM2.5 系列模型的加载方式与此完全一致torch_dtypetorch.float16, trust_remote_codeTrue可视为官方验证过的标准写法。低显存设备可叠加 bitsandbytes 量化8-bit 与 4-bit 两种量化加载方式是可选分支需要在加载前pip install -U bitsandbytes并配合device_mapauto自动切分权重。官方注明 InternLM3 8B 以 4-bit 加载时显存占用约为 8GB适合单卡消费级 GPU 运行。model.eval()切换到推理模式关闭 Dropout 等训练期行为。对话模板与生成截断tokenizer.apply_chat_template(messages, tokenizeTrue, add_generation_promptTrue, return_tensorspt)是组装对话的关键调用它将messages列表按 InternLM 的 ChatML 对话格式序列化为 token并追加生成提示符。关于该格式的详细规范见仓库的 chat/chat_format.md。generated_ids中既包含输入侧 token 也包含新生成的 token因此代码用切片output_ids[len(input_ids):]去掉输入前缀只解码新生成的部分。max_new_tokens512限制本次生成的新 token 数而非总长度。通过 ModelScope 加载 InternLM3-8B-Instruct面向国内环境或 ModelScope 生态用户chat/README.md给出了等价的加载方式。核心差异在于先通过snapshot_download将模型权重同步到本地目录model_dir后续代码与 Transformers 路径几乎一致import torch from modelscope import snapshot_download, AutoTokenizer, AutoModelForCausalLM model_dir snapshot_download(Shanghai_AI_Laboratory/internlm3-8b-instruct) tokenizer AutoTokenizer.from_pretrained(model_dir,trust_remote_codeTrue) # Set torch_dtypetorch.float16 to load model in float16, otherwise it will be loaded as float32 and might cause OOM Error. model AutoModelForCausalLM.from_pretrained(model_dir, trust_remote_codeTrue, torch_dtypetorch.float16) # (Optional) If on low resource devices, you can load model in 4-bit or 8-bit to further save GPU memory via bitsandbytes. # InternLM3 8B in 4bit will cost nearly 8GB GPU memory. # pip install -U bitsandbytes # 8-bit: model AutoModelForCausalLM.from_pretrained(model_dir, device_mapauto, trust_remote_codeTrue, load_in_8bitTrue) # 4-bit: model AutoModelForCausalLM.from_pretrained(model_dir, device_mapauto, trust_remote_codeTrue, load_in_4bitTrue) model model.eval() messages [ {role: system, content: You are an AI assistant whose name is InternLM.}, {role: user, content: Please tell me five scenic spots in Shanghai}, ] tokenized_chat tokenizer.apply_chat_template(messages, tokenizeTrue, add_generation_promptTrue, return_tensorspt) generated_ids model.generate(tokenized_chat, max_new_tokens512) generated_ids [ output_ids[len(input_ids):] for input_ids, output_ids in zip(tokenized_chat, generated_ids) ] response tokenizer.batch_decode(generated_ids)[0]注意两处差异模型 ID 使用 ModelScope 侧的命名空间Shanghai_AI_Laboratory/internlm3-8b-instructAutoTokenizer、AutoModelForCausalLM从modelscope包导入而非transformers。精度、量化、apply_chat_template等其余语义完全相同。通过 Web Demo 进行交互式对话chat/README.md的对话章节给出的启动方式极为简洁pip install streamlit pip install transformers4.48 streamlit run ./chat/web_demo.py其中./chat/web_demo.py即仓库根目录下的 chat/web_demo.py。脚本头部注释提醒务必使用streamlit run启动直接用python chat/web_demo.py可能导致未知问题如需对外访问可追加--server.address0.0.0.0 --server.port 7860。核心源码模型加载与双推理模式从源码看该 Demo 不只是简单聊天框还实现了「Normal Response」与「Deep Thinking」两种推理模式并支持对比输出模型加载load_modelchat/web_demo.py 将internlm/internlm3-8b-instruct以torch.bfloat16精度加载到 CUDA并用st.cache_resource缓存避免页面刷新时反复加载权重。生成参数prepare_generation_config侧边栏提供Max Length滑杆 8~32768默认 32768、Top P默认 0.8、Temperature默认 0.7三个滑杆以及推理模式单选。默认生成配置类GenerationConfigchat/web_demo.py设定max_length32768、top_p0.8、temperature0.8、do_sampleTrue、repetition_penalty1.005。Deep Thinking 模式开启后系统提示会追加一段面向数学竞赛的「深度思考」指令chat/web_demo.py引导模型先充分理解问题、多角度分析、系统化推理、严格论证并反复验证最终以\boxed{}给出结论。回答会以红色「Deep Thinking」前缀标识普通回答则以蓝色「Normal Response」前缀标识postprocess函数。对比功能每条机器人回复旁都有compare开关开启后会在双栏中并行渲染普通回答与深度思考回答便于直接对照两种推理模式的质量差异——这正是 README 中「supports switching between different inference modes and comparing their responses」对应的源码实现chat/web_demo.py。对话历史组装combine_history以s|im_start|system\n{meta_instruction}|im_end|\n开头逐轮按|im_start|user\n...与|im_start|assistant\n...拼接最后以当前用户问题的 assistant 起始符收尾。流式生成generate_interactive逐 token 前向计算经 logits processor 处理后按do_sample决定采样或贪心并在生成时通过additional_eos_token_id92542显式指定对话结束符即|im_end|详见下文对话格式实现边生成边渲染的效果。对话格式与特殊 TokenChatML 扩展Web Demo 中的|im_start|/|im_end|结构来自 InternLM 系列沿用的 ChatML 风格对话格式仓库文档 chat/chat_format.md 对其有完整规范常规对话包含system、user、assistant三个角色为支持智能体应用还引入了environment角色与工具调用相关 token。该格式的核心要点每轮对话以|im_start|role开头、以|im_end|结尾role可取system、user、assistant、environment。词表中维护了 6 个特殊 token 映射|im_start|对话开始符→ token ID92543|im_end|对话结束符→ token ID92542|action_start|调用外部工具开始符→ token ID92541|action_end|调用外部工具结束符→ token ID92540|interpreter|代码解释器→ token ID92539|plugin|外部插件/常规工具→ token ID92538工具调用场景下模型会流式输出「思考文字 |action_start||plugin| JSON 格式调用参数 |action_end|」环境返回结果则以|im_start|environment name|plugin|开头。代码解释器场景使用|interpreter|与 markdown 代码块支持数据分析、复杂计算、文件操作等完整示例见 chat/chat_format.md。理解这套格式有助于排查 Web Demo 中生成截断、结束符缺失等问题——additional_eos_token_id92542正是因为普通 EOS token 无法可靠标识多轮对话的轮次边界。进阶以 LMDeploy 做服务化推理若需更高吞吐的批量推理或 OpenAI 兼容的 HTTP 服务仓库 chat/lmdeploy.md 提供了基于 LMDeploy 的补充路径面向 InternLM2.5 系列Python 3.8pip install lmdeploy0.2.1离线批处理只需四行代码from lmdeploy import pipeline pipe pipeline(internlm/internlm2_5-7b-chat) response pipe([Hi, pls intro yourself, Shanghai is]) print(response)长文本场景可借助 dynamic ntk 将上下文外推到 200Kfrom lmdeploy import pipeline, TurbomindEngineConfig engine_config TurbomindEngineConfig(session_len200000, rope_scaling_factor2.0) pipe pipeline(internlm/internlm2_5-7b-chat, backend_engineengine_config) gen_config GenerationConfig(top_p0.8, top_k40, temperature0.8, max_new_tokens1024) response pipe(prompt, gen_configgen_config) print(response)一键部署服务并测试lmdeploy serve api_server internlm/internlm2_5-7b-chat lmdeploy serve api_client http://0.0.0.0:23333api_server默认监听 23333 端口提供的 RESTful API 兼容 OpenAI 接口也可通过 Swagger UIhttp://0.0.0.0:23333在线试用。仓库测试 tests/test_hf_model.py 还演示了以TurbomindEngineConfig(model_formatawq)加载 AWQ 量化版模型internlm/internlm2-chat-20b-4bits做批量推理的写法可作为低显存服务化的参考。快速验证清单加载时务必带trust_remote_codeTrue并显式指定torch_dtype否则可能 OOM低显存环境按需选择 8-bit/4-bit 量化4-bit 下 InternLM3 8B 显存占用约 8GB对话组装使用tokenizer.apply_chat_template(...)生成后需切片去除输入前缀再解码Web Demo 一律用streamlit run ./chat/web_demo.py启动transformers4.48、streamlit缺一不可多轮对话边界由|im_end|token 92542标识Web Demo 中已将其作为附加 EOS token 传入生成函数生产级推理可迁移至 chat/lmdeploy.md 描述的 LMDeploy pipeline / api_server 方案。以上推理流程与参数均以本仓库 chat/README.md 及其对应源码chat/web_demo.py、chat/chat_format.md、tests/test_hf_model.py为准可直接复制运行并在此基础上二次开发。赞分享大模型人工智能基础模型AI Agent【免费下载链接】InternLMOfficial release of InternLM series (InternLM, InternLM2, InternLM2.5, InternLM3).项目地址https://gitcode.com/gh_mirrors/in/InternLM点击查看免费下载相关推荐InternLM3-8B-Instruct 对话推理实战Transformers 与 ModelScope 加载、Streamlit 网页对话与推理模式切换InternLM3 8B Instruct 对话推理实战Transformers 与 ModelScope 加载、Streamlit 网页对话与推理模式切换大模型人工智能基础模型AI AgentInternLM-Chat-7B 对话 Web Demo 部署指南基于 Streamlit 的本地 Web 对话全流程实战InternLM Chat 7B 对话 Web Demo 部署指南基于 Streamlit 的本地 Web 对话全流程实战 本篇指南以 Datawhale「开大模型人工智能教程本地部署微调Datawhale Self-LLM 实战基于 transformers 与 peft 的 InternLM3-8B-Instruct LoRA 微调教程Datawhale Self LLM 实战基于 transformers 与 peft 的 InternLM3 8B Instruct LoRA 微调教程 本大模型人工智能教程本地部署微调上一篇PowerSploit 域渗透侦察使用 Get-DomainGPOUserLocalGroupMapping 通过 GPO 关联枚举用户本地组成员关系下一篇IQInvision 摄像头 Telnet 默认凭据检测routersploit 字典攻击模块实战指南创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考