RKLLM-Server 实战指南:用 Flask/Gradio 在 RK 板卡上搭建 OpenAI 兼容的大模型推理服务

发布时间:2026/10/5 15:09:24
RKLLM-Server 实战指南:用 Flask/Gradio 在 RK 板卡上搭建 OpenAI 兼容的大模型推理服务 大模型本地部署推理引擎模型量化嵌入式【免费下载链接】rknn-llm项目地址https://gitcode.com/gh_mirrors/rk/rknn-llm点击查看免费下载导读本文聚焦 rknn-llm 仓库中的examples/rkllm_server_demo示例讲解如何将转换好的 RKLLM 模型部署为可对外服务的网络服务一套是基于 Flask 的OpenAI 兼容 HTTP API支持流式输出与 Function Calling 工具调用另一套是基于 Gradio 的浏览器聊天界面。读完本文你将掌握完整的部署命令、参数含义、HTTP 调用方式与函数调用实现并理解从 HTTP 请求到 NPU 推理的底层调用链。一、运行前的准备在启动任何服务之前需要完成两件事参见 examples/rkllm_server_demo/README.md准备好已转换的 RKLLM 模型文件即通过 rkllm-toolkit 转换得到的.rkllm模型需要拷贝到 Linux 板卡上并记录其板卡上的绝对路径。确认板卡 IP在板卡上执行ifconfig命令查看 IP 地址后续客户端与服务端通信、浏览器访问都要用到它。此外仓库为 RK3588、RK3576、RK3562、RV1126B 提供了频率锁定脚本scripts 下的fix_freq_rk*.sh部署脚本会自动将其拷贝到板卡并执行。二、一键构建与部署示例目录下提供了两个自动化部署脚本均通过ADB完成板卡依赖安装、文件推送与服务启动脚本服务类型说明build_rkllm_server_flask.shFlask 服务提供 OpenAI 兼容 API面向程序化调用build_rkllm_server_gradio.shGradio 服务提供网页聊天 UI面向人机交互2.1 Flask ServerOpenAI 兼容 API# Usage: ./build_rkllm_server_flask.sh --workshop [Working Path] --model_path [Absolute Path of RKLLM Model on Board] --platform [Target Platform: rk3588/rk3576] [--lora_model_path [Lora Model Path]] [--prompt_cache_path [Prompt Cache File Path]] [--adb_device [ADB Device Serial]] ./build_rkllm_server_flask.sh --workshop /userdata --model_path /userdata/model.rkllm --platform rk3588 --adb_device 1234567890abcdef启动成功后Flask 服务默认监听0.0.0.0:8080见 flask_server.py 中的app.run(host0.0.0.0, port8080, threadedFalse, debugFalse)对外提供两个 OpenAI 兼容端点EndpointMethodDescription/v1/modelsGET列出可用模型/v1/chat/completionsPOST聊天补全支持流式与非流式2.2 Gradio ServerWeb UI# Usage: ./build_rkllm_server_gradio.sh --workshop [Working Path] --model_path [Absolute Path of RKLLM Model on Board] --platform [Target Platform: rk3588/rk3576] [--lora_model_path [Lora Model Path]] [--prompt_cache_path [Prompt Cache File Path]] [--adb_device [ADB Device Serial]] ./build_rkllm_server_gradio.sh --workshop /userdata --model_path /userdata/model.rkllm --platform rk3588 --adb_device 1234567890abcdef部署完成后在浏览器中打开http://board-ip:8080即可进入 Gradio 聊天界面。2.3 命令行参数详解两个脚本的参数完全一致含义如下参数必选含义--workshop是板卡上的工作目录服务文件会被推送到该目录下--model_path是RKLLM 模型在板卡上的绝对路径--platform是目标平台rk3588/rk3576脚本同时兼容 rv1126b、rk3562--lora_model_path否LoRA 模型在板卡上的绝对路径--prompt_cache_path否Prompt Cache 文件的绝对路径--adb_device否ADB 设备序列号多设备连接时用于指定目标设备--help否打印帮助信息2.4 部署脚本内部做了什么以 build_rkllm_server_flask.sh 为例脚本分为三个阶段依赖检查与安装通过 ADB 进入板卡 shell先pkill掉旧的flask_server.py/gradio_server.py进程随后检查pip3是否存在不存在则apt install python3-pip再检查 Flask 是否可用否则通过清华源安装flask2.2.2与Werkzeug2.2.2Gradio 脚本则安装gradio4.24.0并带--break-system-packages选项。文件推送工作目录不存在时自动mkdir -p随后将 rkllm-runtime 中的librkllmrt.so来自Linux/librkllm_api/aarch64/拷贝进./rkllm_server/lib/把 scripts 下的fix_freq_rk3576.sh、fix_freq_rk3588.sh、fix_freq_rv1126b.sh、fix_freq_rk3562.sh一并拷入服务目录最后整体adb push到板卡。启动服务拼装启动命令python3 flask_server.py --rkllm_model_path MODEL_PATH --target_platform PLATFORM若传入 LoRA 或 Prompt Cache 路径则追加对应参数再通过 ADB shell 在板卡上后台执行。脚本中还注释保留了export RKLLM_LOG_LEVEL2可用于调整运行时日志级别。2.5 服务启动时的底层流程在板卡上直接执行flask_server.py时flask_server.py启动序列依次为参数校验检查--rkllm_model_path、--target_platform合法值rk3588/rk3576/rv1126b/rk3562以及可选的 LoRA / Prompt Cache 路径是否真实存在锁频执行sudo bash fix_freq_platform.sh将 CPU/NPU 频率固定在高频保证推理性能稳定资源限制resource.setrlimit(resource.RLIMIT_NOFILE, (102400, 102400))提高进程可打开文件数上限模型初始化通过ctypes.CDLL(lib/librkllmrt.so)加载运行时动态库并调用rkllm_init完成模型加载模型名取自模型文件主文件名os.path.basename(model_path).rsplit(., 1)[0]随后展示在/v1/models返回中信号处理注册SIGINT/SIGTERM处理器收到中断信号时调用rkllm_model.release()底层即rkllm_destroy释放模型资源后优雅退出。三、Flask API 使用指南Flask 服务暴露的是 OpenAI 兼容的 HTTP API同时支持常规对话与 Function Calling工具调用。3.1 常规对话常规对话即标准的问答交互客户端发送消息服务端返回模型回复。快速开始推荐使用RKLLMClient封装类最小代码即可完成调用from chat_api_flask import RKLLMClient client RKLLMClient(base_urlhttp://x.x.x.x:8080) # 非流式 reply client.chat_simple(Hello, introduce yourself please.) print(reply) # 流式 for chunk in client.chat(messages[{role: user, content: Hello}], streamTrue): print(chunk[content], end, flushTrue)也可以直接运行内置的交互式演示chat_api_flask.py# 交互式对话Demo 1 python chat_api_flask.py --server http://board-ip:8080 # 关闭流式 python chat_api_flask.py --server http://board-ip:8080 --no-stream注意将board-ip替换为板卡实际 IP用ifconfig查看。交互式 demo 中支持输入exit退出、clear清空历史记录并自动维护完整的多轮消息历史。手动 HTTP 调用如果需要自定义实现按以下步骤1) 设置服务器地址server_url http://x.x.x.x:8080/v1/chat/completions2) 创建会话import requests session requests.Session() session.keep_alive False adapter requests.adapters.HTTPAdapter(max_retries3) session.mount(https://, adapter) session.mount(http://, adapter)3) 构造请求体data { model: rkllm, # 模型标识可自定义 messages: [ {role: user, content: Hello} ], stream: False, # True 流式False 非流式 temperature: 0.8, # 采样温度默认0.8 top_p: 0.9, # 核采样阈值默认0.9 top_k: 1, # Top-k 采样默认1 max_tokens: 4096, # 最大生成 token 数默认4096 repeat_penalty: 1.1, # 重复惩罚默认1.1 frequency_penalty: 0.0, # 频率惩罚默认0.0 presence_penalty: 0.0, # 存在惩罚默认0.0 enable_thinking: False, # 是否启用深度思考模式 }4) 发送请求headers {Content-Type: application/json, Authorization: not_required} resp session.post(server_url, jsondata, headersheaders, streamdata[stream])5) 解析响应import json # 非流式 if resp.status_code 200: result resp.json() print(A:, result[choices][0][message][content]) else: print(Error:, resp.text) # 流式SSE if resp.status_code 200: for line in resp.iter_lines(decode_unicodeTrue): if line.startswith(data: ) and line[6:].strip() ! [DONE]: chunk json.loads(line[6:]) delta chunk[choices][0].get(delta, {}) if delta.get(content): print(delta[content], end, flushTrue)请求参数与采样默认值对照上述请求体中的采样参数并非写死服务端会在 flask_server.py 中逐项读取请求 JSON 并按默认值兜底然后构造成RKLLMSamplingParam与 rkllm.h 中定义的 C 结构体一一对应再在单次推理时注入。各参数对推理行为的影响temperature控制采样的随机性值越大输出越发散top_p核采样阈值只从累计概率达到该值的 token 集合中采样top_k只从概率最高的 k 个 token 中采样max_tokens单次生成的最大新 token 数repeat_penalty/frequency_penalty/presence_penalty三种不同的重复/频率抑制策略enable_thinking对应底层RKLLMInput.enable_thinking字段为 True 时模型输出思考过程。3.2 Function Calling工具调用注意本节以Qwen3系列模型为例。其他模型如 DeepSeek、LLaMA输出的工具调用格式可能不同——请按需调整parse_tool_calls()。Function Calling 允许模型根据用户请求自动调用预定义函数从而获取实时或精确的信息如天气、计算。场景说明以天气查询为例用户问旧金山今天和明天的气温是多少模型本身无法回答实时问题但可以通过调用get_current_temperature和get_temperature_date两个工具获取准确数据。快速开始推荐from chat_api_flask import RKLLMClient, TOOLS, parse_tool_calls, execute_tool_calls client RKLLMClient(base_urlhttp://x.x.x.x:8080) messages [ {role: system, content: You are Qwen, created by Alibaba Cloud. You are a helpful assistant.\nCurrent Date: 2024-09-30}, {role: user, content: Whats the temperature in San Francisco now? How about tomorrow?}, ] # 步骤 1模型返回要调用哪些工具及参数 resp client.chat(messagesmessages, toolsTOOLS, streamFalse) tool_calls parse_tool_calls(resp[choices][0][message][content]) # 步骤 2执行工具调用并把结果追加到消息列表 assistant_msg, tool_msgs execute_tool_calls(tool_calls) messages.append(assistant_msg) messages.extend(tool_msgs) # 步骤 3模型根据工具结果综合出最终回答 resp client.chat(messagesmessages, toolsNone, streamFalse) print(A:, resp[choices][0][message][content])运行内置函数调用 demopython chat_api_flask.py --server http://board-ip:8080 --demo 2手动 HTTP 调用1) 定义工具函数def get_current_temperature(location: str, unit: str celsius): return {temperature: 26.1, location: location, unit: unit} def get_temperature_date(location: str, date: str, unit: str celsius): return {temperature: 25.9, location: location, date: date, unit: unit} FUNCTION_MAP { get_current_temperature: get_current_temperature, get_temperature_date: get_temperature_date, }2) 定义工具描述OpenAI Function Calling 格式TOOLS [ { type: function, function: { name: get_current_temperature, description: Get current temperature at a location., parameters: { type: object, properties: { location: {type: string, description: City, State, Country}, unit: {type: string, enum: [celsius, fahrenheit]}, }, required: [location], }, }, }, { type: function, function: { name: get_temperature_date, description: Get temperature at a location and date., parameters: { type: object, properties: { location: {type: string, description: City, State, Country}, date: {type: string, description: YYYY-MM-DD}, unit: {type: string, enum: [celsius, fahrenheit]}, }, required: [location, date], }, }, }, ]3) 第一次调用模型返回工具调用指令messages [ {role: system, content: You are Qwen, created by Alibaba Cloud. You are a helpful assistant.\nCurrent Date: 2024-09-30}, {role: user, content: Whats the temperature in San Francisco now? How about tomorrow?}, ] data { model: rkllm, messages: messages, stream: False, tools: TOOLS, } resp session.post(server_url, jsondata, headersheaders) server_answer resp.json()[choices][0][message][content] # 模型输出示例 # tool_call # {name: get_current_temperature, arguments: {location: San Francisco}} # /tool_call # tool_call # {name: get_temperature_date, arguments: {location: San Francisco, date: 2024-10-01}} # /tool_call4) 解析工具调用并执行import re, json # 同时支持 JSON 与 XML-like 的 tool_call 格式 matches re.findall(rtool_call\s*(\{.*?\})\s*/tool_call, server_answer, re.DOTALL) tool_calls [json.loads(m) for m in matches] # 执行函数并构造消息 for tc in tool_calls: name, args tc[name], tc[arguments] result FUNCTION_MAPname messages.append({role: tool, name: name, content: json.dumps(result)})5) 第二次调用获取最终回答data[messages] messages resp session.post(server_url, jsondata, headersheaders) print(A:, resp.json()[choices][0][message][content]) # 输出示例 # A: The current temperature in San Francisco is 26.1°C. # Tomorrow, the temperature is expected to be 25.9°C.重要注意事项项目说明模型兼容性该 demo 基于 Qwen3 系列。其他模型需调整parse_tool_calls()的正则以及tool_response_str参数工具调用格式Qwen3 输出tool_call.../tool_call包裹的 JSONXML-like 格式functionxxxparameterxxx.../parameter/function同样支持多工具结果合并多个工具结果会在发给模型前合并为 JSON 数组[{...}, {...}]模型能力能力较弱的模型如 Qwen3-0.6B可能无法返回准确的工具调用参数系统提示词系统提示词模板是 Qwen 专属的使用其他模型时需替换为对应模板四、客户端脚本脚本说明chat_api_flask.py面向 Flask 服务的 OpenAI 兼容客户端支持对话 函数调用chat_api_gradio.py面向 Gradio 服务的客户端通过 Gradio API 聊天chat_api_flask.py中的RKLLMClient封装了GET /v1/models与POST /v1/chat/completions两个调用内部使用带 3 次重试的requests.Sessionchat()方法在非流式时返回 OpenAI 风格的字典流式时返回按data:前缀解析出的增量内容生成器parse_tool_calls()依次尝试 JSON 与 XML-like 两种工具调用格式execute_tool_calls()同时兼容{name:..., arguments:...}与 OpenAI 风格{function: {...}}两种输入。五、底层原理从 HTTP 请求到 NPU 推理的调用链5.1 通过 ctypes 直接绑定 C 运行时两个服务端都没有依赖独立的 Python 绑定库而是用 Python 标准库ctypes直接加载lib/librkllmrt.soflask_server.py并在 Python 侧完整复刻了 rkllm.h 中定义的枚举与结构体包括输入类型枚举RKLLMInputTypeRKLLM_INPUT_PROMPT文本提示、RKLLM_INPUT_TOKENtoken 序列、RKLLM_INPUT_EMBEDembedding 向量、RKLLM_INPUT_MULTIMODAL图文多模态推理模式枚举RKLLMInferModeRKLLM_INFER_GENERATE文本生成、RKLLM_INFER_GET_LAST_HIDDEN_LAYER取末层隐藏状态、RKLLM_INFER_GET_LOGITS取 logits回调状态枚举LLMCallStateRKLLM_RUN_NORMAL/WAITING/FINISH/ERROR参数结构体RKLLMParam、RKLLMExtendParam、RKLLMSamplingParam、RKLLMLoraAdapter、RKLLMInferParam、RKLLMResult等。RKLLM类flask_server.py封装了初始化、推理、释放的全过程并暴露set_chat_template、set_function_tools、rkllm_load_lora、rkllm_load_prompt_cache等能力。5.2 默认参数与 CPU 亲和性RKLLM.__init__中设置了初始化默认值flask_server.pymax_context_len 4096、max_new_tokens 4096top_k 1、top_p 0.9、temperature 0.8、repeat_penalty 1.1、frequency_penalty 0、presence_penalty 0skip_special_token True、ignore_eos_token False、is_async False扩展参数embed_flash 1embedding 从 flash 查询、n_batch 1、use_cross_attn 0、enabled_cpus_num 4CPU 亲和性差异rk3576/rk3588平台启用 CPU4–CPU7大核掩码(14)|(15)|(16)|(17)其他平台启用 CPU0–CPU3与rkllm.h中的CPU0~CPU7宏定义一一对应。5.3 回调驱动与流式输出推理采用回调驱动模式C 侧每次生成一段文本就回调 Python 侧注册的callback_implflask_server.pyPython 侧把累积的文本存入全局global_text列表。generate_stream()在一个工作线程中执行rkllm_run主线程循环轮询global_text并逐段 yield由 Flask 的stream_with_context以 SSEtext/event-stream格式输出每个 chunk 遵循 OpenAI 流式协议先发带role的首块中间逐段发送delta.content最后以finish_reason: stop和data: [DONE]收尾flask_server.py。5.4 OpenAI 兼容响应与并发控制非流式响应由build_openai_response()构造包含id、object: chat.completion、created、model、choices[0].message与usageprompt/completion tokens 数完全对齐 OpenAI 返回格式。服务端用一个全局threading.Lock加is_blocking标志控制并发推理期间其他请求会收到 503RKLLM_Server is busy!响应避免多用户同时抢占 NPU 资源flask_server.py。5.5 多轮消息追踪与多工具结果合并服务端通过全局_last_messages记录上一次请求的消息列表用新消息 本次消息 - 上次消息的方式提取增量输入get_last_input()flask_server.py。当检测到新一批连续tool角色消息时会将其内容合并为 JSON 数组后一次性送入模型这与execute_tool_calls()中把多个工具结果合并为[{...}, {...}]的逻辑闭环支撑了多工具并行调用的场景。5.6 Gradio 服务的实现要点gradio_server.py 与 Flask 版共用同一套 ctypes 绑定代码差异在于 UI 层它设置GRADIO_SERVER_NAME0.0.0.0、GRADIO_SERVER_PORT8080用gr.Blocks构建包含gr.Chatbot、gr.Textbox、清空按钮的聊天界面get_RKLLM_output()生成器在模型逐段输出时实时追加到 assistant 消息实现打字机效果同时兼容新版[{role: ..., content: ...}]与旧版[[user, assistant], ...]两种历史记录格式。客户端 chat_api_gradio.py 则通过gradio_client.Client依次调用/get_user_input与/get_RKLLM_output两个端点完成一轮问答。六、小结与延伸部署链路转换模型 → 推送到板卡 →build_rkllm_server_flask.sh或 Gradio 版一键安装依赖、推送运行时与启动服务 → 浏览器或 HTTP 客户端访问:8080。两套服务分工Flask 版面向程序化集成OpenAI 兼容、SSE 流式、Function CallingGradio 版面向快速演示与人工交互。性能前置条件服务启动时会执行fix_freq_platform.sh锁频并依赖 NPU 驱动rknpu-driver确保在 RK3588/RK3576 等平台上获得稳定的推理表现。可扩展方向修改 flask_server.py 中的采样默认值或增加鉴权逻辑即可适配更多生产场景若要接入其他模型家族重点是替换系统提示词模板并调整parse_tool_calls()的解析正则。赞分享大模型本地部署推理引擎模型量化嵌入式【免费下载链接】rknn-llm项目地址https://gitcode.com/gh_mirrors/rk/rknn-llm点击查看免费下载相关推荐PaddleNLP 推理服务化快速上手基于 Flask Gradio 的动态图 OpenAI 兼容 API 部署实战PaddleNLP 推理服务化快速上手基于 Flask Gradio 的动态图 OpenAI 兼容 API 部署实战 导读 本文围绕 PaddleNLP人工智能大模型预训练微调LoRARLHF强化学习分布式训练模型推理服务推理引擎模型量化模型压缩本地部署NLP产品管理工作流实战指南Product-Management 插件把 Claude 变成你的专属产品经理产品管理工作流实战指南Product Management 插件把 Claude 变成你的专属产品经理 从一个周五下午开始 周五下午你要同时交出一份工程在等AI 技能AI 插件OpenCore Legacy Patcher实操指南在老款Intel Mac上安装最新macOSOpenCore Legacy Patcher实操指南在老款Intel Mac上安装最新macOS 如果一台2012年的MacBook还留在High Sier操作系统固件驱动开发上一篇终极指南如何在墨水屏设备上打造流畅的Android启动器体验下一篇高效视频下载利器yt-dlp-gui完整使用指南创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考