)
1. 为什么 GPU 训练完的模型换台机器就加载失败你大概率遇到过这个场景在带显卡的服务器上把模型训练好torch.save存成.pth或.pkl然后拿到一台没有 GPU 的笔记本或者纯 CPU 的推理服务器上跑结果一执行torch.load就报错提示类似Attempting to deserialize object on a CUDA device but torch.cuda.is_available() is False。这不是模型坏了而是保存时张量里记录了它所在的设备位置加载时 PyTorch 默认想把它放回原来的 CUDA 设备可目标机器根本没有 CUDA于是直接抛异常。PyTorch 训练与推理分离是工程里非常常见的部署形态训练侧追求算力用 GPU 集群跑几天推理侧追求成本和稳定往往跑在 CPU 机器、容器或边缘设备上。两者之间靠一个模型文件传递而设备映射device mapping就是这条链路上最容易踩的坑。这篇内容聚焦的就是这个分离场景交付可复制的map_location配置片段、torch.save/load的验证动作以及一套 TaoToken 统一 Key/API 通道的settings.json骨架让你在没有 GPU 的环境里也能稳定跑通已经训练好的模型。适合谁看刚接触 PyTorch 部署、被设备报错卡住的初学者需要把训练产物交付给 CPU 推理服务的工程师以及想给推理服务加一层统一模型调用通道的人。下面从设备选择、保存加载、验证到排错一步步来。2. TaoToken 前置给推理服务准备统一调用通道模型在 CPU 上跑起来之后通常还要接一个模型服务或对话接口。如果你的推理服务需要调用外部大模型能力或者想把本地模型和云端模型统一管理TaoToken 可以作为统一 Key/API 通道来用。它的官网入口是 https://taotoken.net/?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_content API 基址是 https://taotoken.net/api 这个地址不加 UTM 参数。你需要先拿到一个 API Key在控制台里创建即可https://taotoken.net/console?utm_sourcetaotoken_aicg_blog_endutm_contentconsoleutm_campaignrewrite 。创建完 Key 之后可以在 API Keys 页面管理https://taotoken.net/api-keys?utm_sourcetaotoken_aicg_blog_endutm_contentapi-keysutm_campaignrewrite 。接入文档在 https://taotoken.net/doc?utm_sourcetaotoken_aicg_blog_endutm_contentdocutm_campaignrewrite 里面有各语言的调用示例。这里要强调一点TaoToken 是合规的 API 通道不是让你去搞什么灰色中转也不是替代你的编辑器或训练框架。它的定位是给推理服务提供一个统一的模型调用入口本地 CPU 推理和云端模型调用可以共用一套配置。如果你只是想验证模型效果可以直接用模型对话页面https://taotoken.net/model-chat?utm_sourcetaotoken_aicg_blog_endutm_contentmodel-chatutm_campaignrewrite 。如果你长期做编码或 Agent 类任务可以看 Coding Planhttps://taotoken.net/coding-plan?utm_sourcetaotoken_aicg_blog_endutm_contentcoding-planutm_campaignrewrite 。拿到 Key 之后我们把它写进一个settings.json骨架里后面推理脚本读这个文件就行。这样训练侧和推理侧解耦Key 不硬编码在代码里。{ taotoken: { api_base: https://taotoken.net/api, api_key: sk-你的Key填这里, default_model: claude-sonnet-4-20250514, timeout: 60, max_retries: 3 }, inference: { device: cpu, model_path: ./checkpoints/best_model.pth, batch_size: 4, num_workers: 0 } }这个骨架里inference.device固定为cpunum_workers在纯 CPU 环境建议设 0避免多进程加载模型时出现句柄问题。taotoken段就是统一通道的配置推理脚本里读出来直接用。3. 可复制配置GPU 训练保存与 CPU 加载的完整片段先看训练侧。GPU 训练时设备选择要写得健壮一点不要硬编码cuda:0否则换机器就崩。import torch import torch.nn as nn # 设备选择有 GPU 用 GPU没有就退回 CPU device torch.device(cuda if torch.cuda.is_available() else cpu) print(f训练设备: {device}) if device.type cuda: print(f可用 GPU 数量: {torch.cuda.device_count()}) print(f当前 GPU: {torch.cuda.get_device_name(0)}) class Net(nn.Module): def __init__(self): super().__init__() self.fc1 nn.Linear(784, 256) self.fc2 nn.Linear(256, 10) def forward(self, x): x torch.relu(self.fc1(x)) return self.fc2(x) net Net().to(device)多卡场景下如果你只想用第二块卡可以写torch.device(cuda:1)GPU 编号从 0 开始。但保存模型时建议统一把参数搬到 CPU 再存这样最省心。# 训练循环省略假设已经训练完成 # 关键动作保存前把 state_dict 搬到 CPU state_dict_cpu {k: v.cpu() for k, v in net.state_dict().items()} torch.save(state_dict_cpu, ./checkpoints/best_model.pth) print(模型已保存到 CPU 张量跨设备加载无压力)如果你已经存了 GPU 张量的模型也不用重新训练加载时用map_location映射即可。下面是 CPU 推理侧的完整加载片段。import json import torch import torch.nn as nn # 读取 settings.json with open(./settings.json, r, encodingutf-8) as f: settings json.load(f) infer_cfg settings[inference] device torch.device(infer_cfg[device]) # 这里是 cpu print(f推理设备: {device}) class Net(nn.Module): def __init__(self): super().__init__() self.fc1 nn.Linear(784, 256) self.fc2 nn.Linear(256, 10) def forward(self, x): x torch.relu(self.fc1(x)) return self.fc2(x) # 定义相同结构的网络不要 to(device) 之前就 load net Net() # 核心map_location 把 GPU 张量映射到 CPU state_dict torch.load( infer_cfg[model_path], map_locationdevice, weights_onlyTrue ) net.load_state_dict(state_dict) net.to(device) net.eval() print(模型加载完成已切换到 eval 模式)这里有几个细节值得说清楚。第一map_locationdevice里的device是torch.device(cpu)PyTorch 会把原来记录在 CUDA 上的张量重新映射到 CPU 内存。第二weights_onlyTrue是较新版本 PyTorch 的推荐做法只加载权重避免反序列化执行任意代码安全性更好。第三net.eval()一定要调用否则 BatchNorm 和 Dropout 在推理时行为不对结果会飘。推理时数据也要注意设备一致# 模拟一条输入 x torch.randn(1, 784).to(device) with torch.no_grad(): output net(x) pred output.argmax(dim1) print(f预测类别: {pred.item()})torch.no_grad()关闭梯度计算CPU 推理能省不少内存和时间。4. 验证请求确认模型真的在 CPU 上跑通加载完不代表跑通得验证。最直接的办法是检查模型参数的设备类型。# 验证 1检查参数所在设备 for name, param in net.named_parameters(): print(f{name}: device{param.device}, dtype{param.dtype}) break # 看第一个就够了 # 验证 2确认没有 CUDA 依赖 assert not any(p.is_cuda for p in net.parameters()), 还有参数在 GPU 上 print(验证通过所有参数都在 CPU) # 验证 3跑一次前向确认输出形状 with torch.no_grad(): dummy torch.randn(2, 784) out net(dummy) print(f输出形状: {out.shape}) # 期望 torch.Size([2, 10])如果输出形状对得上参数设备也对基本就稳了。接下来把 TaoToken 通道接上验证统一调用是否可用。下面是一个最小请求示例用requests直接打 API。import json import requests with open(./settings.json, r, encodingutf-8) as f: settings json.load(f) cfg settings[taotoken] headers { Authorization: fBearer {cfg[api_key]}, Content-Type: application/json } payload { model: cfg[default_model], messages: [ {role: user, content: 用一句话说明 CPU 推理的优势} ], max_tokens: 128 } resp requests.post( f{cfg[api_base]}/v1/messages, headersheaders, jsonpayload, timeoutcfg[timeout] ) print(f状态码: {resp.status_code}) print(resp.json())成功的话你会看到状态码 200返回体里有模型输出。这一步验证的是统一 Key/API 通道通了和本地 CPU 模型是两条并行的能力本地模型负责你自己的网络TaoToken 通道负责需要大模型能力的部分。两者在同一个推理服务里共存配置都从settings.json读。如果你更想先手动试一下模型对话效果可以直接打开 https://taotoken.net/model-chat?utm_sourcetaotoken_aicg_blog_endutm_contentmodel-chatutm_campaignrewrite 在网页里发一条消息确认 Key 和模型名没问题再回到代码里调。5. 本篇常见错排查报错一RuntimeError: Attempting to deserialize object on a CUDA device but torch.cuda.is_available() is False这是最典型的。原因就是torch.load没加map_locationPyTorch 想还原到 CUDA。解决加上map_locationtorch.device(cpu)或map_locationcpu。如果你已经按第 3 节保存时搬到 CPU就不会遇到这个。报错二KeyError: unexpected key module.xxx这通常是用nn.DataParallel或 DDP 训练保存的state_dict 的 key 带了module.前缀。加载时要么去掉前缀要么给模型也包一层。处理方式state_dict torch.load(path, map_locationcpu) # 去掉 module. 前缀 new_state_dict {k.replace(module., ): v for k, v in state_dict.items()} net.load_state_dict(new_state_dict)报错三RuntimeError: Expected all tensors to be on the same device推理时输入数据还在 CPU但模型某层被意外搬到了别处或者反过来。检查net.to(device)和x.to(device)是否用了同一个 device 对象。CPU 场景下两者都应该是cpu。报错四加载后推理结果和训练时差很多八成是忘了net.eval()。训练模式下 Dropout 会随机丢弃BatchNorm 用 batch 统计量推理时结果自然不稳定。加上net.eval()和torch.no_grad()再试。报错五num_workers导致卡死或报句柄错误纯 CPU 环境、Windows 或容器里DataLoader的num_workers设大于 0 有时会出问题。按settings.json里设成 0先保证跑通再按需调大。报错六TaoToken 请求返回 401 或 403检查settings.json里的api_key是否填对有没有多余空格。Key 在 https://taotoken.net/api-keys?utm_sourcetaotoken_aicg_blog_endutm_contentapi-keysutm_campaignrewrite 管理如果泄露了及时重置。请求头里Authorization的格式是Bearer sk-xxx别漏了 Bearer。报错七weights_onlyTrue导致加载失败老版本 PyTorch 不支持这个参数或者你的模型文件里存了非张量对象。可以先去掉weights_onlyTrue试确认能加载后再评估是否需要用新格式重新保存。生产环境建议用state_dict方式保存天然兼容weights_only。6. 把训练和推理彻底解耦的落地建议走到这里你已经有了完整的链路GPU 侧训练、保存时搬到 CPU、CPU 侧用map_location加载、验证参数设备、接上 TaoToken 统一通道。最后给几条实操建议都是踩过坑之后总结的。保存模型时统一用state_dict而不是整个模型对象这样加载侧只需要相同的网络结构定义不依赖训练时的代码路径。保存前把张量.cpu()一下能省掉加载侧很多麻烦。推理脚本里设备从settings.json读不要写死这样同一份代码在 GPU 机器和 CPU 机器上都能跑只是配置不同。TaoToken 的 Key 也放settings.json不要提交到 Git。可以用环境变量覆盖比如启动时读TAOTOKEN_API_KEY优先级高于文件里的值。这样本地开发用文件线上用环境变量安全又灵活。如果你后面要把推理服务容器化记得基础镜像选 CPU 版 PyTorch别选 CUDA 版镜像能小好几个 G。启动命令里挂载settings.json和模型文件容器里不需要 GPU 运行时。需要长期跑编码或 Agent 任务的话Coding Plan 那条通道可以单独配一个 Key和推理服务的 Key 分开管理方便做配额和审计https://taotoken.net/coding-plan?utm_sourcetaotoken_aicg_blog_endutm_contentcoding-planutm_campaignrewrite 。接入细节都在文档里遇到报错先翻文档再排查https://taotoken.net/doc?utm_sourcetaotoken_aicg_blog_endutm_contentdocutm_campaignrewrite 。整套配置跑通之后你会发现 GPU 训练和 CPU 推理之间就隔了一个map_location的距离。把这个骨架固化到项目模板里下次换机器部署就不用再折腾设备映射了。