📖 Day 6 学习页:Agent 架构原理:手写 ReAct
Day 6:Agent 架构原理——纯 Python 手写 ReAct
🎯 今日目标:不依赖任何框架,手写一个 ReAct Agent 循环。写完你会发现: LangChain/AutoGen 这些框架的核心,就是一个几十行的循环 + 一段格式约定。
⏱ 今日安排(约 2 小时)
| 时间 | 内容 | 方式 |
|---|---|---|
| 60min | 理论学习:Agent 四大模块 + ReAct 循环 | 阅读 + 手推 |
| 60min | 动手实操:完成 5 个任务,跑通 ReAct | 编码 |
📖 第一步:理论学习
1. Agent 到底是什么(四大模块)
| 模块 | 职责 | 本项目对应 |
|---|---|---|
| 规划 Planning | 把目标拆成步骤,决定下一步做什么 | ReAct 循环里的 Thought |
| 记忆 Memory | 记住对话历史和长期知识 | Day 9 实现 |
| 工具 Tools | 调用计算器/搜索/文件等扩展能力 | TOOLS_REGISTRY |
| 执行 Action | 真正去做(调 API、执行代码) | execute_tool |
LLM 本身只会生成文本。Agent = LLM + 循环 + 工具:让 LLM 的"文本输出"能触发真实的函数执行,再把执行结果喂回去。
2. ReAct:Reasoning + Acting 交替
核心格式约定——LLM 按这个格式输出,你的程序按这个格式解析:
Thought: 我需要先算出 3 + 5 * 2 的结果 ← LLM 推理"下一步干什么"
Action: Calculator ← 选择哪个工具
Action Input: 3 + 5 * 2 ← 工具参数
Observation: 13 ← 【你的程序】执行工具后填回去
Thought: 我已经拿到结果了
Final Answer: 3 + 5 * 2 = 13 ← 循环出口
一轮完整流程:构建 prompt(含工具清单 + 历史步骤)→ LLM 输出 → 解析 → 是 Final Answer 就结束,否则执行工具 → 把 Action/Observation 追加进历史 → 再来一轮。
3. 手写循环(伪代码,练习就照这个骨架写)
async def react_agent(user_input, tools, max_steps=10):
steps = []
for step in range(max_steps): # ← 必须有上限,防死循环
prompt = build_prompt(user_input, steps) # 工具清单 + 历史 Thought/Action/Observation
response = await llm_call(prompt)
thought, action, action_input = parse(response)
if action 为空: # LLM 给了 Final Answer
return action_input # ← 循环唯一正常出口
result = execute_tool(action, action_input)
steps.append(AgentStep(thought, action, action_input, observation=result))
return "达到最大步数限制" # 防御性兜底
模板内置的 mock LLM 规则:prompt 里最后一个 Action 后面还没有 Observation → 返回工具调用; 已经有了 → 返回 Final Answer(引用工具结果)。所以只要你的循环正确把历史拼进 prompt, 第 2 轮就会拿到最终答案。
🛠 第二步:动手实操
- 打开模板:day06_practice_template.py,5 个任务:
- TODO 1.2:
get_tools_description()—— 生成放进 prompt 的工具清单 - TODO 1.3:
execute_tool()—— 返回 ToolResult(成功/失败都要兜住) - TODO 2.1:
build_react_prompt()—— 系统说明 + 工具清单 + 历史 steps + Question - TODO 3.1:
parse_react_response()—— 正则提取 Thought/Action/Action Input;有 Final Answer 时 action 返回空串、答案放 action_input - TODO 4.1:
react_agent()主循环(按伪代码骨架写)
- TODO 1.2:
- 运行自测:
python day06_practice_template.py,两个测试都应打出最终回答而不是"达到最大步数"。 - 验收:
python day06_practice_validator.py—— 验收会真跑你的 Agent:计算题答案里必须有 13、 天气题必须有查询结果,Agent 卡死循环会被直接判负。
⚠️ 避坑
- ❌ 不设
max_steps→ LLM 反复要求调工具,无限循环烧钱 - ❌ 历史步骤(Action/Observation)不拼回 prompt → LLM 看不到自己做过什么,永远重复第一步
- ❌ 只解析成功路径 → 工具报错时 Agent 崩溃;把错误信息作为 Observation 喂回去,LLM 会自我纠正
- ✅ 每步记录 token 消耗(Day 10 会做),成本要可见