📖 Day 6 学习页:Agent 架构原理:手写 ReAct

Day 6:Agent 架构原理——纯 Python 手写 ReAct

🎯 今日目标:不依赖任何框架,手写一个 ReAct Agent 循环。写完你会发现: LangChain/AutoGen 这些框架的核心,就是一个几十行的循环 + 一段格式约定。

📚 本日资源:📝 练习模板 | ✅ 验收脚本

⏱ 今日安排(约 2 小时)

时间内容方式
60min理论学习:Agent 四大模块 + ReAct 循环阅读 + 手推
60min动手实操:完成 5 个任务,跑通 ReAct编码

📖 第一步:理论学习

1. Agent 到底是什么(四大模块)

模块职责本项目对应
规划 Planning把目标拆成步骤,决定下一步做什么ReAct 循环里的 Thought
记忆 Memory记住对话历史和长期知识Day 9 实现
工具 Tools调用计算器/搜索/文件等扩展能力TOOLS_REGISTRY
执行 Action真正去做(调 API、执行代码)execute_tool

LLM 本身只会生成文本。Agent = LLM + 循环 + 工具:让 LLM 的"文本输出"能触发真实的函数执行,再把执行结果喂回去。

2. ReAct:Reasoning + Acting 交替

核心格式约定——LLM 按这个格式输出,你的程序按这个格式解析:

Thought: 我需要先算出 3 + 5 * 2 的结果        ← LLM 推理"下一步干什么"
Action: Calculator                            ← 选择哪个工具
Action Input: 3 + 5 * 2                       ← 工具参数
Observation: 13                               ← 【你的程序】执行工具后填回去
Thought: 我已经拿到结果了
Final Answer: 3 + 5 * 2 = 13                  ← 循环出口

一轮完整流程:构建 prompt(含工具清单 + 历史步骤)→ LLM 输出 → 解析 → 是 Final Answer 就结束,否则执行工具 → 把 Action/Observation 追加进历史 → 再来一轮。

3. 手写循环(伪代码,练习就照这个骨架写)

async def react_agent(user_input, tools, max_steps=10):
    steps = []
    for step in range(max_steps):            # ← 必须有上限,防死循环
        prompt = build_prompt(user_input, steps)   # 工具清单 + 历史 Thought/Action/Observation
        response = await llm_call(prompt)
        thought, action, action_input = parse(response)

        if action 为空:                      # LLM 给了 Final Answer
            return action_input              # ← 循环唯一正常出口

        result = execute_tool(action, action_input)
        steps.append(AgentStep(thought, action, action_input, observation=result))

    return "达到最大步数限制"                 # 防御性兜底

模板内置的 mock LLM 规则:prompt 里最后一个 Action 后面还没有 Observation → 返回工具调用; 已经有了 → 返回 Final Answer(引用工具结果)。所以只要你的循环正确把历史拼进 prompt, 第 2 轮就会拿到最终答案。

🛠 第二步:动手实操

  1. 打开模板:day06_practice_template.py,5 个任务:
    • TODO 1.2:get_tools_description() —— 生成放进 prompt 的工具清单
    • TODO 1.3:execute_tool() —— 返回 ToolResult(成功/失败都要兜住)
    • TODO 2.1:build_react_prompt() —— 系统说明 + 工具清单 + 历史 steps + Question
    • TODO 3.1:parse_react_response() —— 正则提取 Thought/Action/Action Input;有 Final Answer 时 action 返回空串、答案放 action_input
    • TODO 4.1:react_agent() 主循环(按伪代码骨架写)
  2. 运行自测:python day06_practice_template.py,两个测试都应打出最终回答而不是"达到最大步数"。
  3. 验收:python day06_practice_validator.py —— 验收会真跑你的 Agent:计算题答案里必须有 13、 天气题必须有查询结果,Agent 卡死循环会被直接判负。

⚠️ 避坑

  • ❌ 不设 max_steps → LLM 反复要求调工具,无限循环烧钱
  • ❌ 历史步骤(Action/Observation)不拼回 prompt → LLM 看不到自己做过什么,永远重复第一步
  • ❌ 只解析成功路径 → 工具报错时 Agent 崩溃;把错误信息作为 Observation 喂回去,LLM 会自我纠正
  • ✅ 每步记录 token 消耗(Day 10 会做),成本要可见

🔗 延伸资源