本章目標與前置條件
讀完後,你應能追蹤一次 agent run 從 API 建立、outbox event、worker claim 到 ADK/tool execution 的路徑,並知道目前 web assistant 仍是 fixture、唯一的 tool 仍是 stub。前置條件是理解非同步工作、owner scope、replay、tool authorization 與安全錯誤輸出。
功能現況
| 區塊 | 後端已核實功能 | 前端現況 |
|---|---|---|
| Run API | POST /v1/agents/runs、GET /v1/agents/runs、GET /v1/agents/runs/:runId;point 為 agent.run.create/agent.run.read,無 page、非 tenant-wide | 沒有對應 live API adapter;路由 registry 宣告 /assistants(agent.assistant)與 /history(agent.history),兩頁都是 RoutePlaceholder |
| 建立與 ownership | input、response_language validation;queued snapshot;creator/tenant ownership | 浮動面板的 fixture session 只保存本地 messages,不建立 run |
| 非同步執行 | agent requested outbox event、worker claim、ADK session、deadline/retry/terminal failure;agent_run_reaper 收斂逾時 run | 尚未接入 /v1/agents/** |
| replay/安全閘門 | PostgreSQL step checkpoint、fingerprint、cycle/budget/tool decision gate、tool 熔斷 | UI 只呈現 pending 與示範回覆 |
| 模型與 tool | worker 經 LiteLLM gateway(或開發用 fake)呼叫模型;唯一 tool knowledge_search 是 stub | 表單頁的 FormAssistantPane 是唯讀的「AI 表單助理尚未提供服務」面板,不呼叫任何 API |
| output/usage | state、output/error、prompt/completion tokens、model/tool calls、step count、timestamps;output.evidence 目前恆為 [] | 無 live history/run detail 頁面 |
前端入口由 PlatformShell 決定:首頁的 AI 鈕只導向 /assistants(HOME_AI_TARGET,佔位頁),不開面板;其他非工作區頁面開關浮動面板 AiChatPanel;工作區外殼不顯示 AI 鈕。浮動面板的會話來自 useAiChatFixtureSession,送出後 400ms 回固定示範文字。前端 fixture 與元件註解仍寫「/v1/agents/** 不在 registry.golden」,這已過時:後端已註冊三條路由,目前缺的是前端 adapter。
核心概念
API 只建立 immutable request snapshot
create 的 body 是 closed object,只收 input 與 response_language。handler 只檢查 input 必須出現且非 null、response_language 若出現不得為 null;內容由 service 驗證:input 為 1–4000 個 Unicode 字元且不得全為空白,response_language 為 1–35 字,違反回 422/10430。service 取一次持久化精度時間(UTC、微秒),寫入 queued run(deadline_at 為建立後 180 秒),再在同一 tenant transaction publish agent.run.requested。回應是 HTTP 201 的 queued snapshot,不是模型已完成。
Run 由建立者讀取
list 只接受 page、page_size(1–100,預設 20)與 sort,每個 key 只能出現一次,sort 只接受 created_at desc;其他 key 或重複 key 回 400/10001。service 以 creator ownership 讀取,依 created_at DESC, id DESC 排序。get 也按 tenant/creator 過濾:非建立者(包含管理員)讀取回 404/10050,與不存在的 id 相同;403 代表授權本身未通過,例如沒有 agent.run.read。失敗的 run 仍回 HTTP 200,失敗內容在 state 與 error。管理員要看別人的 run 必須另立權限與 API,不可把 agent.run.read 任意擴成 tenant-wide。
Worker 以 durable evidence 恢復
executor claim 後載入 run steps,先 restore replay checkpoint,再執行 ADK iterator。每個 model、tool 與 circuit step 都是 PostgreSQL checkpoint;tool 與 circuit step 保存 fingerprint、decision、result code 與 sequence。以下情況會讓 run 以 failed 結束並寫入可分類的 error,不回傳「空輸出」混淆失敗:
- step budget 用盡(60401)、tool 呼叫的 fingerprint 形成重複循環(60400,附
tool_name、cycle_length、step_no)、超過deadline_at(60402)。 - 模型不可用且已是最後一次投遞(60300;還有投遞次數時交由 outbox 重投)、輸出被截斷(60301)、輸出或 replay 無法安全解析或沒有最終文字(60302)、呼叫未註冊的 tool(60200)。
- 設有
AbortOnDeny的 tool 被 PDP 拒絕,run 以該拒絕碼結束。
tool 執行本身出錯不會直接終止 run:錯誤會轉成受控結果(error_code 60302、note tool failed,不含原始錯誤)回給模型;同一 tool 連續失敗 3 次後熔斷,寫入 circuit step 並回傳 step result code 60403。另外,worker 的 agent_run_reaper job 每分鐘把超過 deadline_at 仍是 queued/running 的 run 收斂為 failed/60402,不依賴 outbox dead letter。
這些上限都是程式常數(internal/service/agent/limits.go),不是環境設定:step budget 25、deadline 180 秒、循環偵測週期 1–4 且連續重複 3 次、熔斷門檻 3。
Tool gate 與 elevation 是逐次判定
outbox payload 只有 run_id、tenant_id、created_by 與可選的 elevation_session_id(durable session 引用),不攜帶 role、grant 或權限快照;worker 以 DisallowUnknownFields 解碼,多出欄位即視為 permanent failure。worker 在每次 tool decision 都用當下事實重新詢問 PDP,包含 elevation session 是否仍有效;agent 不能因為建立時有權限就永久持有。輸出 error 只保留安全診斷欄位:code、message,以及視情況出現的 tool_name、step_no、cycle_length,不公開 token、provider 原始錯誤或秘密。
模型、tool 與設定
API process 不建立 model client;只有 worker(cmd/worker/agent_runtime.go)組裝 executor 與 reaper。LLM_PROVIDER 預設 litellm,經 LiteLLM gateway 呼叫,設定 LITELLM_BASE_URL、LITELLM_API_KEY(secret)與 LITELLM_CHAT_MODEL(預設 alias nexus-agent-fallback);fake 是 scripted provider,只供開發與測試(E2E 使用它),staging/production 會拒絕啟動。LITELLM_EMBEDDING_MODEL 與 LITELLM_MASTER_KEY 目前沒有 consumer。
tool registry 目前只有 knowledge_search(policy point knowledge.document.read,設有 Idempotent 與 AbortOnDeny),實作是 stub:只回傳新的 retrieval_id、空 evidence 與 note knowledge retrieval lands with KB-2。因此 succeeded run 的 output.evidence 恆為 [],回答只是模型產生的文字。registry 會拒絕設定 RequireConfirmation 的 tool(確認流程尚未支援);tool 的 policy point 必須由既有 route 宣告,API 與 worker 啟動時都會驗證;tool 沒有個別 timeout,整體受 run deadline 限制。
一條非同步請求鏈
正在繪製架構圖…
查看圖表原始碼
flowchart LR
Client["assistant client(前端尚未接入)"] --> API["POST /v1/agents/runs"]
API --> Run[("queued agent_runs row")]
API --> Outbox[("agent.run.requested")]
Outbox --> Worker["agent worker"]
Worker --> Checkpoint[("agent_run_steps checkpoint")]
Worker --> ADK["ADK + LiteLLM + knowledge_search stub"]
ADK --> Finish[("succeeded / failed snapshot")]
Reaper["agent_run_reaper"] --> Finish
Client --> History["GET owned runs(輪詢)"]
History --> Finish實作步驟與示例
- 先確認 agent run 是新 API、worker tool、replay schema 還是 web adapter,分別讀 API/service/worker/domain 檔案;路由 policy(point、page、tenant_wide)以
internal/route/testdata/registry.golden為準。 - 新增 request/output 欄位時,同步 OpenAPI、renderAgentRun、前端 zod schema(接入後)與 contract test;不要讓
nil同時表示尚未完成與解析失敗。 - 新增 tool 時先定義 input schema、normalization(
NormalizedOut)、fingerprint、policy point、Idempotent/AbortOnDeny與失敗分類,確認該 point 已由 route 宣告,再接入 executor;不要在 ADK callback 直接寫跨域業務資料。 - 驗證要檢查 queued → running → succeeded/failed 的持久化 snapshot、step replay、audit 與同一 event 的重放行為。
POST /v1/agents/runs
Content-Type: application/json
{"input":"請整理本月已核准的請假資料","response_language":"zh-Hant"}此 body 只展示欄位形狀;不要把真實 HR 名單、token、模型 prompt 或內部帳號寫進測試文件。預期第一個 response 是 HTTP 201 的 queued snapshot。目前沒有 SSE、取消或確認端點,只能輪詢合法 owner 的 GET /v1/agents/runs/:runId,直到 state 為 succeeded 或 failed,再用 output/error 和 usage 分辨結果。狀態詞彙另有 awaiting_confirmation 與 cancelled,目前流程不會進入。
驗證矩陣
| 驗證問題 | 程式碼/測試證據 | 仍需另外做的驗收 |
|---|---|---|
| request validation/owner scope | internal/api/v1/agent_runs_test.go、internal/api/v1/agent_history_test.go、internal/service/agent/service_test.go;tests/integration/postgres/agent_history_test.go、agent_rls_test.go;E2E tests/e2e/agent_runs_test.go(預設角色可建立並輪詢、非建立者 404) | 部署環境的 authenticated API 與權限資料 |
| outbox/worker claim | internal/worker/agent.go、internal/worker/agent_test.go、internal/service/agent/executor_test.go;tests/integration/postgres/agent_fence_test.go、agent_lifecycle_test.go(重試、dead letter 後由 reaper 收斂) | 部署環境的 worker、queue 與實際 deadline/retry |
| replay/fingerprint/gate | internal/service/agent/executor_replay_test.go、internal/service/agent/executor_gates_test.go、internal/service/agent/gates_conversion_test.go;tests/integration/postgres/agent_recovery_test.go、agent_worker_authz_test.go(真實 PDP、撤權後重檢、elevation 引用) | 真實 LiteLLM model 與非 stub tool |
| output rendering/safe error | internal/api/v1/agent_runs_test.go、internal/service/agent/replay_validation_test.go;E2E 的 loop run 驗證 60400 診斷欄位 | 日誌 redaction、部署 runtime 與 observability |
| web assistant | app/(platform)/assistants/page.tsx、app/(platform)/history/page.tsx、app/(platform)/_components/useAiChatFixtureSession.ts、app/(platform)/_fixtures/aiChat.ts | /assistants/[id] 與 live /v1/agents/** adapter 尚待實作/核實 |
PostgreSQL 整合測試使用 scripted 或測試專用的 model,E2E 使用 LLM_PROVIDER=fake;兩者都不代表真實 LiteLLM 或知識檢索已驗收。
常見錯誤
- 收到 201 queued 就把回覆當完成;必須再讀 run state。
- 讓 list 接受任意
account_id或 tenant filter,破壞 creator-owned contract。 - 重試 tool 卻換 fingerprint/input,導致 replay 與副作用去重失效。
- 把 elevation session 內容、權限快照、原始 prompt 或 credential 放入 outbox 或 error message;outbox 只允許
elevation_session_id這個引用。 - 以浮動面板 fixture 的 400ms 罐頭回覆宣稱 AI Agent 已接通;目前前端仍是示範流程,首頁 AI 鈕只會導向
/assistants佔位頁。 - 把
output.evidence為空解讀成「查無相關知識」;目前knowledge_search是 stub,永遠不回 evidence。 - 以為 tool 出錯會讓 run 立刻 failed;錯誤其實會以受控結果回給模型,同一 tool 連續失敗 3 次才熔斷。