Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
f347c150a4 |
@@ -87,12 +87,11 @@ locode is a full-screen terminal app built with [Ink](https://github.com/vadimde
|
||||
|
||||
- **Ctrl+O**: print the full text of the last `/compact` (or auto-compact) summary. The collapsed notice you see right after compacting only shows a one-line hint — press Ctrl+O any time afterward to print the whole thing.
|
||||
- **Ctrl+B**: while a `bash` command is running, detaches it into the background and returns control to you immediately — the turn continues with a `bash_output`-checkable job id instead of waiting for the command to finish. A notice appears in the transcript once the backgrounded command actually completes. Only `bash` supports this today. The model can kill a still-running backgrounded job with `bash_kill`; any jobs still running when locode itself exits are killed too, so they don't outlive the process as orphans.
|
||||
- **Ctrl+F**: open and focus a file panel docked to the right of the chat (hidden by default); press again to close it. It has two tabs — **Files**, the project's collapsible file tree (directories in cyan, same `node_modules`/`.git`/`dist` exclusions as `@` mentions), and **Activity**, the files `read_file`/`write_file`/`edit_file` have touched so far this session, most recent first, with a status glyph (`·` read, `+` written, `~` edited) and a repeat count. **Ctrl+G** switches between the two tabs. While the panel is focused, `↑`/`↓` move the selection (auto-scrolling to keep it in view), `↵`/`←`/`→` expand or collapse the selected folder, and **Esc** hands keyboard focus back to the chat input without closing the panel — typing is disabled while the panel has focus, so the same arrow key doesn't simultaneously recall chat history.
|
||||
|
||||
- **Backends**: `--backend ollama` (default) or `--backend lmstudio`, or `--base-url <url>` for anything else that speaks the same API.
|
||||
- **Tools**: `read_file`, `list_files`, `grep`, `definition`, `references`, `diagnostics`, `web_search`, `web_fetch`, `git_status`, `bash_output`, `todo_write`, `task_create`, `task_list`, `task_get`, `task_update` run automatically. `write_file`, `edit_file`, `multi_edit`, `bash`, `bash_kill`, and `git_commit` show a diff/preview in a bordered box and ask you to pick Yes / Yes-always-this-session / No with the arrow keys before running.
|
||||
- **Tools**: `read_file`, `list_files`, `grep`, `web_search`, `web_fetch`, `git_status`, `bash_output`, `todo_write` run automatically. `write_file`, `edit_file`, `bash`, `bash_kill`, and `git_commit` show a diff/preview in a bordered box and ask you to pick Yes / Yes-always-this-session / No with the arrow keys before running.
|
||||
- **Permission modes**: `default` (ask before every mutating tool), `plan` (research only — every mutating tool is blocked outright, no prompt; the model is expected to describe what it would do in its final answer instead), `auto-edit` (file edits auto-approved, `bash`/`git_commit` still ask), `auto-accept` (everything auto-approved — use with care). Cycle with `Shift+Tab` or set directly with `/perm <mode>`.
|
||||
- **Structured tasks**: for multi-step work the model can call `task_create`/`task_list`/`task_get`/`task_update` to track units of work with a dependency graph (`blocks`/`blockedBy`), ownership (`owner`), status, and free-form metadata — created incrementally rather than replaced wholesale. The older flat `todo_write` live checklist (`☐`/`◐`/`☑`) remains for simpler cases.
|
||||
- **Task checklists**: for multi-step work the model can call `todo_write` to show a live checklist (`☐`/`◐`/`☑`) in the transcript instead of silently working through a list you can't see progress on.
|
||||
- **Project instructions**: a `CLAUDE.md` (or `AGENTS.md`) file in the project root is automatically read at session start and folded into the system prompt — put repo-specific conventions there and every session picks them up without being told.
|
||||
- **Images**: `read_file` returns image files (png, jpg, jpeg, gif, webp, bmp — up to 5MB) as actual image content instead of trying to decode them as text, so vision-capable models can see them when the model itself calls the tool. To attach a file or image to your own message directly, use `/import <path> [caption]`.
|
||||
- **`@` file mentions**: type `@` in the chat input to open a fuzzy file picker (searches the whole project, skipping `node_modules`/`.git`/`dist`) — keep typing to filter, `↑`/`↓` to navigate, `Tab` (or `Enter`) to insert the highlighted path. Any `@path` left in your message when you hit `Enter` for real is resolved against disk and attached to that message (text inlined, images attached as image content) — a stray `@` that isn't an actual file (e.g. an email address) is left as plain text.
|
||||
@@ -101,12 +100,10 @@ locode is a full-screen terminal app built with [Ink](https://github.com/vadimde
|
||||
- **MCP servers**: locode connects to any [MCP](https://modelcontextprotocol.io) servers configured via `locode mcp add` or a project's `.mcp.json` (stdio and remote/streamable-HTTP transports), and adds their tools to every session, namespaced as `mcp__<server>__<tool>`. Every MCP tool is treated as mutating (confirmation required on every call) regardless of what it reports — the MCP `readOnlyHint` annotation is advisory and could be wrong (or set by a malicious server specifically to skip confirmation), so locode never trusts it. One misconfigured server doesn't block the others — check `/mcp` for per-server connection status.
|
||||
- **Claude Code plugins**: `locode plugin add <path-or-git-url>` installs a Claude Code-compatible plugin — locode reads its `.claude-plugin/plugin.json`, then loads it all directly: MCP servers (its `.mcp.json` or manifest `mcpServers`, merged in like any other MCP server), slash commands (`commands/*.md` — frontmatter `description`/`argument-hint`, body is a template expanded with `$ARGUMENTS`/`$1..$9` and submitted as your message), agents (`agents/*.md` — the body becomes a sub-agent's system prompt, exposed as a callable tool named `agent__<plugin>__<agent>`; a `tools:` frontmatter list restricts what it can use, with Claude Code's built-in tool names — Read, Grep, Edit, etc. — automatically mapped to locode's equivalents), hooks (`hooks/hooks.json`, see below), and skills (`skills/*/SKILL.md`, see below). Check `/plugins` for what's loaded.
|
||||
- **Skills**: named instructions the model loads on demand rather than a hook or a sub-agent — every installed skill (`skills/<name>/SKILL.md`) is exposed through one shared `skill` tool, whose own description lists every skill's name and "use this when..." blurb so the model knows when to call it. You can also invoke one directly with `/<skill-name> [request]`, which skips the model's own judgment and submits the skill's instructions (plus your request, if any) as the turn. Check `/skills` for what's installed.
|
||||
- **Hooks**: shell commands that fire on session lifecycle events — `SessionStart`, `UserPromptSubmit`, `PreToolUse`, `PostToolUse`, `PermissionRequest`, `SubagentStart`, `SubagentStop`, `CwdChanged`, `FileChanged`, `ConfigChange`, `Stop`, `SessionEnd`. Configured the same way MCP servers are — plugin-bundled (`hooks/hooks.json`), user-level (`locode hooks path`, hand-edited), and project-level (`.locode/hooks.json`) all merge together, every hook from every source runs. A hook receives a JSON payload on stdin (`session_id`, `cwd`, `hook_event_name`, plus event-specific fields like `prompt` or `tool_name`/`tool_input`); exit 0 allows (stdout becomes injected context for `SessionStart`/`UserPromptSubmit`), exit 2 blocks (stderr is the reason shown), anything else is a non-blocking warning. `PreToolUse`, `UserPromptSubmit`, and `PermissionRequest` can block; the rest are fire-and-forget. `PreToolUse` fires before permission modes apply, so a hook's block can't be bypassed by auto-accept. `command` and `http` hook types run an external process/request; a `prompt` hook just injects its static `message` as additional context instead. Command/http hooks can opt into structured JSON output via `outputSchema: "json"`. `SessionEnd` fires on every exit path (including Ctrl+C and external `SIGTERM`/`SIGHUP`). Check `/hooks` for what's configured.
|
||||
- **Hooks**: shell commands that fire on session lifecycle events — `SessionStart`, `UserPromptSubmit`, `PreToolUse`, `PostToolUse`, `PermissionRequest`, `SubagentStart`, `SubagentStop`, `CwdChanged`, `FileChanged`, `ConfigChange`, `Stop`, `SessionEnd`. Configured the same way MCP servers are — plugin-bundled (`hooks/hooks.json`), user-level (`locode hooks path`, hand-edited), and project-level (`.locode/hooks.json`) all merge together, every hook from every source runs. A hook receives a JSON payload on stdin (`session_id`, `cwd`, `hook_event_name`, plus event-specific fields like `prompt` or `tool_name`/`tool_input`); exit 0 allows (stdout becomes injected context for `SessionStart`/`UserPromptSubmit`), exit 2 blocks (stderr is the reason shown), anything else is a non-blocking warning. `PreToolUse`, `UserPromptSubmit`, and `PermissionRequest` can block; the rest are fire-and-forget. `PreToolUse` fires before permission modes apply, so a hook's block can't be bypassed by auto-accept. `command` and `http` hook types are supported (`prompt` is declared in the config format but not yet executed); command hooks can opt into structured JSON output via `outputSchema: "json"`. `SessionEnd` fires on every exit path (including Ctrl+C and external `SIGTERM`/`SIGHUP`). Check `/hooks` for what's configured.
|
||||
- **Tool-calling mode**: on connect, locode probes whether the model reliably uses native OpenAI-style function calling. If not, it switches to a prompt-based fallback mode where the model is instructed to emit tool calls as fenced ` ```tool_call ``` ` JSON blocks, which locode parses itself. The result is cached per backend+model so future sessions skip the probe. Override with `--tool-mode native|fallback|auto` or the in-session `/mode` command.
|
||||
- **Context tracking & compaction**: the status bar shows `ctx NN%` — context window usage, from real `usage.prompt_tokens` when the backend reports it (requested via `stream_options.include_usage`), or a `~`-prefixed char-based estimate otherwise. The window size itself is auto-detected (Ollama's `/api/show`, then LM Studio's `/api/v0/models`) and cached per backend+model; falls back to a configurable default (`locode config set contextWindow <n>`, or `$LOCODE_CONTEXT_WINDOW`) if neither responds. At 85% usage, locode automatically asks the model to summarize the conversation and replaces the history with that summary (a notice tells you when this happens) — or trigger it yourself anytime with `/compact`.
|
||||
|
||||
- **Code intelligence (LSP)**: `definition`, `references`, and `diagnostics` use a real language server (LSP) for go-to-definition, find-all-references, and type/syntax error checks — the same engine an editor's Problems panel uses, more precise than `grep`. A server is lazily started per language on first use and reused for the whole session: `typescript-language-server` (TypeScript/JavaScript), `pyright-langserver` (Python), `gopls` (Go), `rust-analyzer` (Rust), and `clangd` (C/C++ — one clangd covers both). The relevant server binary must be on your PATH; if it isn't, the tool returns a clear "install X" error. After any edit, locode syncs the file to the live server so a subsequent `diagnostics` call reflects the change (it waits for the server to publish fresh diagnostics rather than reading a stale snapshot). Add or override servers with `locode config set lspServers` (see Config).
|
||||
|
||||
Note: even models with genuine native tool-calling support occasionally emit a tool call as plain text instead of a real structured call — this is model sampling variance, not a bug. If a turn seems to "describe" a tool call instead of running it, just ask again or try `/mode fallback`.
|
||||
|
||||
## Slash commands
|
||||
@@ -115,7 +112,6 @@ Note: even models with genuine native tool-calling support occasionally emit a t
|
||||
/model <name> switch the model used for the current backend
|
||||
/backend <name> switch backend (ollama | lmstudio), keeps current model
|
||||
/mode <name> view or force tool-call mode (native | fallback)
|
||||
/mouse [on|off] toggle mouse tracking (on by default; hold Shift+click/drag for native text selection)
|
||||
/perm [mode] cycle or set permission mode (default | plan | auto-edit | auto-accept)
|
||||
/status show current model, backend, tool-call mode, and cwd
|
||||
/dashboard show session stats: token I/O, elapsed/model time, turns, tool calls
|
||||
@@ -127,7 +123,7 @@ Note: even models with genuine native tool-calling support occasionally emit a t
|
||||
/hooks show configured hooks per lifecycle event
|
||||
/skills show installed skills; /<skill-name> [request] invokes one directly
|
||||
/compact summarize the conversation now to free up context
|
||||
/export [file] save the conversation as markdown (or /export json [file] for a full JSON dump incl. tool calls/results)
|
||||
/export [file] save the conversation as markdown — opens an editable filename prompt (default: locode-export-<timestamp>.md)
|
||||
/import <path> [caption] attach a local file or image to your next message
|
||||
/clear clear conversation history
|
||||
/help show this help
|
||||
@@ -147,11 +143,6 @@ locode config set autoCompactThreshold 0.85 # fraction of context window at whi
|
||||
locode config set requestTimeoutMs 300000 # per-request timeout in ms (default 180000); raise this if
|
||||
# your backend queues requests behind a concurrency limit
|
||||
# (e.g. Ollama's OLLAMA_NUM_PARALLEL) under multi-session load
|
||||
locode config set lspServers '{"java":{"command":"jdtls","extensions":[".java"]}}' # add a language server
|
||||
# (JSON object keyed by language id; built-in ids
|
||||
# override command/args, new ids add support and
|
||||
# require extensions). Built-ins: typescript,
|
||||
# python, go, rust, c (C/C++ share clangd).
|
||||
locode config get
|
||||
locode config path
|
||||
```
|
||||
@@ -160,12 +151,13 @@ locode config path
|
||||
|
||||
- Requires a real interactive terminal (TTY) — you can't pipe input into it or run it from a non-interactive script.
|
||||
- Native tool-calling reliability varies by model and is non-deterministic even for capable models (see above).
|
||||
- No OS-level sandboxing (no container/VM isolation) — mutating tools operate on the real filesystem/shell with the permissions of the user running `locode`. Only approve commands you understand. Two lightweight guardrails run unconditionally regardless of permission mode (including `auto-accept`), as a safety floor rather than a full sandbox: `write_file`/`edit_file`/`bash`'s `cwd` override can't target a path outside the working directory (`../` traversal, an absolute path elsewhere, or — on Windows — a different drive all refuse), and `bash` refuses a short list of unambiguously catastrophic commands (wiping the filesystem root or home directory, a fork bomb, formatting/wiping a whole drive, writing raw data to a block device) before they'd ever run. Neither guard stops a model from doing damage confined to *within* the project directory, or running something merely inadvisable — see `src/tools/pathGuard.ts` and `src/tools/bashGuard.ts`.
|
||||
- In-app scrollback: mouse wheel scrolls the conversation view when mouse tracking is on (the default). PageUp/PageDown also scroll a page at a time. Hold Shift+click/drag for native terminal text selection and copy (when mouse tracking is on). Scrolling back up unpins the view from the latest message; scrolling back to the bottom (or sending a new message) re-pins it so new messages auto-scroll into view. Toggle mouse tracking with `/mouse on|off`.
|
||||
- No sandboxing beyond the confirmation prompts — mutating tools operate on the real filesystem/shell with the permissions of the user running `locode`. Only approve commands you understand.
|
||||
- Session resume replays prior user/assistant text so you can see it, but it doesn't re-display prior tool-call/tool-result lines from before the resume (the model still has that history — it's just not re-rendered).
|
||||
- No in-app scrollback — once a message scrolls off the top of the window it's gone until you resize the terminal taller (the conversation itself is still intact and sent to the model; this only affects what you can visually re-read).
|
||||
- Windows shell quoting for the `bash` tool has only had light testing; behavior may differ from Unix shells for complex quoting.
|
||||
- `git_commit` covers add/commit/create_branch/checkout/push/reset/stash/merge/rebase/delete_branch. Use `bash` for anything beyond that.
|
||||
- MCP tool results support text, image, audio, and resource content blocks. Images are returned in the same shape as `read_file` so vision-capable models can see them; audio and binary resources are summarized. Remote (HTTP) MCP servers support static headers (e.g. a bearer token) but not OAuth flows.
|
||||
- Compaction (`/compact` or automatic) replaces history with a model-generated prose summary — it costs one extra model call and loses tool-call/tool-result detail (the model's own account of what happened survives; the raw record doesn't). The auto-compact threshold defaults to 85% and is configurable via `locode config set autoCompactThreshold` or `LOCODE_AUTO_COMPACT_THRESHOLD`.
|
||||
- Plugin support (`locode plugin add`) now covers every part of a plugin: MCP servers, slash commands, agents, hooks, and skills. A slash command's `allowed-tools` frontmatter restricts that one invocation's toolset (same tool-name translation as an agent's `tools:` — see `/plugins`); a skill invoked directly via `/<skill-name>` isn't restricted this way, since skills have no `allowed-tools` field of their own. Duplicate MCP server names across sources are now detected and surfaced in `/mcp` — project-level wins over user-level wins over plugin-level. Duplicate slash-command and skill names are also surfaced in `/plugins` and `/skills`.
|
||||
- Plugin support (`locode plugin add`) now covers every part of a plugin: MCP servers, slash commands, agents, hooks, and skills. A plugin's `allowed-tools` restriction on a command isn't enforced (the expanded prompt just runs as a normal turn with the full toolset). Duplicate MCP server names across sources are now detected and surfaced in `/mcp` — project-level wins over user-level wins over plugin-level. Duplicate slash-command and skill names are also surfaced in `/plugins` and `/skills`.
|
||||
- Skills are exposed as one shared `skill` tool rather than one tool per skill — if two plugins install a skill with the same name, the first plugin in load order wins and the collision is shown in `/skills`. Skills can bundle sibling `references/*.md` files that are included when the skill is invoked.
|
||||
- Hooks cover the most useful subset of Claude Code's lifecycle events: `SessionStart`, `UserPromptSubmit`, `PreToolUse`, `PostToolUse`, `PermissionRequest`, `SubagentStart`, `SubagentStop`, `CwdChanged`, `FileChanged`, `ConfigChange`, `Stop`, `SessionEnd`. `CwdChanged` is declared in the config format but not yet fired anywhere — a cwd never changes mid-session in locode today, so configuring it is a no-op for now. All three hook types are supported: `command` and `http` run an external process/request, and `prompt` just injects its static `message` as additional context (the same way a command/http hook's stdout does) — it has no process to fail, so it can't block an event the way a command hook's exit code 2 can. Command/http hooks can opt into structured JSON output via `outputSchema: "json"`. `PreToolUse`, `UserPromptSubmit`, and `PermissionRequest` can block; the rest are fire-and-forget. `SessionEnd` fires on every exit path (including Ctrl+C and external `SIGTERM`/`SIGHUP`), so it doesn't always have a real session id to report. All hooks for an event run in parallel with no defined ordering, and every configured hook always runs — there's no way to disable one without editing the file it came from.
|
||||
- Hooks cover the most useful subset of Claude Code's lifecycle events: `SessionStart`, `UserPromptSubmit`, `PreToolUse`, `PostToolUse`, `PermissionRequest`, `SubagentStart`, `SubagentStop`, `CwdChanged`, `FileChanged`, `ConfigChange`, `Stop`, `SessionEnd`. `CwdChanged` is declared in the config format but not yet fired anywhere — a cwd never changes mid-session in locode today, so configuring it is a no-op for now. `command` hooks and `http` hooks are supported; `prompt` hooks are declared in the config format but not yet executed. Command hooks can opt into structured JSON output via `outputSchema: "json"`. `PreToolUse`, `UserPromptSubmit`, and `PermissionRequest` can block; the rest are fire-and-forget. `SessionEnd` fires on every exit path (including Ctrl+C and external `SIGTERM`/`SIGHUP`), so it doesn't always have a real session id to report. All hooks for an event run in parallel with no defined ordering, and every configured hook always runs — there's no way to disable one without editing the file it came from.
|
||||
|
||||
@@ -1,91 +0,0 @@
|
||||
# locode 설계 메모
|
||||
|
||||
## 프로젝트 개요
|
||||
**locode** — Claude Code의 설계 철학을 가져와 구현한 에이전트 코딩 CLI. TypeScript + Ink(React-for-CLI) 기반. 백엔드는 OpenAI 호환 `/v1/chat/completions` 엔드포인트 사용 (Ollama, LM Studio, 클라우드 API 모두 지원).
|
||||
|
||||
- 핵심 철학: "신뢰할 수 없고 느리고 비전/툴콜 지원이 불확실한 모델"이라는 현실에 맞춰 모든 가정을 비관적으로 재단. 로컬/클라우드 자동 감지로 프롬프트 분기.
|
||||
- Claude Code 플러그인 포맷을 직접 소비하는 하위호환 브리지 (`.claude-plugin/plugin.json`, commands/agents/skills/hooks/MCP)
|
||||
|
||||
---
|
||||
|
||||
## 아키텍처 결정 이유
|
||||
|
||||
### Normal Screen + `<Static>` (alternate screen 폐지)
|
||||
이전에는 alternate screen buffer + 인앱 가상 스크롤 + 마우스 트래킹을 직접 구현했으나 완전히 폐지. 이유:
|
||||
- **코드 복잡도**: 마우스 SGR-1006 파싱, 선택 영역, 스크롤 상태 관리가 App.tsx의 절반을 차지
|
||||
- **터미널 호환성**: alternate screen은 SSH, tmux, Windows Terminal 등에서 파편화 심함
|
||||
- **버그 발생률**: 마우스 크래시, 선택 텍스트 깨짐 등 이슈가 끝없이 발생
|
||||
- **현재 방식**: Ink `<Static>`으로 완료된 히스토리를 한 번만 렌더링 → 터미널 스크롤백의 영구 부분. 재렌더링 없음. 마우스는 터미널 네이티브에 의존.
|
||||
|
||||
### 로컬/클라우드 모델 감지: `isSmallLocalModel()`
|
||||
- **이전**: `isLocalBackendURL(baseURL)` — URL이 localhost면 무조건 "로컬 모델"
|
||||
- **문제**: Ollama가 클라우드 라우팅 모델(`glm-5.2:cloud`, `qwen3.5:397b-cloud`)도 같은 localhost에서 서비스함. 이 모델들은 컨텍스트 128K~1M이고 툴콜도 안정적인데 로컬용 보수적 프롬프트가 적용됨.
|
||||
- **해결**: `isSmallLocalModel(baseURL, model)` = `isLocalBackendURL(baseURL) && !isCloudRoutedModelName(model)` — 모델명의 `:cloud`/`:-cloud` 태그로 구분.
|
||||
- **기본값**: `isLocal`의 기본값을 `true`에서 `false`(클라우드)로 변경. 명시적 지정이 없으면 보수적이 아닌 기본 프롬프트 사용.
|
||||
|
||||
### 컨텍스트 윈도우 기본값 분리
|
||||
- **이전**: `DEFAULT_CONTEXT_WINDOW = 8192` (단일, 로컬 모델 기준)
|
||||
- **현재**: `DEFAULT_CONTEXT_WINDOW_LOCAL = 8192`, `DEFAULT_CONTEXT_WINDOW_CLOUD = 131072` — 클라우드/Ollama 클라우드 라우팅 모델은 128K~1M 컨텍스트를 가지므로 8192는 과도하게 보수적.
|
||||
|
||||
### 번인레이트(🔥) 계산
|
||||
- **이전**: `outputTokens / 세션 경과 시간` — 사용자가 방치하면 번인레이트가 0에 수렴해서 의미 없음
|
||||
- **현재**: `outputTokens / modelTimeMs` — 실제 모델 응답 시간으로 계산. "이 모델이 얼마나 빠르게 토큰을 뿜는가"를 정확히 반영.
|
||||
|
||||
### 파일 인코딩: CRLF/LF 혼재
|
||||
- Windows 환경에서는 CRLF, Unix에서는 LF가 섞여 있음. `edit_file`/`multi_edit`은 매칭 전 LF로 정규화하고, 쓰기 전 원래 EOL을 복원. 이것 없이는 Windows에서 거의 모든 edit_file이 실패함.
|
||||
|
||||
---
|
||||
|
||||
## 핵심 파일 맵
|
||||
- `src/ui/ink/index.tsx` — 진입점. alternate screen 없이 Ink render. cleanup 시 flush + 종료.
|
||||
- `src/ui/ink/App.tsx` — 메인 UI 컴포넌트. `<Static>` + 라이브 영역. 상태: starting→connecting→loading-models→model-select/session-select→input.
|
||||
- `src/ui/ink/ChatInput.tsx` — 커스텀 multiline 입력. Shift+Enter 줄바꿈, bracket paste, @멘션 fuzzy picker, IME 커서.
|
||||
- `src/agent/loop.ts` — 메인 에이전트 루프. 턴/스트리밍/툴콜/컴팩션/서브에이전트/병렬 툴 배치/반복 루프 감지.
|
||||
- `src/agent/session.ts` — Session 객체, 통계, 상태, mutation gate.
|
||||
- `src/agent/systemPrompt.ts` — 시스템 프롬프트 빌더 (`isSmallLocalModel` 기반 로컬/클라우드 분기).
|
||||
- `src/config/defaults.ts` — 모든 기본값. `isSmallLocalModel()`, `isCloudRoutedModelName()`, `DEFAULT_CONTEXT_WINDOW_LOCAL/CLOUD` 등.
|
||||
- `src/toolcalling/` — native 어댑터, fallback 파서/프롬프트, partialJson 복구, resolve (Ollama 빈키 복구 포함).
|
||||
|
||||
---
|
||||
|
||||
## 설정 기본값
|
||||
|
||||
| 설정 | 기본값 | 비고 |
|
||||
|---|---|---|
|
||||
| `DEFAULT_CONTEXT_WINDOW_LOCAL` | 8192 | 작은 로컬 모델 폴백 |
|
||||
| `DEFAULT_CONTEXT_WINDOW_CLOUD` | 131072 | 클라우드/클라우드 라우팅 폴백 |
|
||||
| `DEFAULT_MAX_ITERATIONS` | **300** | 50→100→300 상향 |
|
||||
| `DEFAULT_MAX_OUTPUT_TOKENS` | **131072** | 128K. GLM 등 1M 컨텍스트 모델 대응 |
|
||||
| `DEFAULT_MAX_RETRIES` | 0 | SDK 지수 백오프 |
|
||||
| `DEFAULT_AUTO_COMPACT_THRESHOLD` | 0.85 | |
|
||||
| `DEFAULT_REQUEST_TIMEOUT_MS` | 180,000 | 3분 |
|
||||
| `DEFAULT_SUBAGENT_TIMEOUT_MS` | 600,000 | 10분 |
|
||||
| `MAX_EMPTY_RESPONSE_RETRIES` | **3** | 1→3 상향 |
|
||||
| `MAX_SUBAGENT_DEPTH` | 1 | 서브에이전트 중첩 금지 |
|
||||
| `MAX_PRESERVED_TAIL_MESSAGES` | 8 | 컴팩션 시 보존 |
|
||||
| `MAX_PRESERVED_TAIL_FRACTION` | 0.3 | 컴팩션 시 보존 비율 |
|
||||
| `MAX_RETAINED_IMAGES` | 2 | 히스토리 이미지 보존 |
|
||||
|
||||
---
|
||||
|
||||
## 트러블슈팅 힌트
|
||||
- **"자꾸 에러"**: 주요 원인은 max_tokens 잘림 → malformed 툴콜. 동적 max_tokens + CRLF 제어문자 이스케이프로 해결됨.
|
||||
- **CRLF edit_file 매칭 버그**: LF 정규화 공간에서 매칭, 쓰기 전 원래 EOL 복원으로 해결됨.
|
||||
- **"Paused after N steps"**: `locode config set maxIterations <number>` (기본 300)
|
||||
- **클라우드 모델 빈 응답**: `MAX_EMPTY_RESPONSE_RETRIES=3`으로 재시도
|
||||
- **로컬/클라우드 프롬프트 분기**: `isSmallLocalModel(baseURL, model)` — Ollama 클라우드 라우팅 모델(`:cloud` 태그)은 localhost여도 클라우드 프롬프트 사용
|
||||
- **반복 루프 감지**: `detectRepetitionLoop()` — 스트리밍 텍스트에서 짧은 반복 패턴 감지 시 중단
|
||||
- **번인레이트(🔥)**: `outputTokens ÷ modelTimeMs` 기준 (세션 경과 시간이 아닌 실제 모델 응답 시간)
|
||||
- **IPv6 localhost**: `isLocalBackendURL()`은 `[::1]` 형식(WHATWG URL 직렬화)도 인식
|
||||
|
||||
---
|
||||
|
||||
## 의존성 제약
|
||||
- **marked는 15에 고정**. `marked-terminal@7.3.0`의 peer가 `marked >=1 <16`이고 marked-terminal 업데이트가 없음. marked 16+로 올리려면 marked-terminal을 교체하거나 peer를 강제해야 함.
|
||||
- **typescript는 7.x (네이티브 컴파일러 포팅)**. 프로젝트는 `tsc` CLI로 타입체크만 하고 programmatic API를 안 씀 → 네이티브 포트로 안전하게 이전. 빌드는 tsup/esbuild라 tsc와 무관. 플랫폼별 `@typescript/typescript-*` 바이너리가 optional dep으로 붙음.
|
||||
- **wrap-ansi는 10.x (ink와 동일)**. `renderMarkdown`이 히스토리 출력을 터미널 폭으로 하드랩할 때 사용 — ink 내부 래핑과 같은 string-width v8을 공유해야 폭 계산이 어긋나지 않음. v10은 타입 내장(앰비언트 선언 불필요).
|
||||
- `allowScripts`에 트리에 실제로 존재하는 esbuild 버전을 모두 나열해야 함 (현재 `0.27.2` = tsup/vite, `0.28.1` = tsx). 빠지면 postinstall 경고.
|
||||
|
||||
## TODO
|
||||
- LSP 실서버 통합 테스트 (실제 tsserver/pyright 띄워서 검증)
|
||||
- task store 영속화 (세션에 task 저장)
|
||||
Generated
+632
-1393
File diff suppressed because it is too large
Load Diff
+18
-19
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "locode",
|
||||
"version": "0.5.2",
|
||||
"version": "0.6.0",
|
||||
"description": "Agentic coding CLI for local models served via Ollama and LM Studio",
|
||||
"type": "module",
|
||||
"bin": {
|
||||
@@ -10,7 +10,7 @@
|
||||
"dist"
|
||||
],
|
||||
"engines": {
|
||||
"node": ">=22.12.0"
|
||||
"node": ">=22"
|
||||
},
|
||||
"scripts": {
|
||||
"build": "tsup",
|
||||
@@ -20,39 +20,38 @@
|
||||
"prepublishOnly": "npm run build"
|
||||
},
|
||||
"dependencies": {
|
||||
"@modelcontextprotocol/sdk": "^1.30.0",
|
||||
"@anthropic-ai/claude-code": "^2.1.233",
|
||||
"@modelcontextprotocol/sdk": "^1.29.0",
|
||||
"@vscode/ripgrep": "^1.18.0",
|
||||
"commander": "^15.0.0",
|
||||
"commander": "^13.0.0",
|
||||
"diff": "^9.0.0",
|
||||
"env-paths": "^4.0.0",
|
||||
"execa": "^10.0.1",
|
||||
"execa": "^9.6.1",
|
||||
"fast-glob": "^3.3.3",
|
||||
"ink": "^7.1.1",
|
||||
"ink": "^7.1.0",
|
||||
"ink-select-input": "^6.2.0",
|
||||
"ink-spinner": "^5.0.0",
|
||||
"ink-text-input": "^6.0.0",
|
||||
"marked": "^15.0.12",
|
||||
"marked-terminal": "^7.3.0",
|
||||
"openai": "^7.13.0",
|
||||
"react": "^19.3.0",
|
||||
"string-width": "^8.2.2",
|
||||
"openai": "^6.45.0",
|
||||
"react": "^19.2.7",
|
||||
"string-width": "^8.2.1",
|
||||
"tree-kill": "^1.2.2",
|
||||
"vscode-languageserver-protocol": "^3.18.3",
|
||||
"vscode-uri": "^3.2.0",
|
||||
"wrap-ansi": "^10.0.1",
|
||||
"zod": "^4.6.1"
|
||||
"zod": "^4.4.3"
|
||||
},
|
||||
"devDependencies": {
|
||||
"@types/marked-terminal": "^6.1.1",
|
||||
"@types/node": "^26.5.1",
|
||||
"@types/react": "^19.3.0",
|
||||
"@types/node": "^22.10.0",
|
||||
"@types/react": "^19.2.17",
|
||||
"tsup": "^8.3.0",
|
||||
"tsx": "^4.23.13",
|
||||
"typescript": "^7.0.2",
|
||||
"vitest": "^5.0.0"
|
||||
"tsx": "^4.19.0",
|
||||
"typescript": "^5.7.0",
|
||||
"vitest": "^3.0.0"
|
||||
},
|
||||
"allowScripts": {
|
||||
"esbuild@0.27.2": true,
|
||||
"@anthropic-ai/claude-code@2.1.233": true,
|
||||
"esbuild@0.27.7": true,
|
||||
"esbuild@0.28.1": true
|
||||
}
|
||||
}
|
||||
|
||||
+15
-15
@@ -1,33 +1,33 @@
|
||||
import type { TodoItem } from "../tools/types.js";
|
||||
import type { TaskSummary } from "../tools/task.js";
|
||||
|
||||
export type AgentEvent =
|
||||
| { type: "text_delta"; delta: string }
|
||||
| { type: "text_done"; fullText: string }
|
||||
| { type: "thinking_delta"; delta: string }
|
||||
| { type: "thinking_done"; fullThinking: string }
|
||||
/** Throw away any text streamed so far this turn without committing it as an assistant message —
|
||||
* emitted when a partially-streamed native tool-call turn turns out to have malformed args and is
|
||||
* retried non-streaming, so the UI doesn't carry the stale partial into the retry's output. */
|
||||
| { type: "stream_discard" }
|
||||
/** `name`/`args` are only populated for an actually-resolved tool call (not the "unknown tool"
|
||||
* error path) — the file panel's Activity tab (App.tsx) uses them to track which files a
|
||||
* read_file/write_file/edit_file call touched, without having to re-parse the display `label`. */
|
||||
| { type: "tool_call"; label: string; name?: string; args?: unknown }
|
||||
/** `name`/`result` mirror `tool_call`'s — only populated when a tool actually ran (not a
|
||||
* hook-blocked/denied/unknown-tool result), for the same file-panel tracking purpose. */
|
||||
| { type: "tool_result"; summary: string; isError: boolean; name?: string; result?: unknown }
|
||||
| { type: "tool_call"; label: string }
|
||||
| { type: "tool_result"; summary: string; isError: boolean; diff?: string }
|
||||
/** The model finished a turn in plan mode with a prose plan. Rendered as a distinct `plan`
|
||||
* HistoryItem (not a plain assistant message) and followed by an approve/reject prompt built
|
||||
* on the same confirm primitive tool permissions use — see maybePresentPlan in agent/loop.ts. */
|
||||
| { type: "plan_presented"; text: string }
|
||||
/** The user approved the presented plan. loop.ts has already exited plan mode; the UI drives
|
||||
* implementation by injecting a "proceed" follow-up turn (see App.tsx submitTurn). */
|
||||
| { type: "plan_approved"; text: string }
|
||||
/** A sub-agent's tool call or result, forwarded to the parent so its work is visible while it
|
||||
* runs headless. Routed to the dedicated sub-agent panel below the input (not the main
|
||||
* scrollback) — see App.tsx. */
|
||||
* runs headless. Rendered inline in the main scrollback (tagged with the sub-agent's
|
||||
* description) — see App.tsx and HistoryItemView. */
|
||||
| { type: "subagent"; description: string; line: SubagentLine }
|
||||
/** A hook (see hooks/runner.ts) blocked something or failed non-fatally — surfaced as a notice. */
|
||||
| { type: "hook_notice"; text: string; isError: boolean }
|
||||
/** A general informational notice from the loop itself (not tied to a hook) — e.g. a mid-turn
|
||||
* auto-compaction. Surfaced the same way as hook_notice. */
|
||||
| { type: "notice"; text: string; isError: boolean }
|
||||
/** The `todo_write` tool replaced the session's task checklist — carries the full new list so
|
||||
* the UI can render it as a standalone checklist item rather than raw JSON tool output. */
|
||||
| { type: "todos_update"; todos: TodoItem[] };
|
||||
/** A task_create/task_update mutation changed the session's task store — carries a snapshot so
|
||||
* the UI can render the current checklist as a standalone item rather than raw JSON tool output. */
|
||||
| { type: "tasks_update"; tasks: TaskSummary[] };
|
||||
|
||||
export type SubagentLine =
|
||||
| { kind: "call"; label: string }
|
||||
|
||||
+1217
-467
File diff suppressed because it is too large
Load Diff
+674
-651
File diff suppressed because it is too large
Load Diff
@@ -1,229 +0,0 @@
|
||||
import { describe, expect, it, vi } from "vitest";
|
||||
import { z } from "zod";
|
||||
import { runTurn } from "./loop.js";
|
||||
import { agentTool } from "../tools/agentTool.js";
|
||||
import { createSession } from "./session.js";
|
||||
import { buildToolSet } from "../tools/toolset.js";
|
||||
import type { ToolDef } from "../tools/types.js";
|
||||
|
||||
/** A one-shot streaming response: yields `chunk` once, then ends. */
|
||||
function oneShotStream(chunk: any) {
|
||||
let yielded = false;
|
||||
return {
|
||||
[Symbol.asyncIterator]: () => ({
|
||||
next: async () => {
|
||||
if (yielded) return { done: true, value: undefined };
|
||||
yielded = true;
|
||||
return { done: false, value: chunk };
|
||||
},
|
||||
}),
|
||||
};
|
||||
}
|
||||
|
||||
function textChunk(text: string) {
|
||||
return { choices: [{ delta: { content: text }, finish_reason: "stop" }] };
|
||||
}
|
||||
|
||||
function toolCallChunk(name: string, args: string, id = "call_0") {
|
||||
return {
|
||||
choices: [
|
||||
{
|
||||
delta: { tool_calls: [{ index: 0, id, function: { name, arguments: args } }] },
|
||||
finish_reason: "tool_calls",
|
||||
},
|
||||
],
|
||||
};
|
||||
}
|
||||
|
||||
describe("agent tool / parallel sub-agents", () => {
|
||||
it("runs a `tasks` batch and returns each result independently", async () => {
|
||||
const agentArgs = JSON.stringify({
|
||||
description: "parallel research",
|
||||
tasks: [
|
||||
{ description: "task A", prompt: "do A" },
|
||||
{ description: "task B", prompt: "do B" },
|
||||
],
|
||||
});
|
||||
|
||||
// create() is called: 1 parent agent-call, then sub-agent A, then sub-agent B, then parent final.
|
||||
let call = 0;
|
||||
const fakeClient = {
|
||||
chat: {
|
||||
completions: {
|
||||
create: vi.fn(async () => {
|
||||
call++;
|
||||
if (call === 1) return oneShotStream(toolCallChunk("agent", agentArgs));
|
||||
if (call === 2) return oneShotStream(textChunk("result-A"));
|
||||
if (call === 3) return oneShotStream(textChunk("result-B"));
|
||||
return oneShotStream(textChunk("all done"));
|
||||
}),
|
||||
},
|
||||
},
|
||||
} as any;
|
||||
|
||||
const session = createSession(fakeClient, "test-model", process.cwd(), async () => "once", "native", [agentTool]);
|
||||
|
||||
const result = await runTurn(session, "research A and B in parallel", () => {});
|
||||
expect(result).toBe("all done");
|
||||
|
||||
// The agent tool's result must carry both sub-agent answers as a `results` array.
|
||||
const toolResultMsg = session.messages.find(
|
||||
(m) => m.role === "tool" && typeof (m as any).content === "string" && (m as any).content.includes("results"),
|
||||
) as any;
|
||||
expect(toolResultMsg).toBeTruthy();
|
||||
const parsed = JSON.parse(toolResultMsg.content);
|
||||
expect(parsed.results).toHaveLength(2);
|
||||
const byDesc = Object.fromEntries(parsed.results.map((r: any) => [r.description, r]));
|
||||
expect(byDesc["task A"].result).toBe("result-A");
|
||||
expect(byDesc["task B"].result).toBe("result-B");
|
||||
});
|
||||
|
||||
it("a failed sub-agent surfaces as its own error without discarding sibling results", async () => {
|
||||
const agentArgs = JSON.stringify({
|
||||
description: "mixed batch",
|
||||
tasks: [
|
||||
{ description: "ok", prompt: "succeed" },
|
||||
{ description: "boom", prompt: "fail" },
|
||||
],
|
||||
});
|
||||
|
||||
let call = 0;
|
||||
const fakeClient = {
|
||||
chat: {
|
||||
completions: {
|
||||
create: vi.fn(async () => {
|
||||
call++;
|
||||
if (call === 1) return oneShotStream(toolCallChunk("agent", agentArgs));
|
||||
if (call === 2) return oneShotStream(textChunk("ok-result"));
|
||||
if (call === 3) throw new Error("sub-agent boom failed");
|
||||
return oneShotStream(textChunk("done"));
|
||||
}),
|
||||
},
|
||||
},
|
||||
} as any;
|
||||
|
||||
const session = createSession(fakeClient, "test-model", process.cwd(), async () => "once", "native", [agentTool]);
|
||||
|
||||
await runTurn(session, "run mixed batch", () => {});
|
||||
|
||||
const toolResultMsg = session.messages.find(
|
||||
(m) => m.role === "tool" && typeof (m as any).content === "string" && (m as any).content.includes("results"),
|
||||
) as any;
|
||||
expect(toolResultMsg).toBeTruthy();
|
||||
const parsed = JSON.parse(toolResultMsg.content);
|
||||
const byDesc = Object.fromEntries(parsed.results.map((r: any) => [r.description, r]));
|
||||
expect(byDesc.ok.result).toBe("ok-result");
|
||||
expect(byDesc.boom.error).toBeTruthy();
|
||||
expect(byDesc.boom.error).toMatch(/boom/);
|
||||
});
|
||||
});
|
||||
|
||||
describe("mutation gate / serialization", () => {
|
||||
it("serializes mutating tool calls so two never run concurrently", async () => {
|
||||
let inFlight = 0;
|
||||
let maxOverlap = 0;
|
||||
const slowEdit: ToolDef = {
|
||||
name: "slow_edit",
|
||||
description: "slow edit",
|
||||
schema: z.object({}),
|
||||
mutating: true,
|
||||
handler: async () => {
|
||||
inFlight++;
|
||||
maxOverlap = Math.max(maxOverlap, inFlight);
|
||||
await new Promise((r) => setTimeout(r, 20));
|
||||
inFlight--;
|
||||
return { ok: true };
|
||||
},
|
||||
};
|
||||
|
||||
function twoEditsChunk() {
|
||||
return {
|
||||
choices: [
|
||||
{
|
||||
delta: {
|
||||
tool_calls: [
|
||||
{ index: 0, id: "c0", function: { name: "slow_edit", arguments: "{}" } },
|
||||
{ index: 1, id: "c1", function: { name: "slow_edit", arguments: "{}" } },
|
||||
],
|
||||
},
|
||||
finish_reason: "tool_calls",
|
||||
},
|
||||
],
|
||||
};
|
||||
}
|
||||
|
||||
let call = 0;
|
||||
const fakeClient = {
|
||||
chat: {
|
||||
completions: {
|
||||
create: vi.fn(async () => {
|
||||
call++;
|
||||
const chunk = call === 1 ? twoEditsChunk() : textChunk("done");
|
||||
return oneShotStream(chunk);
|
||||
}),
|
||||
},
|
||||
},
|
||||
} as any;
|
||||
|
||||
const session = createSession(fakeClient, "m", process.cwd(), async () => "once", "native", [slowEdit]);
|
||||
session.maxIterations = 5;
|
||||
|
||||
await runTurn(session, "two edits", () => {});
|
||||
expect(maxOverlap).toBe(1);
|
||||
});
|
||||
|
||||
it("lets read-only tools run concurrently (gate only blocks mutating)", async () => {
|
||||
let inFlight = 0;
|
||||
let maxOverlap = 0;
|
||||
const fastRead: ToolDef = {
|
||||
name: "fast_read",
|
||||
description: "fast read",
|
||||
schema: z.object({}),
|
||||
mutating: false,
|
||||
handler: async () => {
|
||||
inFlight++;
|
||||
maxOverlap = Math.max(maxOverlap, inFlight);
|
||||
await new Promise((r) => setTimeout(r, 20));
|
||||
inFlight--;
|
||||
return { ok: true };
|
||||
},
|
||||
};
|
||||
|
||||
function threeReadsChunk() {
|
||||
return {
|
||||
choices: [
|
||||
{
|
||||
delta: {
|
||||
tool_calls: [
|
||||
{ index: 0, id: "c0", function: { name: "fast_read", arguments: "{}" } },
|
||||
{ index: 1, id: "c1", function: { name: "fast_read", arguments: "{}" } },
|
||||
{ index: 2, id: "c2", function: { name: "fast_read", arguments: "{}" } },
|
||||
],
|
||||
},
|
||||
finish_reason: "tool_calls",
|
||||
},
|
||||
],
|
||||
};
|
||||
}
|
||||
|
||||
let call = 0;
|
||||
const fakeClient = {
|
||||
chat: {
|
||||
completions: {
|
||||
create: vi.fn(async () => {
|
||||
call++;
|
||||
const chunk = call === 1 ? threeReadsChunk() : textChunk("done");
|
||||
return oneShotStream(chunk);
|
||||
}),
|
||||
},
|
||||
},
|
||||
} as any;
|
||||
|
||||
const session = createSession(fakeClient, "m", process.cwd(), async () => "once", "native", [fastRead]);
|
||||
session.maxIterations = 5;
|
||||
|
||||
await runTurn(session, "three reads", () => {});
|
||||
// All three reads are read-only → runToolBatch runs them concurrently → they overlap.
|
||||
expect(maxOverlap).toBe(3);
|
||||
});
|
||||
});
|
||||
@@ -1,98 +0,0 @@
|
||||
import { describe, it, expect } from "vitest";
|
||||
import { undoLastTurn, resetSession } from "./session.js";
|
||||
import type { ChatCompletionMessageParam } from "openai/resources/chat/completions";
|
||||
|
||||
function makeSession(messages: ChatCompletionMessageParam[]) {
|
||||
return {
|
||||
messages,
|
||||
lastContextTokens: 0,
|
||||
lastContextTokensIsEstimate: true,
|
||||
} as any;
|
||||
}
|
||||
|
||||
describe("undoLastTurn", () => {
|
||||
it("returns 0 when there are no user messages", () => {
|
||||
const session = makeSession([
|
||||
{ role: "system", content: "You are helpful." },
|
||||
]);
|
||||
expect(undoLastTurn(session)).toBe(0);
|
||||
});
|
||||
|
||||
it("removes a single user turn at the end", () => {
|
||||
const session = makeSession([
|
||||
{ role: "system", content: "You are helpful." },
|
||||
{ role: "user", content: "Hello" },
|
||||
{ role: "assistant", content: "Hi there!" },
|
||||
]);
|
||||
const removed = undoLastTurn(session);
|
||||
expect(removed).toBe(2);
|
||||
expect(session.messages.length).toBe(1);
|
||||
expect(session.messages[0]!.role).toBe("system");
|
||||
});
|
||||
|
||||
it("removes user + assistant + tool results together", () => {
|
||||
const session = makeSession([
|
||||
{ role: "system", content: "You are helpful." },
|
||||
{ role: "user", content: "Read the file" },
|
||||
{ role: "assistant", content: "", tool_calls: [{ id: "tc1", type: "function", function: { name: "read_file", arguments: "{}" } }] } as any,
|
||||
{ role: "tool", content: "file contents here", tool_call_id: "tc1" } as any,
|
||||
{ role: "assistant", content: "The file contains..." },
|
||||
]);
|
||||
const removed = undoLastTurn(session);
|
||||
expect(removed).toBe(4);
|
||||
expect(session.messages.length).toBe(1);
|
||||
});
|
||||
|
||||
it("only removes the last turn, keeping earlier turns", () => {
|
||||
const session = makeSession([
|
||||
{ role: "system", content: "You are helpful." },
|
||||
{ role: "user", content: "First question" },
|
||||
{ role: "assistant", content: "First answer" },
|
||||
{ role: "user", content: "Second question" },
|
||||
{ role: "assistant", content: "Second answer" },
|
||||
]);
|
||||
const removed = undoLastTurn(session);
|
||||
expect(removed).toBe(2);
|
||||
expect(session.messages.length).toBe(3);
|
||||
expect((session.messages[2] as any).content).toBe("First answer");
|
||||
});
|
||||
|
||||
it("handles consecutive user messages (removing only the last one)", () => {
|
||||
const session = makeSession([
|
||||
{ role: "system", content: "You are helpful." },
|
||||
{ role: "user", content: "Message 1" },
|
||||
{ role: "user", content: "Message 2" },
|
||||
]);
|
||||
const removed = undoLastTurn(session);
|
||||
expect(removed).toBe(1);
|
||||
expect(session.messages.length).toBe(2);
|
||||
expect((session.messages[1] as any).content).toBe("Message 1");
|
||||
});
|
||||
|
||||
it("updates context token tracking after undo", () => {
|
||||
const session = makeSession([
|
||||
{ role: "system", content: "You are helpful." },
|
||||
{ role: "user", content: "Hello" },
|
||||
{ role: "assistant", content: "Hi!" },
|
||||
]);
|
||||
session.lastContextTokens = 5000;
|
||||
session.lastContextTokensIsEstimate = false;
|
||||
undoLastTurn(session);
|
||||
expect(session.lastContextTokensIsEstimate).toBe(true);
|
||||
// lastContextTokens should be recalculated (smaller than before)
|
||||
expect(session.lastContextTokens).toBeLessThan(5000);
|
||||
});
|
||||
});
|
||||
|
||||
describe("resetSession", () => {
|
||||
it("clears all messages except system prompt", () => {
|
||||
const session = makeSession([
|
||||
{ role: "system", content: "You are helpful." },
|
||||
{ role: "user", content: "Hello" },
|
||||
{ role: "assistant", content: "Hi!" },
|
||||
]);
|
||||
resetSession(session);
|
||||
expect(session.messages.length).toBe(1);
|
||||
expect(session.messages[0]!.role).toBe("system");
|
||||
});
|
||||
});
|
||||
+132
-83
@@ -2,15 +2,16 @@ import { randomUUID } from "node:crypto";
|
||||
import type OpenAI from "openai";
|
||||
import type { ChatCompletionMessageParam } from "openai/resources/chat/completions";
|
||||
import type { ToolCallMode } from "../backend/capabilityProbe.js";
|
||||
import { DEFAULT_AUTO_COMPACT_THRESHOLD, DEFAULT_CONTEXT_WINDOW_CLOUD, DEFAULT_CONTEXT_WINDOW_LOCAL, DEFAULT_MAX_ITERATIONS } from "../config/defaults.js";
|
||||
import { DEFAULT_AUTO_COMPACT_THRESHOLD, DEFAULT_CONTEXT_WINDOW, DEFAULT_MAX_ITERATIONS, DEFAULT_SUBAGENT_MAX_DEPTH, DEFAULT_SUBAGENT_MAX_ITERATIONS } from "../config/defaults.js";
|
||||
import type { SessionRecord } from "../persistence/sessionStore.js";
|
||||
import { PermissionManager } from "../permissions/permissionManager.js";
|
||||
import type { ConfirmFn } from "../permissions/types.js";
|
||||
import { TOOLS } from "../tools/index.js";
|
||||
import { buildToolSet, type ToolSet } from "../tools/toolset.js";
|
||||
import type { TodoItem, ToolDef } from "../tools/types.js";
|
||||
import type { AskQuestionSpec, AskQuestionAnswer, ToolDef } from "../tools/types.js";
|
||||
import { TaskStore } from "../tools/task.js";
|
||||
import type { TaskStoreSnapshot } from "../tools/task.js";
|
||||
import type { CronStore } from "../scheduler/cron.js";
|
||||
import type { PermissionMode, PermissionRule } from "../permissions/types.js";
|
||||
import { estimateTokens } from "../utils/tokens.js";
|
||||
import { buildSystemPrompt } from "./systemPrompt.js";
|
||||
|
||||
@@ -47,12 +48,34 @@ export interface Session {
|
||||
mode: ToolCallMode;
|
||||
messages: ChatCompletionMessageParam[];
|
||||
maxIterations: number;
|
||||
/** Max tool calls in a single sub-agent turn launched from this session. Intentionally smaller
|
||||
* than maxIterations (sub-agents run one focused task, one round, no auto-continue) so a runaway
|
||||
* sub-agent fails fast and surfaces a "split the task" hint instead of burning a large budget. */
|
||||
subagentMaxIterations: number;
|
||||
/** Max nesting depth for sub-agents launched from this session. The main session is depth 0;
|
||||
* a sub-agent it spawns is depth 1, and so on. A sub-agent at the cap has `agent` excluded from
|
||||
* its toolset (with an explicit depth-check backstop in agent/loop.ts) so it can't delegate
|
||||
* further. Configurable so it can be lowered in tests without touching env/config. */
|
||||
subagentMaxDepth: number;
|
||||
permissions: PermissionManager;
|
||||
confirm: ConfirmFn;
|
||||
/** Optional structured-question callback wired by the UI so the `ask_user_question` tool can
|
||||
* prompt the user with multiple-choice options. Absent in headless/non-UI contexts (sub-agents),
|
||||
* in which case the tool returns a clear "can't ask" error instead of hanging. */
|
||||
askQuestion?: (questions: AskQuestionSpec[]) => Promise<AskQuestionAnswer[]>;
|
||||
/** Promise-chain lock serializing calls to `confirm` across concurrent sub-agents. When a parent
|
||||
* turn fans out multiple `agent` delegations in parallel (see runBatch in agent/loop.ts), each
|
||||
* sub-agent shares this same mutex *holder* (subSession.confirmMutex = parent.confirmMutex, by
|
||||
* reference) so their mutating-tool confirmation prompts queue one at a time instead of racing
|
||||
* for the UI's single PendingPermission slot. The holder wraps a `chain` promise that
|
||||
* withConfirmLock reassigns on each confirm; sharing the holder (rather than the promise itself)
|
||||
* keeps every concurrent sub-agent queued on the same lock even as the chain advances. */
|
||||
confirmMutex: { chain: Promise<void> };
|
||||
/** Local tools plus any dynamically-discovered ones (currently: MCP) available for this session. */
|
||||
toolset: ToolSet;
|
||||
/** 0 for a normal session; incremented for each level of sub-agent nesting (capped at
|
||||
* MAX_SUBAGENT_DEPTH in agent/loop.ts, independent of the toolset already excluding `agent`). */
|
||||
* session.subagentMaxDepth in agent/loop.ts, independent of the toolset already excluding `agent`
|
||||
* for sub-agents at the depth cap). */
|
||||
subAgentDepth: number;
|
||||
/** The model's context window in tokens — auto-detected where possible (backend/contextWindow.ts),
|
||||
* otherwise a configured/hardcoded fallback (see contextWindowIsEstimate). */
|
||||
@@ -85,29 +108,48 @@ export interface Session {
|
||||
* project's CLAUDE.md/AGENTS.md conventions survive everything that regenerates messages[0].
|
||||
* Null when neither file exists. */
|
||||
projectInstructions: string | null;
|
||||
/** Current task checklist shown to the user via the `todo_write` tool — session-scoped state
|
||||
* since checklist items are a snapshot of progress, not part of the model-visible conversation. */
|
||||
todos: TodoItem[];
|
||||
/** Structured task store backing the task_create/list/get/update tools — session-scoped,
|
||||
* in-memory (not persisted), and independent of the flat `todos` checklist. */
|
||||
/** The user's personal memory file (config dir / memory.md), folded into the system prompt
|
||||
* alongside projectInstructions so learned preferences/feedback survive across sessions and repos.
|
||||
* Null when the file is absent or empty. See utils/userMemory.ts. */
|
||||
userMemory: string | null;
|
||||
/** The session's structured task store (dependency graph + ownership), surfaced to the model via
|
||||
* the task_create/list/get/update tools and to the UI via `tasks_update` events. Session-scoped —
|
||||
* tasks are a progress snapshot, not part of the model-visible conversation, so not persisted. */
|
||||
taskStore: TaskStore;
|
||||
/** A serialized-mutation gate shared by EVERY tool call in this session — including sub-agents
|
||||
* spawned in parallel via the `agent` tool's `tasks` array. Parallel sub-agents share their
|
||||
* parent's session, so without a lock two of them could simultaneously call a mutating tool,
|
||||
* race on the single React `permission` slot (makeConfirmFn), and interleave filesystem writes.
|
||||
* This gate lets read-only tools run concurrently (matching runToolBatch) while serializing
|
||||
* mutating ones: each mutating call awaits the previous one before it even prompts, so
|
||||
* permission prompts stay one-at-a-time and edits can't overlap. The chain is per-session,
|
||||
* so a top-level turn and its sub-agents all funnel through the same queue. Initialized as a
|
||||
* resolved promise so the first caller doesn't wait on anything. */
|
||||
mutationGate: Promise<void>;
|
||||
/** True when this is a small/less-capable local model, not just a local *backend* — Ollama's
|
||||
* cloud-routed models (e.g. "glm-5.2:cloud") share a localhost endpoint with genuinely local
|
||||
* ones, so this is more than a baseURL check. Used to tailor the system prompt — small local
|
||||
* models need extra guidance about their limitations; everything else gets a leaner prompt
|
||||
* without self-fulfilling "you may produce empty responses" framing. Derived by
|
||||
* isSmallLocalModel(baseURL, model); recomputed on /model and /backend switches (see App.tsx). */
|
||||
isLocal: boolean;
|
||||
/** The session's cron/wakeup scheduler, set by the App after creating the session (the App owns its
|
||||
* lifecycle: starts the tick with enqueue/isIdle callbacks, stops it on unmount). Absent in
|
||||
* non-UI contexts. Used by the cron_create/list/delete and schedule_wakeup tools. */
|
||||
cronStore?: CronStore;
|
||||
/** Tracks the most recent file mutation (write_file or edit_file) so the user can roll it back
|
||||
* with the /undo slash command. The path is stored as the user-supplied path rather than a
|
||||
* resolved absolute path, so the undo re-uses the same relative path logic as the original edit. */
|
||||
lastEdit: { path: string; previousContent: string } | null;
|
||||
/** Resumable sub-agent sessions keyed by their agentId, so the parent can continue one with a
|
||||
* follow-up message via the `send_message` tool (ctx.resumeSubAgent). Only sub-agents that ran in
|
||||
* the shared cwd are stored here — worktree-isolated parallel agents' cwd is cleaned up after the
|
||||
* batch, so they're fire-and-forget and never added. Cleared on resetSession. */
|
||||
subAgentSessions: Map<string, Session>;
|
||||
/** Named teammates: maps a teammate name (supplied via the `agent` tool's `name` arg) to its
|
||||
* agentId, so `send_message` can address it by name and `list_teammates` can roster it. Backed by
|
||||
* the same resumable sub-agent sessions as subAgentSessions — this is just a name → agentId index
|
||||
* over them. Independent per session (a sub-agent gets its own map for nested teammates). Cleared
|
||||
* on resetSession alongside subAgentSessions. */
|
||||
namedAgents: Map<string, string>;
|
||||
/** The active interactive worktree session, set by `enter_worktree` and cleared by `exit_worktree`.
|
||||
* While set, `cwd` points at the worktree dir (an isolated checkout on its own branch) and file
|
||||
* tools operate there; `originalCwd` is restored on exit. Undefined when not in a worktree session.
|
||||
* Not persisted — worktree sessions don't survive an app restart (the branch stays in git, so the
|
||||
* work itself isn't lost; re-enter via git if needed). */
|
||||
worktree?: { dir: string; branch: string; originalCwd: string };
|
||||
/** App-provided callback invoked when the session's cwd changes (currently only via
|
||||
* `enter_worktree`/`exit_worktree`), so the UI can update its live cwd display, /undo resolution,
|
||||
* git-info refresh, and @mention resolution. Absent in non-UI contexts (a headless sub-agent
|
||||
* switching its own cwd just has no UI to notify). */
|
||||
onCwdChange?: (newCwd: string) => void;
|
||||
/** App-provided callback invoked when the worktree session is entered/exited, so the App can keep a
|
||||
* ref that survives a /model switch (which recreates the session) and re-attach the worktree
|
||||
* tracking to the new session. Absent in non-UI contexts. */
|
||||
onWorktreeChange?: (worktree: { dir: string; branch: string; originalCwd: string } | null) => void;
|
||||
}
|
||||
|
||||
export function createSession(
|
||||
@@ -117,18 +159,19 @@ export function createSession(
|
||||
confirm: ConfirmFn,
|
||||
mode: ToolCallMode,
|
||||
tools: ToolDef[] = TOOLS,
|
||||
contextWindow?: number,
|
||||
contextWindow: number = DEFAULT_CONTEXT_WINDOW,
|
||||
contextWindowIsEstimate: boolean = true,
|
||||
maxIterations: number = DEFAULT_MAX_ITERATIONS,
|
||||
autoCompactThreshold: number = DEFAULT_AUTO_COMPACT_THRESHOLD,
|
||||
projectInstructions: string | null = null,
|
||||
isLocal?: boolean,
|
||||
subagentMaxIterations: number = DEFAULT_SUBAGENT_MAX_ITERATIONS,
|
||||
userMemory: string | null = null,
|
||||
subagentMaxDepth: number = DEFAULT_SUBAGENT_MAX_DEPTH,
|
||||
permissionRules: PermissionRule[] = [],
|
||||
): Session {
|
||||
const resolvedIsLocal = isLocal ?? false;
|
||||
const resolvedContextWindow = contextWindow ?? (resolvedIsLocal ? DEFAULT_CONTEXT_WINDOW_LOCAL : DEFAULT_CONTEXT_WINDOW_CLOUD);
|
||||
const toolset = buildToolSet(tools);
|
||||
const messages: ChatCompletionMessageParam[] = [
|
||||
{ role: "system", content: buildSystemPrompt(toolset.tools, mode, projectInstructions, resolvedIsLocal) },
|
||||
{ role: "system", content: buildSystemPrompt(toolset.tools, mode, projectInstructions, userMemory) },
|
||||
];
|
||||
return {
|
||||
id: randomUUID(),
|
||||
@@ -137,14 +180,16 @@ export function createSession(
|
||||
model,
|
||||
cwd,
|
||||
mode,
|
||||
isLocal: resolvedIsLocal,
|
||||
messages,
|
||||
maxIterations,
|
||||
permissions: new PermissionManager(),
|
||||
subagentMaxIterations,
|
||||
subagentMaxDepth,
|
||||
permissions: new PermissionManager(permissionRules),
|
||||
confirm,
|
||||
confirmMutex: { chain: Promise.resolve() },
|
||||
toolset,
|
||||
subAgentDepth: 0,
|
||||
contextWindow: resolvedContextWindow,
|
||||
contextWindow,
|
||||
contextWindowIsEstimate,
|
||||
lastContextTokens: estimateTokens(messages),
|
||||
lastContextTokensIsEstimate: true,
|
||||
@@ -153,9 +198,11 @@ export function createSession(
|
||||
activeBackground: null,
|
||||
mutationCommitLength: null,
|
||||
projectInstructions,
|
||||
todos: [],
|
||||
userMemory,
|
||||
taskStore: new TaskStore(),
|
||||
mutationGate: Promise.resolve(),
|
||||
lastEdit: null,
|
||||
subAgentSessions: new Map(),
|
||||
namedAgents: new Map(),
|
||||
};
|
||||
}
|
||||
|
||||
@@ -167,35 +214,45 @@ export function createSessionFromRecord(
|
||||
cwd: string,
|
||||
confirm: ConfirmFn,
|
||||
tools: ToolDef[] = TOOLS,
|
||||
contextWindow?: number,
|
||||
contextWindow: number = DEFAULT_CONTEXT_WINDOW,
|
||||
contextWindowIsEstimate: boolean = true,
|
||||
maxIterations: number = DEFAULT_MAX_ITERATIONS,
|
||||
autoCompactThreshold: number = DEFAULT_AUTO_COMPACT_THRESHOLD,
|
||||
projectInstructions: string | null = null,
|
||||
isLocal?: boolean,
|
||||
subagentMaxIterations: number = DEFAULT_SUBAGENT_MAX_ITERATIONS,
|
||||
userMemory: string | null = null,
|
||||
subagentMaxDepth: number = DEFAULT_SUBAGENT_MAX_DEPTH,
|
||||
permissionRules: PermissionRule[] = [],
|
||||
): Session {
|
||||
const resolvedIsLocal = isLocal ?? false;
|
||||
const resolvedContextWindow = contextWindow ?? (resolvedIsLocal ? DEFAULT_CONTEXT_WINDOW_LOCAL : DEFAULT_CONTEXT_WINDOW_CLOUD);
|
||||
const toolset = buildToolSet(tools);
|
||||
// Build the system prompt with the *restored* permission mode (not the default) so a resumed
|
||||
// plan-mode session gets plan instructions in its prompt from the first request, rather than
|
||||
// only learning it's in plan mode from tool-rejection errors.
|
||||
const messages: ChatCompletionMessageParam[] = [
|
||||
{ role: "system", content: buildSystemPrompt(toolset.tools, record.mode, projectInstructions, resolvedIsLocal) },
|
||||
{ role: "system", content: buildSystemPrompt(toolset.tools, record.mode, projectInstructions, userMemory, record.permissionMode) },
|
||||
...record.messages,
|
||||
];
|
||||
const session: Session = {
|
||||
const permissions = new PermissionManager(permissionRules);
|
||||
// Restore the saved permission mode so plan/auto-edit/auto-accept survive a resume instead of
|
||||
// always resetting to default. Older saved sessions omit the field → default.
|
||||
if (record.permissionMode) permissions.setMode(record.permissionMode);
|
||||
return {
|
||||
id: record.id,
|
||||
createdAt: record.createdAt,
|
||||
client,
|
||||
model: record.model,
|
||||
cwd,
|
||||
mode: record.mode,
|
||||
isLocal: resolvedIsLocal,
|
||||
messages,
|
||||
maxIterations,
|
||||
permissions: new PermissionManager(),
|
||||
subagentMaxIterations,
|
||||
subagentMaxDepth,
|
||||
permissions,
|
||||
confirm,
|
||||
confirmMutex: { chain: Promise.resolve() },
|
||||
toolset,
|
||||
subAgentDepth: 0,
|
||||
contextWindow: resolvedContextWindow,
|
||||
contextWindow,
|
||||
contextWindowIsEstimate,
|
||||
lastContextTokens: estimateTokens(messages),
|
||||
lastContextTokensIsEstimate: true,
|
||||
@@ -204,19 +261,12 @@ export function createSessionFromRecord(
|
||||
activeBackground: null,
|
||||
mutationCommitLength: null,
|
||||
projectInstructions,
|
||||
todos: [],
|
||||
taskStore: TaskStore.fromJSON(record.tasks ?? { seq: 0, tasks: [] }),
|
||||
mutationGate: Promise.resolve(),
|
||||
userMemory,
|
||||
taskStore: new TaskStore(),
|
||||
lastEdit: null,
|
||||
subAgentSessions: new Map(),
|
||||
namedAgents: new Map(),
|
||||
};
|
||||
|
||||
// Restore session-allowed tools from the saved record, so /perm approvals survive resume.
|
||||
if (record.allowedTools) {
|
||||
for (const toolName of record.allowedTools) {
|
||||
session.permissions.allowForSession(toolName);
|
||||
}
|
||||
}
|
||||
|
||||
return session;
|
||||
}
|
||||
|
||||
export function toSessionRecord(session: Session, baseURL: string): SessionRecord {
|
||||
@@ -228,9 +278,8 @@ export function toSessionRecord(session: Session, baseURL: string): SessionRecor
|
||||
baseURL,
|
||||
model: session.model,
|
||||
mode: session.mode,
|
||||
permissionMode: session.permissions.getMode(),
|
||||
messages: session.messages.slice(1),
|
||||
allowedTools: session.permissions.listAllowed(),
|
||||
tasks: session.taskStore.toJSON(),
|
||||
};
|
||||
}
|
||||
|
||||
@@ -238,32 +287,32 @@ export function resetSession(session: Session): void {
|
||||
session.messages = [session.messages[0] as ChatCompletionMessageParam];
|
||||
session.lastContextTokens = estimateTokens(session.messages);
|
||||
session.lastContextTokensIsEstimate = true;
|
||||
}
|
||||
|
||||
/** Removes the last complete user turn (the user message + all subsequent assistant/tool messages
|
||||
* up to the next user message or the end of history). Returns the number of messages removed,
|
||||
* or 0 if there's no user message to undo (only the system prompt remains). This is a soft undo \u2014
|
||||
* filesystem changes from tool calls are NOT rolled back, but the model will no longer see the
|
||||
* removed context, so it won't repeat those actions. */
|
||||
export function undoLastTurn(session: Session): number {
|
||||
// Walk backwards from the end to find the last user message.
|
||||
let lastUserIdx = -1;
|
||||
for (let i = session.messages.length - 1; i >= 1; i--) {
|
||||
if (session.messages[i]!.role === "user") {
|
||||
lastUserIdx = i;
|
||||
break;
|
||||
}
|
||||
}
|
||||
if (lastUserIdx === -1) return 0; // No user messages to undo.
|
||||
|
||||
const removed = session.messages.length - lastUserIdx;
|
||||
session.messages.length = lastUserIdx;
|
||||
session.lastContextTokens = estimateTokens(session.messages);
|
||||
session.lastContextTokensIsEstimate = true;
|
||||
return removed;
|
||||
// Drop any resumable sub-agent sessions — their context references the old conversation and would
|
||||
// be stale after a /clear.
|
||||
session.subAgentSessions.clear();
|
||||
// Drop the name → agentId index over those sessions too.
|
||||
session.namedAgents.clear();
|
||||
}
|
||||
|
||||
export function setMode(session: Session, mode: ToolCallMode): void {
|
||||
session.mode = mode;
|
||||
session.messages[0] = { role: "system", content: buildSystemPrompt(session.toolset.tools, mode, session.projectInstructions, session.isLocal) };
|
||||
// Re-apply the current permission mode so plan-mode instructions survive a tool-call-mode switch
|
||||
// (native ↔ fallback) rather than being dropped from the rebuilt prompt.
|
||||
session.messages[0] = {
|
||||
role: "system",
|
||||
content: buildSystemPrompt(session.toolset.tools, mode, session.projectInstructions, session.userMemory, session.permissions.getMode()),
|
||||
};
|
||||
}
|
||||
|
||||
/** Switch the session's permission mode and rebuild the system prompt so the model is told about
|
||||
* the new mode (e.g. entering plan mode injects the plan-only research instructions). This is the
|
||||
* permission-mode counterpart to setMode (which handles tool-call mode). UI paths that change the
|
||||
* permission mode (/perm, Shift+Tab) should call this instead of `session.permissions.setMode` alone,
|
||||
* which would leave the prompt stale. */
|
||||
export function setPermissionMode(session: Session, mode: PermissionMode): void {
|
||||
session.permissions.setMode(mode);
|
||||
session.messages[0] = {
|
||||
role: "system",
|
||||
content: buildSystemPrompt(session.toolset.tools, session.mode, session.projectInstructions, session.userMemory, mode),
|
||||
};
|
||||
}
|
||||
|
||||
+54
-120
@@ -1,131 +1,65 @@
|
||||
import { describe, it, expect } from "vitest";
|
||||
import { describe, expect, it } from "vitest";
|
||||
import { z } from "zod";
|
||||
import { buildSystemPrompt, } from "./systemPrompt.js";
|
||||
import { createSession, setMode, setPermissionMode } from "./session.js";
|
||||
import { buildToolSet } from "../tools/toolset.js";
|
||||
import { FALLBACK_TOOL_INSTRUCTIONS } from "../toolcalling/fallbackPrompt.js";
|
||||
import type { ToolDef } from "../tools/types.js";
|
||||
import { buildSystemPrompt } from "./systemPrompt.js";
|
||||
|
||||
const dummyTool = (name: string, mutating: boolean): ToolDef => ({
|
||||
name,
|
||||
description: `Tool ${name} for testing`,
|
||||
const dummyTool: ToolDef = {
|
||||
name: "noop",
|
||||
description: "does nothing",
|
||||
schema: z.object({}),
|
||||
mutating,
|
||||
handler: async () => null,
|
||||
mutating: false,
|
||||
handler: async () => ({ ok: true }),
|
||||
};
|
||||
const toolset = buildToolSet([dummyTool]);
|
||||
|
||||
function systemPromptOf(session: { messages: { content?: unknown }[] }): string {
|
||||
return String(session.messages[0]!.content);
|
||||
}
|
||||
|
||||
describe("buildSystemPrompt plan-mode injection", () => {
|
||||
it("omits plan instructions in the default mode", () => {
|
||||
const prompt = buildSystemPrompt(toolset.tools, "native");
|
||||
expect(prompt).not.toContain("Plan mode is ACTIVE");
|
||||
});
|
||||
|
||||
it("injects plan instructions when permissionMode is 'plan'", () => {
|
||||
const prompt = buildSystemPrompt(toolset.tools, "native", null, null, "plan");
|
||||
expect(prompt).toContain("Plan mode is ACTIVE");
|
||||
expect(prompt).toContain("present a concrete implementation plan");
|
||||
});
|
||||
|
||||
it("does not inject plan instructions for auto-edit/auto-accept", () => {
|
||||
expect(buildSystemPrompt(toolset.tools, "native", null, null, "auto-edit")).not.toContain("Plan mode is ACTIVE");
|
||||
expect(buildSystemPrompt(toolset.tools, "native", null, null, "auto-accept")).not.toContain("Plan mode is ACTIVE");
|
||||
});
|
||||
});
|
||||
|
||||
describe("buildSystemPrompt", () => {
|
||||
const readTool = dummyTool("read_file", false);
|
||||
const writeTool = dummyTool("write_file", true);
|
||||
describe("setPermissionMode / setMode prompt rebuild", () => {
|
||||
// A fake OpenAI client — these tests never make requests, createSession just needs a client.
|
||||
const fakeClient = {} as any;
|
||||
|
||||
it("includes tool names grouped by category", () => {
|
||||
const prompt = buildSystemPrompt([readTool, writeTool], "native");
|
||||
expect(prompt).toContain("read_file");
|
||||
expect(prompt).toContain("write_file");
|
||||
expect(prompt).toContain("Read-only");
|
||||
expect(prompt).toContain("Mutating");
|
||||
it("setPermissionMode('plan') rebuilds the system prompt with plan instructions, and back to default removes them", () => {
|
||||
const session = createSession(fakeClient, "m", process.cwd(), async () => "once", "native", [dummyTool]);
|
||||
expect(systemPromptOf(session)).not.toContain("Plan mode is ACTIVE");
|
||||
|
||||
setPermissionMode(session, "plan");
|
||||
expect(systemPromptOf(session)).toContain("Plan mode is ACTIVE");
|
||||
|
||||
setPermissionMode(session, "default");
|
||||
expect(systemPromptOf(session)).not.toContain("Plan mode is ACTIVE");
|
||||
});
|
||||
|
||||
it("includes core principles for local models", () => {
|
||||
const prompt = buildSystemPrompt([readTool], "native", null, true);
|
||||
expect(prompt).toContain("Inspect before answering");
|
||||
expect(prompt).toContain("Prefer small, targeted edits");
|
||||
expect(prompt).toContain("Recovery over retry");
|
||||
expect(prompt).toContain("Respect confirmation");
|
||||
});
|
||||
|
||||
it("includes core principles for cloud models", () => {
|
||||
const prompt = buildSystemPrompt([readTool], "native", null, false);
|
||||
expect(prompt).toContain("Inspect before answering");
|
||||
expect(prompt).toContain("Prefer small, targeted edits");
|
||||
expect(prompt).toContain("Recovery over retry");
|
||||
expect(prompt).toContain("Respect confirmation");
|
||||
});
|
||||
|
||||
it("includes tool usage guide", () => {
|
||||
const prompt = buildSystemPrompt([readTool], "native");
|
||||
expect(prompt).toContain("read_file");
|
||||
expect(prompt).toContain("edit_file");
|
||||
expect(prompt).toContain("bash");
|
||||
expect(prompt).toContain("agent");
|
||||
});
|
||||
|
||||
it("includes local model guidance when isLocal is true", () => {
|
||||
const prompt = buildSystemPrompt([readTool], "native", null, true);
|
||||
expect(prompt).toContain("Working with local models");
|
||||
expect(prompt).toContain("Tool-call formatting can be unreliable");
|
||||
expect(prompt).toContain("Empty or malformed responses can happen");
|
||||
});
|
||||
|
||||
it("excludes local model guidance when isLocal is false (cloud)", () => {
|
||||
const prompt = buildSystemPrompt([readTool], "native", null, false);
|
||||
expect(prompt).not.toContain("Working with local models");
|
||||
expect(prompt).not.toContain("Tool-call formatting can be unreliable");
|
||||
expect(prompt).not.toContain("Empty or malformed responses can happen");
|
||||
});
|
||||
|
||||
it("defaults to cloud prompt when isLocal is not specified", () => {
|
||||
const prompt = buildSystemPrompt([readTool], "native");
|
||||
expect(prompt).toContain("coding assistant");
|
||||
expect(prompt).not.toContain("Working with local models");
|
||||
});
|
||||
|
||||
it("includes safety guidelines for both local and cloud", () => {
|
||||
const localPrompt = buildSystemPrompt([readTool], "native", null, true);
|
||||
const cloudPrompt = buildSystemPrompt([readTool], "native", null, false);
|
||||
expect(localPrompt).toContain(".git");
|
||||
expect(localPrompt).toContain("destructive");
|
||||
expect(cloudPrompt).toContain(".git");
|
||||
expect(cloudPrompt).toContain("destructive");
|
||||
});
|
||||
|
||||
it("includes fallback instructions when mode is fallback", () => {
|
||||
const prompt = buildSystemPrompt([readTool], "fallback");
|
||||
expect(prompt).toContain("tool_call");
|
||||
expect(prompt.toLowerCase()).toContain("fallback");
|
||||
});
|
||||
|
||||
it("includes native mode instructions when mode is native (local)", () => {
|
||||
const prompt = buildSystemPrompt([readTool], "native", null, true);
|
||||
expect(prompt).toContain("native tool-call mode");
|
||||
});
|
||||
|
||||
it("includes native mode instructions when mode is native (cloud)", () => {
|
||||
const prompt = buildSystemPrompt([readTool], "native", null, false);
|
||||
expect(prompt).toContain("native tool-call mode");
|
||||
});
|
||||
|
||||
it("appends project instructions", () => {
|
||||
const prompt = buildSystemPrompt([readTool], "native", "Always use TypeScript strict mode.");
|
||||
expect(prompt).toContain("Always use TypeScript strict mode.");
|
||||
// Project instructions should be at the end
|
||||
const idx = prompt.indexOf("Always use TypeScript strict mode.");
|
||||
const safetyIdx = prompt.indexOf("## Safety");
|
||||
expect(idx).toBeGreaterThan(safetyIdx);
|
||||
});
|
||||
|
||||
it("works without project instructions", () => {
|
||||
const prompt = buildSystemPrompt([readTool], "native", null);
|
||||
expect(prompt).not.toContain("Project instructions");
|
||||
});
|
||||
|
||||
it("handles empty tool list", () => {
|
||||
const prompt = buildSystemPrompt([], "native");
|
||||
expect(prompt).toContain("Available tools");
|
||||
expect(prompt).toContain("Core principles");
|
||||
});
|
||||
|
||||
it("handles all read-only tools", () => {
|
||||
const tools = [dummyTool("read_file", false), dummyTool("grep", false), dummyTool("definition", false)];
|
||||
const prompt = buildSystemPrompt(tools, "native");
|
||||
// Tool list section should only have Read-only
|
||||
const toolSection = prompt.split("## Core principles")[0];
|
||||
expect(toolSection).toContain("Read-only: read_file, grep, definition");
|
||||
expect(toolSection).not.toContain("Mutating");
|
||||
});
|
||||
|
||||
it("handles all mutating tools", () => {
|
||||
const tools = [dummyTool("write_file", true), dummyTool("edit_file", true)];
|
||||
const prompt = buildSystemPrompt(tools, "native");
|
||||
const toolSection = prompt.split("## Core principles")[0];
|
||||
expect(toolSection).toContain("Mutating (requires confirmation): write_file, edit_file");
|
||||
expect(toolSection).not.toContain("Read-only");
|
||||
it("setMode (tool-call mode switch) preserves plan instructions while plan mode is active", () => {
|
||||
const session = createSession(fakeClient, "m", process.cwd(), async () => "once", "native", [dummyTool]);
|
||||
setPermissionMode(session, "plan");
|
||||
// Switch tool-call mode to fallback while still in plan mode — the rebuilt prompt must carry
|
||||
// BOTH the fallback tool instructions and the plan instructions.
|
||||
setMode(session, "fallback");
|
||||
const prompt = systemPromptOf(session);
|
||||
expect(prompt).toContain(FALLBACK_TOOL_INSTRUCTIONS);
|
||||
expect(prompt).toContain("Plan mode is ACTIVE");
|
||||
});
|
||||
});
|
||||
+72
-124
@@ -1,144 +1,92 @@
|
||||
import type { ToolCallMode } from "../backend/capabilityProbe.js";
|
||||
import type { PermissionMode } from "../permissions/types.js";
|
||||
import { FALLBACK_TOOL_INSTRUCTIONS } from "../toolcalling/fallbackPrompt.js";
|
||||
import type { ToolDef } from "../tools/types.js";
|
||||
|
||||
function formatToolList(tools: ToolDef[]): string {
|
||||
const readWrite = new Map<string, string[]>();
|
||||
for (const t of tools) {
|
||||
const category = t.mutating ? "Mutating (requires confirmation)" : "Read-only";
|
||||
const list = readWrite.get(category) ?? [];
|
||||
list.push(t.name);
|
||||
readWrite.set(category, list);
|
||||
}
|
||||
const parts: string[] = [];
|
||||
for (const [category, names] of readWrite) {
|
||||
parts.push(`${category}: ${names.join(", ")}`);
|
||||
}
|
||||
return parts.join("\n");
|
||||
}
|
||||
/** Injected into the system prompt only while the session is in plan mode. Tells the model it must
|
||||
* research read-only and present a plan rather than attempt changes (which would be blocked at the
|
||||
* tool gate anyway). Without this, the model only learns it's in plan mode from tool-rejection
|
||||
* errors, after it has already tried (and failed) to mutate. */
|
||||
const PLAN_INSTRUCTIONS = `Plan mode is ACTIVE. In this mode:
|
||||
- Do NOT call any mutating tool (write_file, edit_file, multi_edit, notebook_edit, git_commit, or a bash command that changes state). They are blocked and will return an error — that is expected.
|
||||
- Explore read-only first: use grep, list_files, read_file, and git_status to fully understand the request and the code it touches.
|
||||
- Then present a concrete implementation plan: the files you would change, the approach for each, and the key edits. Do not make the changes yet. Call the \`exit_plan_mode\` tool with the plan once you're ready for the user to approve it — on approval plan mode ends and you implement in this same turn; on rejection, refine and call it again. (If you can't call tools, present the plan as prose instead.)
|
||||
- Keep the plan focused and actionable so the user can review it.
|
||||
- Only ask a clarifying question if the request is still genuinely ambiguous after you've explored.`;
|
||||
|
||||
export function buildSystemPrompt(tools: ToolDef[], mode: ToolCallMode, projectInstructions?: string | null, isLocal?: boolean): string {
|
||||
const toolList = formatToolList(tools);
|
||||
export function buildSystemPrompt(
|
||||
tools: ToolDef[],
|
||||
mode: ToolCallMode,
|
||||
projectInstructions?: string | null,
|
||||
userMemory?: string | null,
|
||||
permissionMode?: PermissionMode,
|
||||
): string {
|
||||
const toolList = tools.map((t) => `- ${t.name}: ${t.description}`).join("\n");
|
||||
|
||||
// Auto-detect: if not explicitly specified, use the cloud prompt by default.
|
||||
// Callers (App.tsx) always pass isSmallLocalModel(baseURL, model) explicitly, so this
|
||||
// default only affects tests or edge cases without a baseURL/model.
|
||||
const useLocal = isLocal ?? false;
|
||||
const base = useLocal
|
||||
? buildLocalPrompt(toolList, mode)
|
||||
: buildCloudPrompt(toolList, mode);
|
||||
|
||||
return projectInstructions ? `${base}\n\n${projectInstructions}` : base;
|
||||
}
|
||||
|
||||
/** System prompt for local models (Ollama / LM Studio) — includes extra guidance about their
|
||||
* limitations (unreliable tool-call formatting, occasional empty/malformed responses).
|
||||
* Cloud models get a leaner prompt (buildCloudPrompt) that omits these assumptions. */
|
||||
function buildLocalPrompt(toolList: string, mode: ToolCallMode): string {
|
||||
return `You are a helpful coding assistant with access to tools for exploring and editing a codebase on the user's machine.
|
||||
|
||||
## Available tools
|
||||
const base = `You are a capable local coding assistant with access to tools on the user's machine.
|
||||
|
||||
Available tools:
|
||||
${toolList}
|
||||
|
||||
## Core principles
|
||||
Core workflow:
|
||||
1. Understand the user's goal before acting. Ask clarifying questions if the request is ambiguous or could destroy data.
|
||||
2. Explore the codebase efficiently: use grep to locate symbols/patterns, list_files to understand structure, and read_file only on the files or page ranges you actually need.
|
||||
3. Read files before editing them. Make the smallest change that solves the problem.
|
||||
4. Test your assumptions when possible (run typecheck/tests, read related code, verify file contents).
|
||||
5. Respond in plain text once you have enough information. Be concise; avoid restating obvious context.
|
||||
|
||||
1. **Inspect before answering.** Never guess file contents, function signatures, or directory structures — use read_file, list_files, grep, or definition to verify. Stale assumptions are worse than an extra tool call.
|
||||
Editing guidelines:
|
||||
- Prefer edit_file for small, targeted changes. Include enough surrounding context in old_string to make the match unique.
|
||||
- Use write_file for new files or when you are replacing most of a file's content.
|
||||
- Never invent file contents you haven't read; if unsure, read the file first.
|
||||
- When edit_file fails with "old_string not found", re-read the file and try again with a more precise match.
|
||||
|
||||
2. **Prefer small, targeted edits.** Use edit_file (or multi_edit for several changes in one file) for surgical changes. Use write_file only for new files or full rewrites. edit_file requires old_string to match exactly — copy the exact text from the file (read it first), including indentation and blank lines.
|
||||
Task tracking:
|
||||
- For non-trivial multi-step work (3+ steps), create tasks with task_create so progress is visible. Mark a task in_progress when you start it and completed when done.
|
||||
- Express dependencies with addBlocks/addBlockedBy (task ids) when one step must finish before another can start; check a task's blockedBy via task_get before starting it.
|
||||
- Set owner when a sub-agent or teammate will claim a specific task. Keep task subjects short and imperative.
|
||||
|
||||
3. **One tool call per response in fallback mode.** If you are in fallback mode (see below), call at most one tool per response and wait for the result before proceeding. In native mode you may call multiple read-only tools in parallel.
|
||||
Named teammates:
|
||||
- When you'll send a sub-agent several messages across the conversation, give the 'agent' call a 'name' (e.g. "researcher", "implementer") to create a named teammate. Then continue it with send_message using that 'name' instead of tracking its agentId, and use list_teammates to see your roster.
|
||||
- A named teammate runs in the shared working directory (a single, non-parallel delegation) and is resumable; parallel (worktree-isolated) delegations are fire-and-forget and ignore the name.
|
||||
- Teammate names must be unique per session — re-using an existing name returns an error (don't clobber a teammate in use); address the existing one via send_message instead.
|
||||
- Note: a named teammate still runs synchronously within your turn (it blocks until it answers). It is a stable, re-addressable handle, not a truly background process.
|
||||
|
||||
4. **Preserve existing style.** Match the surrounding code's indentation, naming conventions, quotes, and formatting. Don't reformat code outside the change scope.
|
||||
Worktree sessions:
|
||||
- To try changes without touching the main working tree, call enter_worktree (optionally with a name). It creates an isolated git worktree on a new branch at the current HEAD and switches your working directory into it — every file tool then operates there, and the worktree starts from the last commit so the user's uncommitted changes aren't carried over.
|
||||
- When done, call exit_worktree. Use action "keep" to preserve the work on its branch (recoverable later via git), or "remove" to discard it entirely (deletes the worktree and the branch). A remove is refused if the worktree has uncommitted changes unless you pass discardChanges: true.
|
||||
- You can only be in one worktree session at a time — exit before entering another. Don't switch models mid-worktree-session. /worktree shows the current worktree session.
|
||||
|
||||
5. **Keep answers concise.** When you have enough information, respond in plain text — don't pad with pleasantries or restated context. Code explanations should be brief and focused on the "why", not the "what" (the code already says what).
|
||||
Bash guidelines:
|
||||
- Destructive commands (rm, git push, git reset --hard, etc.) require explicit user confirmation via the tool's confirmation prompt.
|
||||
- For long-running commands, increase timeout_ms or press Ctrl+B while the command is running to background it; then use bash_output with the returned jobId.
|
||||
- Prefer git_status and git_commit for git work rather than raw git commands in bash.
|
||||
|
||||
6. **Recovery over retry.** If a tool call fails (edit_file "not found", bash non-zero exit, etc.), read the file or check the error output before retrying — don't repeat the same call. If edit_file suggests a closest match, use that text exactly.
|
||||
Tool-use discipline:
|
||||
- Call at most one tool at a time. Exception: when a single response delegates several independent sub-tasks to sub-agents, you may emit multiple 'agent' or 'agent__*' calls together — they run in parallel and their results come back in order. Writable parallel delegations (general-purpose, debugger, test-writer, or plugin agents) each run in an isolated throwaway git worktree at the last commit, so their file changes are discarded and only the returned answer matters. Read-only parallel delegations (explore, code-reviewer, planner) run in the shared working directory so they see current uncommitted state. Use parallel batches for research/review/planning; for implementation that must persist, delegate a single (sequential) agent call. Do not batch any other tool combinations. A single (non-parallel) delegation runs in the shared working directory, persists its edits, and is resumable — its result includes an agentId you can pass to send_message to continue it (refine the answer, ask a follow-up, or resume one that ran out of budget) without re-delegating from scratch. Parallel/worktree-isolated delegations are fire-and-forget and don't return an agentId.
|
||||
- Mutating tools (write_file, edit_file, bash, git_commit) require user confirmation unless the user has changed the permission mode or chosen "Yes, and don't ask again this session".
|
||||
- Use grep first when searching across many files; do not read_file dozens of files blindly.
|
||||
- If a tool returns an error or empty result, adapt: refine your grep pattern, check the path, or ask the user.
|
||||
- The session auto-continues a long tool chain internally up to a large per-turn budget. If it ever pauses because the budget was exhausted, briefly report progress and the user can continue with another message.
|
||||
- For deterministic multi-agent orchestration the user explicitly asks for ('use a workflow', 'fan out agents'), use the 'workflow' tool with a JS script that calls agent/parallel/pipeline/phase/log. Don't invoke it for ordinary single delegations — that's just the 'agent' tool. When fanning out writable agents (general-purpose, debugger, test-writer, or plugin agents) in parallel, pass each agent() call opts.isolation: 'worktree' so each runs in its own throwaway git worktree and their file writes can't collide (edits are discarded; only the returned answer matters). Read-only agents (explore, code-reviewer, planner) should NOT use isolation — they run in the shared cwd to see current uncommitted state.
|
||||
|
||||
7. **Respect confirmation.** Mutating tools (write_file, edit_file, multi_edit, notebook_edit, bash, git_commit) require user confirmation — you will see a permission prompt. Plan your edits so the user sees a clear, concise preview.
|
||||
Slash commands the user can type:
|
||||
- /undo — roll back your most recent write_file or edit_file to its previous content.
|
||||
- /summary — ask you to summarize the conversation without replacing history.
|
||||
|
||||
## Tool usage guide
|
||||
When to ask the user:
|
||||
- The request is ambiguous or underspecified.
|
||||
- A change would delete or overwrite significant user data.
|
||||
- You are about to push commits, force-delete branches, or run commands with side effects outside the project.
|
||||
- You cannot complete the task with the available tools or information.
|
||||
- For a decision that is genuinely the user's to make and that you can't resolve from the code or a sensible default, call the 'ask_user_question' tool with a short multiple-choice question (1-4 questions, 2-4 options each). Don't offload decisions you could make yourself — explore and pick a reasonable default first, and only ask when the choice truly changes what you do next.
|
||||
|
||||
- **read_file**: Start here. Use offset/limit for large files. Always read before editing.
|
||||
- **list_files**: Explore directory structure. Supports glob patterns like "src/**/*.ts".
|
||||
- **grep**: Search file contents. Prefer over read_file when you know what you're looking for.
|
||||
- **definition / references / diagnostics**: LSP-powered code intelligence. Use definition to find where a symbol is declared, references for all usages, diagnostics for type errors.
|
||||
- **edit_file**: For small changes to existing files. old_string must match exactly — include enough surrounding context to be unique. On mismatch, the tool suggests the closest similar text.
|
||||
- **multi_edit**: Apply several edits to the same file in one call. Each edit sees the result of previous edits, so adjust old_string for context shifts.
|
||||
- **write_file**: For new files or complete rewrites. Overwrites the entire file — use with care.
|
||||
- **bash**: Run shell commands. Prefer targeted tools (grep, definition) over broad shell commands when possible. Use timeout_ms for long-running commands. Background with Ctrl+B for very long commands.
|
||||
- **git_status / git_commit**: Inspect repo state and commit changes. Always check status before committing.
|
||||
- **web_search / web_fetch**: Look up information not in the local codebase. For API docs, error messages, or unfamiliar libraries.
|
||||
- **agent**: Delegate a sub-task to a focused sub-agent. Good for researching many files in parallel. Sub-agents cannot spawn further sub-agents.
|
||||
- **task_create / task_list / task_get / task_update**: Track structured work items with dependencies. Use for multi-step tasks (3+ steps) so progress is visible.
|
||||
- **todo_write**: Simple checklist for progress tracking. Good for linear step-by-step work.
|
||||
Personal memory:
|
||||
- A lightweight index of saved memory facts is included below when present; each line is 'name (type) — description'. Call the 'memory' tool with a name to read that fact's full body when its hook looks relevant to the current task.
|
||||
- 'memory_write' saves (action='write') or removes (action='delete') a typed fact: pick a kebab-case 'name', a one-line 'description' (the recall hook), a 'type' (user/feedback/project/reference), and the 'content' body. Proactively save durable facts the user states — preferences, working-style feedback, corrections worth remembering. Do not save transient per-task notes.`;
|
||||
|
||||
## Working with local models
|
||||
|
||||
- **Tool-call formatting can be unreliable.** If you're in fallback mode, follow the tool_call format strictly. If native mode produces errors, the system will automatically retry with fallback parsing.
|
||||
- **Empty or malformed responses can happen.** The system retries automatically, but if you see repeated failures, simplify your request.
|
||||
- **Output length may be limited.** For large file generations, prefer edit_file over write_file when possible — it uses fewer output tokens.
|
||||
|
||||
## Fallback mode
|
||||
|
||||
${mode === "fallback" ? FALLBACK_TOOL_INSTRUCTIONS : "You are in native tool-call mode. Call tools using the standard function-calling format. You may call multiple read-only tools in parallel, but mutating tools are always run sequentially."}
|
||||
|
||||
## Safety
|
||||
|
||||
- Do not modify .git directories or other version-control internals.
|
||||
- Do not delete large sections of code without clear justification and user confirmation.
|
||||
- When running bash commands, prefer read-only inspections (ls, cat, git status) over destructive operations (rm, git reset --hard).
|
||||
- If unsure about a destructive action, ask the user first rather than proceeding.`;
|
||||
const withToolMode = mode === "fallback" ? `${base}\n\n${FALLBACK_TOOL_INSTRUCTIONS}` : base;
|
||||
const withPlan = permissionMode === "plan" ? `${withToolMode}\n\n${PLAN_INSTRUCTIONS}` : withToolMode;
|
||||
const withProject = projectInstructions ? `${withPlan}\n\n${projectInstructions}` : withPlan;
|
||||
return userMemory ? `${withProject}\n\n${userMemory}` : withProject;
|
||||
}
|
||||
|
||||
/** System prompt for cloud models (large context window, reliable tool calls, no local-model quirks).
|
||||
* Leaner than the local prompt — skips the "Working with local models" section entirely and uses
|
||||
* a more direct tone, since cloud models don't need hand-holding about their own limitations. */
|
||||
function buildCloudPrompt(toolList: string, mode: ToolCallMode): string {
|
||||
return `You are a coding assistant with access to tools for exploring and editing a codebase on the user's machine.
|
||||
|
||||
## Available tools
|
||||
|
||||
${toolList}
|
||||
|
||||
## Core principles
|
||||
|
||||
1. **Inspect before answering.** Never guess file contents, function signatures, or directory structures — use read_file, list_files, grep, or definition to verify. Stale assumptions are worse than an extra tool call.
|
||||
|
||||
2. **Prefer small, targeted edits.** Use edit_file (or multi_edit for several changes in one file) for surgical changes. Use write_file only for new files or full rewrites. edit_file requires old_string to match exactly — copy the exact text from the file (read it first), including indentation and blank lines.
|
||||
|
||||
3. **Preserve existing style.** Match the surrounding code's indentation, naming conventions, quotes, and formatting. Don't reformat code outside the change scope.
|
||||
|
||||
4. **Keep answers concise.** When you have enough information, respond in plain text — don't pad with pleasantries or restated context. Code explanations should be brief and focused on the "why", not the "what" (the code already says what).
|
||||
|
||||
5. **Recovery over retry.** If a tool call fails (edit_file "not found", bash non-zero exit, etc.), read the file or check the error output before retrying — don't repeat the same call. If edit_file suggests a closest match, use that text exactly.
|
||||
|
||||
6. **Respect confirmation.** Mutating tools (write_file, edit_file, multi_edit, notebook_edit, bash, git_commit) require user confirmation — you will see a permission prompt. Plan your edits so the user sees a clear, concise preview.
|
||||
|
||||
## Tool usage guide
|
||||
|
||||
- **read_file**: Start here. Use offset/limit for large files. Always read before editing.
|
||||
- **list_files**: Explore directory structure. Supports glob patterns like "src/**/*.ts".
|
||||
- **grep**: Search file contents. Prefer over read_file when you know what you're looking for.
|
||||
- **definition / references / diagnostics**: LSP-powered code intelligence. Use definition to find where a symbol is declared, references for all usages, diagnostics for type errors.
|
||||
- **edit_file**: For small changes to existing files. old_string must match exactly — include enough surrounding context to be unique. On mismatch, the tool suggests the closest similar text.
|
||||
- **multi_edit**: Apply several edits to the same file in one call. Each edit sees the result of previous edits, so adjust old_string for context shifts.
|
||||
- **write_file**: For new files or complete rewrites. Overwrites the entire file — use with care.
|
||||
- **bash**: Run shell commands. Prefer targeted tools (grep, definition) over broad shell commands when possible. Use timeout_ms for long-running commands. Background with Ctrl+B for very long commands.
|
||||
- **git_status / git_commit**: Inspect repo state and commit changes. Always check status before committing.
|
||||
- **web_search / web_fetch**: Look up information not in the local codebase. For API docs, error messages, or unfamiliar libraries.
|
||||
- **agent**: Delegate a sub-task to a focused sub-agent. Good for researching many files in parallel. Sub-agents cannot spawn further sub-agents.
|
||||
- **task_create / task_list / task_get / task_update**: Track structured work items with dependencies. Use for multi-step tasks (3+ steps) so progress is visible.
|
||||
- **todo_write**: Simple checklist for progress tracking. Good for linear step-by-step work.
|
||||
|
||||
## ${mode === "fallback" ? "Fallback mode" : "Tool calling"}
|
||||
|
||||
${mode === "fallback" ? FALLBACK_TOOL_INSTRUCTIONS : "You are in native tool-call mode. Call tools using the standard function-calling format. You may call multiple read-only tools in parallel, but mutating tools are always run sequentially."}
|
||||
|
||||
## Safety
|
||||
|
||||
- Do not modify .git directories or other version-control internals.
|
||||
- Do not delete large sections of code without clear justification and user confirmation.
|
||||
- When running bash commands, prefer read-only inspections (ls, cat, git status) over destructive operations (rm, git reset --hard).
|
||||
- If unsure about a destructive action, ask the user first rather than proceeding.`;
|
||||
}
|
||||
@@ -7,7 +7,7 @@ import type { ToolCallMode } from "./capabilityProbe.js";
|
||||
const paths = envPaths("locode", { suffix: "" });
|
||||
const cacheFile = path.join(paths.config, "model-capabilities.json");
|
||||
|
||||
type Cache = Record<string, { mode: ToolCallMode; cachedAt?: number }>;
|
||||
type Cache = Record<string, ToolCallMode>;
|
||||
|
||||
function keyFor(baseURL: string, model: string): string {
|
||||
return `${baseURL}::${model}`;
|
||||
@@ -23,19 +23,7 @@ function load(): Cache {
|
||||
// Always re-read from disk if possible, so concurrent processes' writes aren't overwritten.
|
||||
if (existsSync(cacheFile)) {
|
||||
try {
|
||||
const raw = JSON.parse(readFileSync(cacheFile, "utf-8")) as Record<string, unknown>;
|
||||
// Migrate legacy format: bare string values ("native" | "fallback") become { mode, cachedAt }.
|
||||
const cache: Cache = {};
|
||||
for (const [k, v] of Object.entries(raw)) {
|
||||
if (typeof v === "string" && (v === "native" || v === "fallback")) {
|
||||
// Legacy entry — no cachedAt, so it can be re-validated on next probe.
|
||||
cache[k] = { mode: v };
|
||||
} else if (typeof v === "object" && v !== null && "mode" in v) {
|
||||
cache[k] = v as Cache[string];
|
||||
}
|
||||
// Silently drop unrecognized entries.
|
||||
}
|
||||
memoryCache = cache;
|
||||
memoryCache = JSON.parse(readFileSync(cacheFile, "utf-8")) as Cache;
|
||||
return memoryCache;
|
||||
} catch {
|
||||
memoryCache = {};
|
||||
@@ -53,36 +41,12 @@ function save(cache: Cache): void {
|
||||
void writeFileAtomic(cacheFile, JSON.stringify(cache, null, 2));
|
||||
}
|
||||
|
||||
/** How long a cached tool-call mode detection stays fresh before locode re-probes. A model's
|
||||
* tool-call capability rarely changes, but a transient probe failure (network timeout, 5xx)
|
||||
* can leave a stale "fallback" entry that permanently disables native tool calls. The TTL
|
||||
* ensures periodic re-validation. Set to 0 via LOCODE_CAPABILITY_CACHE_TTL_DAYS=0 to force
|
||||
* a re-probe every session. */
|
||||
const DEFAULT_CACHE_TTL_DAYS = 30;
|
||||
|
||||
function resolveCacheTtlDays(): number {
|
||||
const envValue = Number(process.env.LOCODE_CAPABILITY_CACHE_TTL_DAYS);
|
||||
if (Number.isFinite(envValue) && envValue >= 0 && envValue <= 365) return envValue;
|
||||
return DEFAULT_CACHE_TTL_DAYS;
|
||||
}
|
||||
|
||||
function isStale(entry: { cachedAt?: number }, ttlDays: number): boolean {
|
||||
if (ttlDays <= 0) return true;
|
||||
if (typeof entry.cachedAt !== "number") return true; // legacy entry — always re-probe
|
||||
const ageMs = Date.now() - entry.cachedAt;
|
||||
return ageMs > ttlDays * 24 * 60 * 60 * 1000;
|
||||
}
|
||||
|
||||
export function getCachedMode(baseURL: string, model: string): ToolCallMode | undefined {
|
||||
const entry = load()[keyFor(baseURL, model)];
|
||||
if (entry === undefined) return undefined;
|
||||
// Treat stale or legacy entries as a miss so the backend is re-probed.
|
||||
if (isStale(entry, resolveCacheTtlDays())) return undefined;
|
||||
return entry.mode;
|
||||
return load()[keyFor(baseURL, model)];
|
||||
}
|
||||
|
||||
export function setCachedMode(baseURL: string, model: string, mode: ToolCallMode): void {
|
||||
const cache = load();
|
||||
cache[keyFor(baseURL, model)] = { mode, cachedAt: Date.now() };
|
||||
cache[keyFor(baseURL, model)] = mode;
|
||||
save(cache);
|
||||
}
|
||||
@@ -1,5 +1,5 @@
|
||||
import OpenAI from "openai";
|
||||
import { resolveMaxRetries, resolveRequestTimeoutMs } from "../config/config.js";
|
||||
import { resolveRequestTimeoutMs } from "../config/config.js";
|
||||
import type { AppConfig } from "../config/types.js";
|
||||
|
||||
export function makeClient(cfg: AppConfig): OpenAI {
|
||||
@@ -15,10 +15,6 @@ export function makeClient(cfg: AppConfig): OpenAI {
|
||||
// requests behind a concurrency limit (e.g. Ollama's OLLAMA_NUM_PARALLEL) can legitimately take
|
||||
// longer than the 180s default to even start serving a request under contention.
|
||||
timeout: resolveRequestTimeoutMs(),
|
||||
// Configurable retries on transient failures (connection errors, 429, 5xx) with exponential
|
||||
// backoff. Defaults to 0 (fail immediately) to preserve the old behavior, since a local
|
||||
// backend's slow response usually means the model is stuck rather than a transient blip — but
|
||||
// raise via `maxRetries` / LOCODE_MAX_RETRIES for setups with occasional connection drops.
|
||||
maxRetries: resolveMaxRetries(),
|
||||
maxRetries: 0,
|
||||
});
|
||||
}
|
||||
|
||||
@@ -61,14 +61,9 @@ async function detectLmStudioContextWindow(baseURL: string, model: string): Prom
|
||||
* assume which one is actually running behind an OpenAI-compatible baseURL. Returns null (rather
|
||||
* than guessing) if neither responds usefully — callers should fall back to a configured default. */
|
||||
export async function detectContextWindow(baseURL: string, model: string): Promise<number | null> {
|
||||
// Try both backends in parallel to halve detection latency.
|
||||
const [ollama, lmStudio] = await Promise.allSettled([
|
||||
detectOllamaContextWindow(baseURL, model),
|
||||
detectLmStudioContextWindow(baseURL, model),
|
||||
]);
|
||||
if (ollama.status === "fulfilled" && ollama.value !== null) return ollama.value;
|
||||
if (lmStudio.status === "fulfilled" && lmStudio.value !== null) return lmStudio.value;
|
||||
return null;
|
||||
const ollama = await detectOllamaContextWindow(baseURL, model);
|
||||
if (ollama !== null) return ollama;
|
||||
return detectLmStudioContextWindow(baseURL, model);
|
||||
}
|
||||
|
||||
export interface ResolvedContextWindow {
|
||||
@@ -90,5 +85,5 @@ export async function resolveContextWindow(baseURL: string, model: string): Prom
|
||||
return { value: detected, isEstimate: false };
|
||||
}
|
||||
|
||||
return { value: resolveContextWindowDefault(baseURL, model), isEstimate: true };
|
||||
return { value: resolveContextWindowDefault(), isEstimate: true };
|
||||
}
|
||||
|
||||
@@ -4,8 +4,6 @@ import envPaths from "env-paths";
|
||||
import { afterEach, beforeEach, describe, expect, it } from "vitest";
|
||||
import { getCachedContextWindow } from "./contextWindowCache.js";
|
||||
|
||||
const KEY = "http://localhost:11434/v1::ttl-model";
|
||||
|
||||
const cacheFile = path.join(envPaths("locode", { suffix: "" }).config, "context-windows.json");
|
||||
|
||||
describe("contextWindowCache", () => {
|
||||
@@ -21,18 +19,16 @@ describe("contextWindowCache", () => {
|
||||
} else if (existsSync(cacheFile)) {
|
||||
rmSync(cacheFile);
|
||||
}
|
||||
delete process.env.LOCODE_CONTEXT_WINDOW_CACHE_TTL_DAYS;
|
||||
});
|
||||
|
||||
it("reads a well-formed cached entry", () => {
|
||||
// Written directly (rather than via setCachedContextWindow, whose write is fire-and-forget
|
||||
// async and would race this file-backed cache's always-read-from-disk load()) so the test is
|
||||
// deterministic and can't leak a pending write past its own afterEach cleanup. Includes a
|
||||
// fresh cachedAt so the TTL check treats it as current.
|
||||
// deterministic and can't leak a pending write past its own afterEach cleanup.
|
||||
mkdirSync(path.dirname(cacheFile), { recursive: true });
|
||||
writeFileSync(cacheFile, JSON.stringify({ "http://localhost:11434/v1::test-model": { value: 32768, isEstimate: false, cachedAt: Date.now() } }), "utf-8");
|
||||
writeFileSync(cacheFile, JSON.stringify({ "http://localhost:11434/v1::test-model": { value: 32768, isEstimate: false } }), "utf-8");
|
||||
|
||||
expect(getCachedContextWindow("http://localhost:11434/v1", "test-model")).toEqual({ value: 32768, isEstimate: false, cachedAt: expect.any(Number) });
|
||||
expect(getCachedContextWindow("http://localhost:11434/v1", "test-model")).toEqual({ value: 32768, isEstimate: false });
|
||||
});
|
||||
|
||||
it("treats a legacy bare-number cache entry as a miss instead of returning {value: undefined}", () => {
|
||||
@@ -46,41 +42,4 @@ describe("contextWindowCache", () => {
|
||||
const result = getCachedContextWindow("http://localhost:11434/v1", "legacy-model");
|
||||
expect(result).toBeUndefined();
|
||||
});
|
||||
|
||||
it("treats an entry without cachedAt as expired (re-detect)", () => {
|
||||
mkdirSync(path.dirname(cacheFile), { recursive: true });
|
||||
writeFileSync(cacheFile, JSON.stringify({ [KEY]: { value: 32768, isEstimate: false } }), "utf-8");
|
||||
// No cachedAt field — legacy entry from before TTL was added; should be a miss.
|
||||
expect(getCachedContextWindow("http://localhost:11434/v1", "ttl-model")).toBeUndefined();
|
||||
});
|
||||
|
||||
it("treats a fresh entry (recent cachedAt) as a hit", () => {
|
||||
mkdirSync(path.dirname(cacheFile), { recursive: true });
|
||||
writeFileSync(cacheFile, JSON.stringify({ [KEY]: { value: 32768, isEstimate: false, cachedAt: Date.now() } }), "utf-8");
|
||||
expect(getCachedContextWindow("http://localhost:11434/v1", "ttl-model")).toEqual({ value: 32768, isEstimate: false, cachedAt: expect.any(Number) });
|
||||
});
|
||||
|
||||
it("treats an old entry (cachedAt beyond TTL) as a miss", () => {
|
||||
mkdirSync(path.dirname(cacheFile), { recursive: true });
|
||||
// 30 days ago, default TTL is 7 days — stale.
|
||||
const old = Date.now() - 30 * 24 * 60 * 60 * 1000;
|
||||
writeFileSync(cacheFile, JSON.stringify({ [KEY]: { value: 32768, isEstimate: false, cachedAt: old } }), "utf-8");
|
||||
expect(getCachedContextWindow("http://localhost:11434/v1", "ttl-model")).toBeUndefined();
|
||||
});
|
||||
|
||||
it("respects a configured TTL of 0 (always re-detect)", () => {
|
||||
process.env.LOCODE_CONTEXT_WINDOW_CACHE_TTL_DAYS = "0";
|
||||
mkdirSync(path.dirname(cacheFile), { recursive: true });
|
||||
writeFileSync(cacheFile, JSON.stringify({ [KEY]: { value: 32768, isEstimate: false, cachedAt: Date.now() } }), "utf-8");
|
||||
expect(getCachedContextWindow("http://localhost:11434/v1", "ttl-model")).toBeUndefined();
|
||||
});
|
||||
|
||||
it("respects a longer configured TTL", () => {
|
||||
process.env.LOCODE_CONTEXT_WINDOW_CACHE_TTL_DAYS = "365";
|
||||
mkdirSync(path.dirname(cacheFile), { recursive: true });
|
||||
// 30 days ago, but TTL is now 365 days — fresh.
|
||||
const old = Date.now() - 30 * 24 * 60 * 60 * 1000;
|
||||
writeFileSync(cacheFile, JSON.stringify({ [KEY]: { value: 32768, isEstimate: false, cachedAt: old } }), "utf-8");
|
||||
expect(getCachedContextWindow("http://localhost:11434/v1", "ttl-model")).toBeDefined();
|
||||
});
|
||||
});
|
||||
|
||||
@@ -6,23 +6,9 @@ import { writeFileAtomic } from "../utils/writeFileAtomic.js";
|
||||
const paths = envPaths("locode", { suffix: "" });
|
||||
const cacheFile = path.join(paths.config, "context-windows.json");
|
||||
|
||||
/** How long a cached context-window detection stays fresh before locode re-detects it. A model's
|
||||
* context window rarely changes, but a backend can be reconfigured (quantization swapped, a
|
||||
* different model loaded under the same id, Ollama's `num_ctx` raised) — a TTL avoids pinning a
|
||||
* stale value forever. Set to 0 to disable caching (re-detect every session). */
|
||||
const DEFAULT_CACHE_TTL_DAYS = 7;
|
||||
|
||||
export function resolveCacheTtlDays(): number {
|
||||
const envValue = Number(process.env.LOCODE_CONTEXT_WINDOW_CACHE_TTL_DAYS);
|
||||
if (Number.isFinite(envValue) && envValue >= 0 && envValue <= 365) return envValue;
|
||||
return DEFAULT_CACHE_TTL_DAYS;
|
||||
}
|
||||
|
||||
export interface CachedContextWindow {
|
||||
value: number;
|
||||
isEstimate: boolean;
|
||||
/** Unix epoch ms when this entry was cached. Absent on legacy entries (treated as expired). */
|
||||
cachedAt?: number;
|
||||
}
|
||||
|
||||
type Cache = Record<string, CachedContextWindow>;
|
||||
@@ -53,17 +39,7 @@ function save(cache: Cache): void {
|
||||
void writeFileAtomic(cacheFile, JSON.stringify(cache, null, 2));
|
||||
}
|
||||
|
||||
/** Returns true when the entry is stale given the configured TTL. A TTL of 0 means "always
|
||||
* re-detect", so every entry is stale; a missing `cachedAt` (legacy entry) is also stale. */
|
||||
function isStale(entry: CachedContextWindow, ttlDays: number): boolean {
|
||||
if (ttlDays <= 0) return true;
|
||||
if (typeof entry.cachedAt !== "number") return true;
|
||||
const ageMs = Date.now() - entry.cachedAt;
|
||||
return ageMs > ttlDays * 24 * 60 * 60 * 1000;
|
||||
}
|
||||
|
||||
export function getCachedContextWindow(baseURL: string, model: string): CachedContextWindow | undefined {
|
||||
const ttlDays = resolveCacheTtlDays();
|
||||
const entry = load()[keyFor(baseURL, model)];
|
||||
if (entry === undefined) return undefined;
|
||||
// An older locode version cached a bare number instead of { value, isEstimate }. Treat that
|
||||
@@ -74,13 +50,11 @@ export function getCachedContextWindow(baseURL: string, model: string): CachedCo
|
||||
if (typeof entry !== "object" || entry === null || typeof (entry as CachedContextWindow).value !== "number") {
|
||||
return undefined;
|
||||
}
|
||||
// Expired entries are treated as a miss so the backend is re-queried and the entry refreshed.
|
||||
if (isStale(entry, ttlDays)) return undefined;
|
||||
return entry;
|
||||
}
|
||||
|
||||
export function setCachedContextWindow(baseURL: string, model: string, contextWindow: CachedContextWindow): void {
|
||||
const cache = load();
|
||||
cache[keyFor(baseURL, model)] = { ...contextWindow, cachedAt: Date.now() };
|
||||
cache[keyFor(baseURL, model)] = contextWindow;
|
||||
save(cache);
|
||||
}
|
||||
}
|
||||
|
||||
+18
-41
@@ -2,14 +2,13 @@ import { execa } from "execa";
|
||||
import { existsSync, mkdirSync } from "node:fs";
|
||||
import path from "node:path";
|
||||
import { Command } from "commander";
|
||||
import pkg from "../package.json" with { type: "json" };
|
||||
import { makeClient } from "./backend/client.js";
|
||||
import { ConfigError, resolveBackendConfig, resolveModel } from "./config/config.js";
|
||||
import { configFilePath, loadStoredConfig, saveStoredConfig, type StoredConfig } from "./config/store.js";
|
||||
import { loadMergedHooks, userHooksFilePath } from "./hooks/config.js";
|
||||
import { loadMergedServers, removeUserServer, saveUserServer, userMcpFilePath } from "./mcp/config.js";
|
||||
import { isHttpServerConfig } from "./mcp/types.js";
|
||||
import { deleteSession, listSessions, loadSession, mostRecentSessionId, sessionsDir } from "./persistence/sessionStore.js";
|
||||
import { deleteSession, listSessions, mostRecentSessionId, sessionsDir } from "./persistence/sessionStore.js";
|
||||
import { addInstalledPlugin, loadInstalledPlugins, pluginsDir, removeInstalledPlugin } from "./plugins/config.js";
|
||||
import { loadPlugin } from "./plugins/loader.js";
|
||||
import { runInkApp } from "./ui/ink/index.js";
|
||||
@@ -25,7 +24,7 @@ function addBackendOptions(cmd: Command): Command {
|
||||
program
|
||||
.name("locode")
|
||||
.description("Agentic coding CLI for local models via Ollama and LM Studio")
|
||||
.version(pkg.version);
|
||||
.version("0.3.1");
|
||||
|
||||
addBackendOptions(program)
|
||||
.option("-m, --model <name>", "model name as known to the backend")
|
||||
@@ -115,54 +114,44 @@ configCmd
|
||||
|
||||
configCmd
|
||||
.command("set <key> <value>")
|
||||
.description(
|
||||
"Persist a config value (backend, model, baseUrl, contextWindow, maxOutputTokens, maxIterations, autoCompactThreshold, requestTimeoutMs, subagentTimeoutMs, maxRetries, lspServers)",
|
||||
)
|
||||
.description("Persist a config value (backend, model, baseUrl, contextWindow, maxIterations, subagentMaxIterations, subagentMaxDepth, autoCompactThreshold, requestTimeoutMs, subagentTimeoutMs)")
|
||||
.action((key: string, value: string) => {
|
||||
if (
|
||||
key !== "backend" &&
|
||||
key !== "model" &&
|
||||
key !== "baseUrl" &&
|
||||
key !== "contextWindow" &&
|
||||
key !== "maxOutputTokens" &&
|
||||
key !== "maxIterations" &&
|
||||
key !== "subagentMaxIterations" &&
|
||||
key !== "subagentMaxDepth" &&
|
||||
key !== "autoCompactThreshold" &&
|
||||
key !== "requestTimeoutMs" &&
|
||||
key !== "subagentTimeoutMs" &&
|
||||
key !== "maxRetries" &&
|
||||
key !== "lspServers"
|
||||
key !== "subagentTimeoutMs"
|
||||
) {
|
||||
console.error(
|
||||
`Unknown config key "${key}". Valid keys: backend, model, baseUrl, contextWindow, maxOutputTokens, maxIterations, autoCompactThreshold, requestTimeoutMs, subagentTimeoutMs, maxRetries, lspServers`,
|
||||
`Unknown config key "${key}". Valid keys: backend, model, baseUrl, contextWindow, maxIterations, subagentMaxIterations, subagentMaxDepth, autoCompactThreshold, requestTimeoutMs, subagentTimeoutMs`,
|
||||
);
|
||||
process.exit(1);
|
||||
}
|
||||
const stored = loadStoredConfig();
|
||||
if (key === "lspServers") {
|
||||
// lspServers is a JSON object: { "<languageId>": { "command": "...", "args": [...], "extensions": [...] } }
|
||||
let parsed: unknown;
|
||||
try {
|
||||
parsed = JSON.parse(value);
|
||||
} catch {
|
||||
console.error(`lspServers must be a JSON object, got invalid JSON: ${value}`);
|
||||
process.exit(1);
|
||||
}
|
||||
if (typeof parsed !== "object" || parsed === null || Array.isArray(parsed)) {
|
||||
console.error(`lspServers must be a JSON object keyed by language id, got: ${value}`);
|
||||
process.exit(1);
|
||||
}
|
||||
stored.lspServers = parsed as Record<string, { command: string; args?: string[]; extensions?: string[] }>;
|
||||
} else if (key === "contextWindow" || key === "maxIterations") {
|
||||
if (key === "contextWindow" || key === "maxIterations") {
|
||||
const n = Number(value);
|
||||
if (!Number.isFinite(n) || n <= 0) {
|
||||
console.error(`${key} must be a positive number, got "${value}".`);
|
||||
process.exit(1);
|
||||
}
|
||||
stored[key] = n;
|
||||
} else if (key === "maxOutputTokens") {
|
||||
} else if (key === "subagentMaxIterations") {
|
||||
const n = Number(value);
|
||||
if (!Number.isFinite(n) || n < 256 || n > 1_000_000) {
|
||||
console.error(`maxOutputTokens must be between 256 and 1000000, got "${value}".`);
|
||||
if (!Number.isFinite(n) || n < 1 || n > 1000) {
|
||||
console.error(`subagentMaxIterations must be between 1 and 1000, got "${value}".`);
|
||||
process.exit(1);
|
||||
}
|
||||
stored[key] = n;
|
||||
} else if (key === "subagentMaxDepth") {
|
||||
const n = Number(value);
|
||||
if (!Number.isFinite(n) || n < 0 || n > 10) {
|
||||
console.error(`subagentMaxDepth must be between 0 and 10, got "${value}".`);
|
||||
process.exit(1);
|
||||
}
|
||||
stored[key] = n;
|
||||
@@ -187,13 +176,6 @@ configCmd
|
||||
process.exit(1);
|
||||
}
|
||||
stored[key] = n;
|
||||
} else if (key === "maxRetries") {
|
||||
const n = Number(value);
|
||||
if (!Number.isFinite(n) || n < 0 || n > 10) {
|
||||
console.error(`maxRetries must be between 0 and 10, got "${value}".`);
|
||||
process.exit(1);
|
||||
}
|
||||
stored[key] = n;
|
||||
} else {
|
||||
stored[key] = value;
|
||||
}
|
||||
@@ -230,11 +212,6 @@ sessionsCmd
|
||||
.action((id: string) => {
|
||||
if (deleteSession(id)) {
|
||||
console.log(`Deleted session ${id}.`);
|
||||
} else if (loadSession(id)) {
|
||||
// The file exists but deleteSession() couldn't actually remove it (e.g. locked by another
|
||||
// process) — a different situation from "no such session", so say so distinctly.
|
||||
console.error(`Could not delete session "${id}" — the file may be in use by another process.`);
|
||||
process.exit(1);
|
||||
} else {
|
||||
console.error(`No saved session found with id "${id}".`);
|
||||
process.exit(1);
|
||||
|
||||
@@ -1,60 +0,0 @@
|
||||
import { afterEach, beforeEach, describe, expect, it } from "vitest";
|
||||
import { configureLanguageSpecs, _resetSpecsForTests, _specsForTests } from "./lspManager.js";
|
||||
|
||||
// configureLanguageSpecs mutates the module's LANGUAGE_SPECS (forward-only by design — production
|
||||
// applies it once at startup). Tests restore the built-in defaults via _resetSpecsForTests so they
|
||||
// stay independent, then assert the merged spec list through _specsForTests (no server spawned).
|
||||
|
||||
function specFor(ext: string) {
|
||||
const specs = _specsForTests();
|
||||
return specs.find((s) => s.extensions.includes(ext)) ?? null;
|
||||
}
|
||||
|
||||
describe("configureLanguageSpecs", () => {
|
||||
beforeEach(() => _resetSpecsForTests());
|
||||
afterEach(() => _resetSpecsForTests());
|
||||
|
||||
it("leaves the built-in specs untouched for an empty override", () => {
|
||||
configureLanguageSpecs({});
|
||||
expect(_specsForTests().map((s) => s.languageId)).toEqual([
|
||||
"typescript",
|
||||
"python",
|
||||
"go",
|
||||
"rust",
|
||||
"c",
|
||||
]);
|
||||
// C and C++ share one clangd spec (no separate "cpp" entry).
|
||||
expect(specFor(".cpp")?.languageId).toBe("c");
|
||||
expect(specFor(".h")?.languageId).toBe("c");
|
||||
});
|
||||
|
||||
it("adds a brand-new language with extensions", () => {
|
||||
configureLanguageSpecs({ java: { command: "jdtls", extensions: [".java"] } });
|
||||
expect(specFor(".java")?.command).toBe("jdtls");
|
||||
expect(specFor(".java")?.languageId).toBe("java");
|
||||
});
|
||||
|
||||
it("ignores a new-language entry without extensions (can't route files to it)", () => {
|
||||
configureLanguageSpecs({ ruby: { command: "solargraph" } });
|
||||
expect(specFor(".rb")).toBeNull();
|
||||
});
|
||||
|
||||
it("overrides a built-in server's command and args", () => {
|
||||
configureLanguageSpecs({ typescript: { command: "my-tsserver", args: ["--stdio"] } });
|
||||
expect(specFor(".ts")?.command).toBe("my-tsserver");
|
||||
expect(specFor(".ts")?.args).toEqual(["--stdio"]);
|
||||
});
|
||||
|
||||
it("keeps a built-in's extensions when an override omits them", () => {
|
||||
configureLanguageSpecs({ python: { command: "basedpyright", args: ["--stdio"] } });
|
||||
expect(specFor(".py")?.command).toBe("basedpyright");
|
||||
expect(specFor(".pyi")?.languageId).toBe("python"); // extensions unchanged
|
||||
});
|
||||
|
||||
it("rewrites a built-in language's extensions when provided", () => {
|
||||
configureLanguageSpecs({ go: { command: "gopls", args: ["serve"], extensions: [".rs"] } });
|
||||
// .rs now routes to "go", not "rust".
|
||||
expect(specFor(".rs")?.languageId).toBe("go");
|
||||
expect(specFor(".go")).toBeNull(); // .go no longer claimed by go
|
||||
});
|
||||
});
|
||||
@@ -1,452 +0,0 @@
|
||||
import { spawn, type ChildProcess } from "node:child_process";
|
||||
import path from "node:path";
|
||||
import { readFile as fsReadFile } from "node:fs/promises";
|
||||
import {
|
||||
createProtocolConnection,
|
||||
DidChangeTextDocumentNotification,
|
||||
DidOpenTextDocumentNotification,
|
||||
DefinitionRequest,
|
||||
ReferencesRequest,
|
||||
type ProtocolConnection,
|
||||
type TextDocumentIdentifier,
|
||||
type Position,
|
||||
type Location,
|
||||
type Diagnostic,
|
||||
} from "vscode-languageserver-protocol";
|
||||
import { StreamMessageReader, StreamMessageWriter } from 'vscode-languageserver-protocol/node';
|
||||
import { URI } from "vscode-uri";
|
||||
import type { MarkupContent } from "vscode-languageserver-protocol";
|
||||
|
||||
/** LSP diagnostic messages can be either a plain string or a { kind, value } MarkupContent object.
|
||||
* locode's tool surface deals in plain strings, so flatten either form to text. */
|
||||
function messageToString(message: string | MarkupContent): string {
|
||||
if (typeof message === "string") return message;
|
||||
return message?.value ?? "";
|
||||
}
|
||||
|
||||
/** A connected LSP server for one language, plus its child process so we can clean it up. */
|
||||
interface LspHandle {
|
||||
connection: ProtocolConnection;
|
||||
child: ChildProcess;
|
||||
languageId: string;
|
||||
/** Open documents we've already sent didOpen for, so we send didChange (not didOpen) on edits. */
|
||||
openDocs: Set<string>;
|
||||
}
|
||||
|
||||
/** Maps a file extension to a language id (the LSP "languageId" string) and the server command to
|
||||
* spawn for it. Only one server per language is ever spawned (lazy, on first use). A missing entry
|
||||
* means locode has no built-in mapping — the user can still point a server at it via config in a
|
||||
* future extension. The command is resolved on the PATH; if it isn't installed the spawn fails and
|
||||
* the tool returns a clear "install X" error rather than a silent no-op. */
|
||||
interface LanguageSpec {
|
||||
languageId: string;
|
||||
extensions: string[];
|
||||
/** The server command (no args). Must be on PATH. */
|
||||
command: string;
|
||||
/** Args passed to the server command. */
|
||||
args?: string[];
|
||||
}
|
||||
|
||||
// The built-in language→server mappings. Mutable so `configureLanguageSpecs` can merge in user
|
||||
// overrides/additions from config (see config.ts `lspServers`). One clangd spec covers both C
|
||||
// and C++ — clangd handles both, and merging avoids spawning a second clangd for a mixed C/C++
|
||||
// project (two servers keyed by separate languageIds would each index the same headers twice).
|
||||
let LANGUAGE_SPECS: LanguageSpec[] = [
|
||||
// TypeScript / JavaScript — `typescript-language-server` wraps tsserver and speaks LSP. The most
|
||||
// common local-model codebase shape, so it's the first one locode wires up.
|
||||
{
|
||||
languageId: "typescript",
|
||||
extensions: [".ts", ".tsx", ".mts", ".cts", ".js", ".jsx", ".mjs", ".cjs"],
|
||||
command: "typescript-language-server",
|
||||
args: ["--stdio"],
|
||||
},
|
||||
{
|
||||
languageId: "python",
|
||||
extensions: [".py", ".pyi"],
|
||||
command: "pyright-langserver",
|
||||
args: ["--stdio"],
|
||||
},
|
||||
{
|
||||
languageId: "go",
|
||||
extensions: [".go"],
|
||||
command: "gopls",
|
||||
args: ["serve"],
|
||||
},
|
||||
{
|
||||
languageId: "rust",
|
||||
extensions: [".rs"],
|
||||
command: "rust-analyzer",
|
||||
},
|
||||
// C and C++ share clangd. The languageId is "c" (clangd treats .cpp/.hpp the same way);
|
||||
// all C/C++ extensions route to the single clangd process.
|
||||
{
|
||||
languageId: "c",
|
||||
extensions: [".c", ".h", ".cpp", ".cc", ".cxx", ".hpp", ".hh", ".hxx"],
|
||||
command: "clangd",
|
||||
},
|
||||
];
|
||||
|
||||
/** Merge user-configured LSP server entries (from `locode config set lspServers`) into the
|
||||
* built-in specs. An entry keyed by a built-in languageId overrides that spec's command/args
|
||||
* and, if `extensions` is provided, which file extensions route to it. An entry keyed by a new
|
||||
* languageId (e.g. "java", "ruby") adds a brand-new mapping — it MUST supply `extensions` so
|
||||
* files can be routed to it. Call once at startup; idempotent against the built-in list.
|
||||
*
|
||||
* Entries missing a `command` are ignored (a server we can't spawn is useless), and entries for
|
||||
* new ids without `extensions` are ignored too (no way to route files to them). */
|
||||
export function configureLanguageSpecs(overrides: Record<string, { command: string; args?: string[]; extensions?: string[] }>): void {
|
||||
const merged: LanguageSpec[] = LANGUAGE_SPECS.map((spec) => {
|
||||
const ov = overrides[spec.languageId];
|
||||
if (!ov) return spec;
|
||||
return {
|
||||
languageId: spec.languageId,
|
||||
extensions: ov.extensions ?? spec.extensions,
|
||||
command: ov.command,
|
||||
args: ov.args,
|
||||
};
|
||||
});
|
||||
for (const [languageId, ov] of Object.entries(overrides)) {
|
||||
if (merged.some((s) => s.languageId === languageId)) continue; // already a built-in we overrode
|
||||
if (!ov.command || !ov.extensions || ov.extensions.length === 0) continue;
|
||||
merged.push({ languageId, extensions: ov.extensions, command: ov.command, args: ov.args });
|
||||
}
|
||||
LANGUAGE_SPECS = merged;
|
||||
}
|
||||
|
||||
/** Picks the LanguageSpec for a file path, or null if no extension matches. */
|
||||
function specForFile(filePath: string): LanguageSpec | null {
|
||||
const ext = path.extname(filePath).toLowerCase();
|
||||
if (!ext) return null;
|
||||
return LANGUAGE_SPECS.find((s) => s.extensions.includes(ext)) ?? null;
|
||||
}
|
||||
|
||||
/** A per-workspace (cwd) registry of live LSP servers, keyed by language id. One server per
|
||||
* language per cwd — a second project gets its own manager (locode is single-session-per-process
|
||||
* today, but keying on cwd keeps it correct if that ever changes). */
|
||||
const handles = new Map<string, LspHandle>();
|
||||
|
||||
/** Convert an absolute filesystem path to an LSP file:// URI string. */
|
||||
function toUri(absPath: string): string {
|
||||
return URI.file(absPath).toString();
|
||||
}
|
||||
|
||||
interface LocResult {
|
||||
path: string;
|
||||
line: number;
|
||||
column: number;
|
||||
}
|
||||
|
||||
function toLocation(loc: Location): LocResult {
|
||||
return {
|
||||
path: URI.parse(loc.uri).fsPath,
|
||||
line: loc.range.start.line + 1,
|
||||
column: loc.range.start.character + 1,
|
||||
};
|
||||
}
|
||||
|
||||
/** Spawns the LSP server for `spec`, initializes it, and returns a live handle. Throws a clear,
|
||||
* actionable error if the server binary isn't on the PATH (the most common failure) so the tool
|
||||
* can surface "install typescript-language-server" instead of an opaque spawn ENOENT. */
|
||||
async function startServer(spec: LanguageSpec, cwd: string): Promise<LspHandle> {
|
||||
let child: ChildProcess;
|
||||
try {
|
||||
// npm installs global CLI packages on Windows as .cmd/.ps1 shims, not raw .exe files — spawn()
|
||||
// can't resolve those without shell:true, so a genuinely-installed server would otherwise ENOENT.
|
||||
child = spawn(spec.command, spec.args ?? [], {
|
||||
cwd,
|
||||
stdio: ["pipe", "pipe", "pipe"],
|
||||
shell: process.platform === "win32",
|
||||
});
|
||||
} catch (err) {
|
||||
throw new Error(
|
||||
`Could not start the LSP server "${spec.command}" for ${spec.languageId}. Is it installed and on your PATH? (${(err as Error).message})`,
|
||||
);
|
||||
}
|
||||
// spawn() itself rarely throws synchronously — a missing binary (ENOENT) instead fires an
|
||||
// async 'error' event on the child process. With no listener, that event is unhandled and
|
||||
// crashes the whole process, so wait for either a successful spawn or that error before
|
||||
// proceeding.
|
||||
await new Promise<void>((resolve, reject) => {
|
||||
child.once("spawn", () => resolve());
|
||||
child.once("error", (err) => {
|
||||
reject(
|
||||
new Error(
|
||||
`Could not start the LSP server "${spec.command}" for ${spec.languageId}. Is it installed and on your PATH? (${(err as Error).message})`,
|
||||
),
|
||||
);
|
||||
});
|
||||
});
|
||||
// After startup, a late 'error' (e.g. the process dying unexpectedly) must not go unhandled
|
||||
// either — the 'exit' handler below already disposes the connection, so just swallow it here.
|
||||
child.on("error", () => {});
|
||||
if (!child.stdin || !child.stdout) {
|
||||
child.kill();
|
||||
throw new Error(`LSP server "${spec.command}" did not open stdio streams.`);
|
||||
}
|
||||
|
||||
const reader = new StreamMessageReader(child.stdout);
|
||||
const writer = new StreamMessageWriter(child.stdin);
|
||||
const connection = createProtocolConnection(reader, writer);
|
||||
// vscode-jsonrpc buffers all messages until listen() starts pumping them — without this,
|
||||
// sendRequest hangs (nothing is ever written) or throws "Call listen() first."
|
||||
connection.listen();
|
||||
|
||||
// Surface stderr so a crashing server isn't a silent void (matches locode's MCP stdio policy).
|
||||
child.stderr?.on("data", () => {
|
||||
// Discard by default; a future debug mode could surface this. Don't let it back up.
|
||||
});
|
||||
|
||||
await connection.sendRequest("initialize", {
|
||||
processId: process.pid,
|
||||
rootUri: URI.file(cwd).toString(),
|
||||
capabilities: {
|
||||
// locode consumes definition/references/diagnostics; declare only those so a server doesn't
|
||||
// waste effort enabling features we'll never query. Full text sync (change=1) is simplest and
|
||||
// correct — we always resend the whole file, never a range edit.
|
||||
textDocumentSync: { openClose: true, change: 1 },
|
||||
definitionProvider: true,
|
||||
referencesProvider: true,
|
||||
},
|
||||
workspaceFolders: [{ uri: URI.file(cwd).toString(), name: path.basename(cwd) || cwd }],
|
||||
});
|
||||
// Per LSP spec, the client must send `initialized` after the initialize response.
|
||||
await connection.sendNotification("initialized", {});
|
||||
|
||||
// A server crash should reject any in-flight request rather than hanging forever — listen for
|
||||
// exit and dispose the connection so the next call throws instead of awaiting a dead process.
|
||||
child.on("exit", () => {
|
||||
connection.dispose();
|
||||
handles.delete(`${cwd}::${spec.languageId}`);
|
||||
});
|
||||
|
||||
return { connection, child, languageId: spec.languageId, openDocs: new Set() };
|
||||
}
|
||||
|
||||
/** Returns the (lazily-started) LSP handle for the language owning `filePath`, or throws if no
|
||||
* server is configured/can't start. The first call for a language pays the initialize round-trip;
|
||||
* every later call reuses the live server. */
|
||||
async function handleForFile(filePath: string, cwd: string): Promise<LspHandle> {
|
||||
const spec = specForFile(filePath);
|
||||
if (!spec) {
|
||||
throw new Error(`No LSP server configured for "${path.extname(filePath)}" (code intelligence supports: ${LANGUAGE_SPECS.map((s) => s.extensions[0]).join(", ")}).`);
|
||||
}
|
||||
const key = `${cwd}::${spec.languageId}`;
|
||||
let handle = handles.get(key);
|
||||
if (!handle) {
|
||||
handle = await startServer(spec, cwd);
|
||||
handles.set(key, handle);
|
||||
}
|
||||
return handle;
|
||||
}
|
||||
|
||||
/** Ensures the LSP server knows the current on-disk contents of `filePath`. Sends didOpen the
|
||||
* first time a file is touched, didChange on subsequent syncs (the file was edited on disk since).
|
||||
* Reads the file fresh each time — locode's tools write to disk before this runs, so the disk is
|
||||
* the source of truth, not any in-memory buffer. */
|
||||
async function syncDocument(handle: LspHandle, absPath: string, cwd: string): Promise<void> {
|
||||
const uri = toUri(absPath);
|
||||
const content = await fsReadFile(absPath, "utf-8");
|
||||
if (!handle.openDocs.has(uri)) {
|
||||
await handle.connection.sendNotification(DidOpenTextDocumentNotification.type, {
|
||||
textDocument: { uri, languageId: handle.languageId, version: 1, text: content },
|
||||
});
|
||||
handle.openDocs.add(uri);
|
||||
} else {
|
||||
await handle.connection.sendNotification(DidChangeTextDocumentNotification.type, {
|
||||
textDocument: { uri, version: Date.now() },
|
||||
contentChanges: [{ text: content }],
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
export interface DefinitionResult {
|
||||
/** The file/line/column of the symbol's definition. Multiple entries if the symbol has more than
|
||||
* one definition (interface implementations, overloads, partial classes). Empty if the server
|
||||
* found none (undefined symbol, or the server couldn't resolve it). */
|
||||
definitions: LocResult[];
|
||||
}
|
||||
|
||||
/** Resolves where the symbol at `line`/`column` (1-indexed) in `filePath` is defined. Syncs the
|
||||
* document first so the server's view matches disk. Returns an empty list (not an error) when the
|
||||
* server has no definition to offer — that's a legitimate "not found", not a failure. */
|
||||
export async function getDefinition(
|
||||
filePath: string,
|
||||
line: number,
|
||||
column: number,
|
||||
cwd: string,
|
||||
): Promise<DefinitionResult> {
|
||||
const absPath = path.resolve(cwd, filePath);
|
||||
const handle = await handleForFile(filePath, cwd);
|
||||
await syncDocument(handle, absPath, cwd);
|
||||
const pos: Position = { line: line - 1, character: column - 1 };
|
||||
const result = (await handle.connection.sendRequest(DefinitionRequest.type, {
|
||||
textDocument: { uri: toUri(absPath) } as TextDocumentIdentifier,
|
||||
position: pos,
|
||||
})) as Location | Location[] | null;
|
||||
const locs = Array.isArray(result) ? result : result ? [result] : [];
|
||||
return { definitions: locs.map(toLocation) };
|
||||
}
|
||||
|
||||
export interface ReferencesResult {
|
||||
/** Every place the symbol at `line`/`column` is referenced (including its definition). */
|
||||
references: LocResult[];
|
||||
}
|
||||
|
||||
/** Finds every reference to the symbol at `line`/`column` in `filePath`. `includeDeclaration`
|
||||
* defaults to true (matches most IDE "find all references" behavior). */
|
||||
export async function getReferences(
|
||||
filePath: string,
|
||||
line: number,
|
||||
column: number,
|
||||
cwd: string,
|
||||
includeDeclaration = true,
|
||||
): Promise<ReferencesResult> {
|
||||
const absPath = path.resolve(cwd, filePath);
|
||||
const handle = await handleForFile(filePath, cwd);
|
||||
await syncDocument(handle, absPath, cwd);
|
||||
const pos: Position = { line: line - 1, character: column - 1 };
|
||||
const result = (await handle.connection.sendRequest(ReferencesRequest.type, {
|
||||
textDocument: { uri: toUri(absPath) } as TextDocumentIdentifier,
|
||||
position: pos,
|
||||
context: { includeDeclaration },
|
||||
})) as Location[] | null;
|
||||
return { references: (result ?? []).map(toLocation) };
|
||||
}
|
||||
|
||||
/** Notifies the LSP server that `filePath` changed on disk, so a subsequent `diagnostics` call
|
||||
* reflects the new content. Called from the FileChanged hook path after edit_file/write_file. If no
|
||||
* server is running for this language (or the file isn't one we manage), this is a no-op — it must
|
||||
* never throw from a hook context, since hooks fire on every mutating tool. */
|
||||
export async function notifyFileChanged(filePath: string, cwd: string): Promise<void> {
|
||||
try {
|
||||
const spec = specForFile(filePath);
|
||||
if (!spec) return;
|
||||
const key = `${cwd}::${spec.languageId}`;
|
||||
const handle = handles.get(key);
|
||||
if (!handle) return; // No server started yet — diagnostics will sync on first query.
|
||||
await syncDocument(handle, path.resolve(cwd, filePath), cwd);
|
||||
} catch {
|
||||
// Best-effort: a hook context can't propagate errors into the turn.
|
||||
}
|
||||
}
|
||||
|
||||
type Severity = "error" | "warning" | "information" | "hint";
|
||||
|
||||
export interface DiagnosticsResult {
|
||||
diagnostics: { path: string; line: number; column: number; severity: Severity; message: string; source?: string }[];
|
||||
}
|
||||
|
||||
/** The most recent diagnostics the server has published for `filePath`. LSP pushes diagnostics via
|
||||
* `textDocument/publishDiagnostics` notifications; locode collects them per-URI as they arrive and
|
||||
* returns the latest snapshot here. Forces a document sync first so the snapshot is current. */
|
||||
const diagnosticsByUri = new Map<string, Diagnostic[]>();
|
||||
|
||||
// Per-URI resolvers waiting on the next publishDiagnostics notification. getDiagnostics arms one
|
||||
// for the file it just synced, then races it against a timeout — so a slow server (tsserver on a
|
||||
// large file) still gets a chance to publish the fresh snapshot rather than the caller reading a
|
||||
// stale one after a single event-loop turn. Resolved and cleared by the publishDiagnostics handler.
|
||||
const diagWaiters = new Map<string, () => void>();
|
||||
|
||||
/** Wait for the next publishDiagnostics for `uri`, or give up after `timeoutMs`. Resolves true if
|
||||
* a publish arrived, false on timeout. The waiter is removed either way. */
|
||||
function waitForDiagnostics(uri: string, timeoutMs: number): Promise<boolean> {
|
||||
return new Promise((resolve) => {
|
||||
const timer = setTimeout(() => {
|
||||
diagWaiters.delete(uri);
|
||||
resolve(false);
|
||||
}, timeoutMs);
|
||||
diagWaiters.set(uri, () => {
|
||||
clearTimeout(timer);
|
||||
diagWaiters.delete(uri);
|
||||
resolve(true);
|
||||
});
|
||||
});
|
||||
}
|
||||
|
||||
const SEVERITY_MAP: Record<number, Severity> = {
|
||||
1: "error",
|
||||
2: "warning",
|
||||
3: "information",
|
||||
4: "hint",
|
||||
};
|
||||
|
||||
export async function getDiagnostics(filePath: string, cwd: string): Promise<DiagnosticsResult> {
|
||||
const absPath = path.resolve(cwd, filePath);
|
||||
const handle = await handleForFile(filePath, cwd);
|
||||
const uri = toUri(absPath);
|
||||
// Attach a per-connection diagnostic collector the first time we use this handle.
|
||||
if (!(handle as unknown as { __diagWired?: boolean }).__diagWired) {
|
||||
(handle as unknown as { __diagWired?: boolean }).__diagWired = true;
|
||||
handle.connection.onNotification("textDocument/publishDiagnostics", (params: { uri: string; diagnostics: Diagnostic[] }) => {
|
||||
diagnosticsByUri.set(params.uri, params.diagnostics);
|
||||
// Wake a getDiagnostics call waiting on this URI, if any.
|
||||
diagWaiters.get(params.uri)?.();
|
||||
});
|
||||
}
|
||||
// Clear any stale snapshot for this URI before syncing so a timeout fallthrough can't return
|
||||
// diagnostics from before the edit. The server publishes asynchronously after didChange; race
|
||||
// its next publish against a short timeout so a slow server (tsserver on a large file) still
|
||||
// gets a chance to compute fresh diagnostics rather than us reading a stale snapshot after one
|
||||
// event-loop turn. Fall through to whatever's cached on timeout (possibly empty).
|
||||
diagnosticsByUri.delete(uri);
|
||||
await syncDocument(handle, absPath, cwd);
|
||||
await waitForDiagnostics(uri, 1500);
|
||||
const diags = diagnosticsByUri.get(uri) ?? [];
|
||||
return {
|
||||
diagnostics: diags.map((d) => ({
|
||||
path: URI.parse(uri).fsPath,
|
||||
line: (d.range?.start.line ?? 0) + 1,
|
||||
column: (d.range?.start.character ?? 0) + 1,
|
||||
severity: SEVERITY_MAP[d.severity ?? 1] ?? "information",
|
||||
message: messageToString(d.message ?? ""),
|
||||
source: d.source,
|
||||
})),
|
||||
};
|
||||
}
|
||||
|
||||
/** Shuts down every live LSP server. Call on locode exit so spawned servers (tsserver, pyright,
|
||||
* gopls, …) don't outlive the process as orphans. Awaits each shutdown so the signals land before
|
||||
* teardown. Best-effort: a stuck server can't block exit forever (the child kill still fires). */
|
||||
export async function shutdownAll(): Promise<void> {
|
||||
const all = [...handles.values()];
|
||||
handles.clear();
|
||||
await Promise.allSettled(
|
||||
all.map(async (h) => {
|
||||
try {
|
||||
await h.connection.sendRequest("shutdown", null);
|
||||
h.connection.sendNotification("exit", {});
|
||||
} catch {
|
||||
// Already dead — fall through to kill.
|
||||
}
|
||||
h.child.kill();
|
||||
}),
|
||||
);
|
||||
}
|
||||
|
||||
/** For tests only: clear the live-handle registry and diagnostic cache without spawning/killing. */
|
||||
export function _resetForTests(): void {
|
||||
handles.clear();
|
||||
diagnosticsByUri.clear();
|
||||
diagWaiters.clear();
|
||||
}
|
||||
|
||||
/** For tests only: a snapshot of the currently configured language specs (after any
|
||||
* configureLanguageSpecs merge), so tests can assert the merge without spawning a server. */
|
||||
export function _specsForTests(): readonly LanguageSpec[] {
|
||||
return LANGUAGE_SPECS;
|
||||
}
|
||||
|
||||
// The immutable built-in spec list, kept so tests can restore LANGUAGE_SPECS to defaults after a
|
||||
// configureLanguageSpecs call (the merge is forward-only by design — production applies it once).
|
||||
const BUILTIN_LANGUAGE_SPECS: readonly LanguageSpec[] = [
|
||||
{ languageId: "typescript", extensions: [".ts", ".tsx", ".mts", ".cts", ".js", ".jsx", ".mjs", ".cjs"], command: "typescript-language-server", args: ["--stdio"] },
|
||||
{ languageId: "python", extensions: [".py", ".pyi"], command: "pyright-langserver", args: ["--stdio"] },
|
||||
{ languageId: "go", extensions: [".go"], command: "gopls", args: ["serve"] },
|
||||
{ languageId: "rust", extensions: [".rs"], command: "rust-analyzer" },
|
||||
{ languageId: "c", extensions: [".c", ".h", ".cpp", ".cc", ".cxx", ".hpp", ".hh", ".hxx"], command: "clangd" },
|
||||
];
|
||||
|
||||
/** For tests only: restore the built-in language specs (undo any configureLanguageSpecs merge). */
|
||||
export function _resetSpecsForTests(): void {
|
||||
LANGUAGE_SPECS = BUILTIN_LANGUAGE_SPECS.map((s) => ({ ...s }));
|
||||
}
|
||||
+52
-60
@@ -3,8 +3,8 @@ import { existsSync, mkdtempSync, rmSync } from "node:fs";
|
||||
import path from "node:path";
|
||||
import os from "node:os";
|
||||
import { _setConfigFilePathForTest, loadStoredConfig, saveStoredConfig } from "./store.js";
|
||||
import { resolveAutoCompactThreshold, resolveContextWindowDefault, resolveMaxIterations, resolveMaxOutputTokens, resolveMaxRetries, resolveRequestTimeoutMs, resolveSubagentTimeoutMs } from "./config.js";
|
||||
import { DEFAULT_AUTO_COMPACT_THRESHOLD, DEFAULT_CONTEXT_WINDOW_LOCAL, DEFAULT_CONTEXT_WINDOW_CLOUD, DEFAULT_MAX_ITERATIONS, DEFAULT_MAX_OUTPUT_TOKENS, DEFAULT_MAX_RETRIES, DEFAULT_REQUEST_TIMEOUT_MS, DEFAULT_SUBAGENT_TIMEOUT_MS } from "./defaults.js";
|
||||
import { resolveAutoCompactThreshold, resolveContextWindowDefault, resolveMaxIterations, resolveRequestTimeoutMs, resolveSubagentMaxDepth, resolveSubagentMaxIterations, resolveSubagentTimeoutMs } from "./config.js";
|
||||
import { DEFAULT_AUTO_COMPACT_THRESHOLD, DEFAULT_CONTEXT_WINDOW, DEFAULT_MAX_ITERATIONS, DEFAULT_REQUEST_TIMEOUT_MS, DEFAULT_SUBAGENT_MAX_DEPTH, DEFAULT_SUBAGENT_MAX_ITERATIONS, DEFAULT_SUBAGENT_TIMEOUT_MS } from "./defaults.js";
|
||||
|
||||
// Isolate the persisted config to a temp directory so the suite never reads or overwrites the
|
||||
// user's real ~/.config/locode/config.json (the previous afterEach { saveStoredConfig({}) } wiped
|
||||
@@ -25,10 +25,9 @@ describe("config resolution", () => {
|
||||
delete process.env.LOCODE_AUTO_COMPACT_THRESHOLD;
|
||||
delete process.env.LOCODE_CONTEXT_WINDOW;
|
||||
delete process.env.LOCODE_MAX_ITERATIONS;
|
||||
delete process.env.LOCODE_MAX_OUTPUT_TOKENS;
|
||||
delete process.env.LOCODE_MAX_OUTPUT_TOKENS;
|
||||
delete process.env.LOCODE_MAX_RETRIES;
|
||||
delete process.env.LOCODE_REQUEST_TIMEOUT_MS;
|
||||
delete process.env.LOCODE_SUBAGENT_MAX_ITERATIONS;
|
||||
delete process.env.LOCODE_SUBAGENT_MAX_DEPTH;
|
||||
delete process.env.LOCODE_SUBAGENT_TIMEOUT_MS;
|
||||
});
|
||||
|
||||
@@ -53,46 +52,14 @@ describe("config resolution", () => {
|
||||
expect(resolveAutoCompactThreshold()).toBe(DEFAULT_AUTO_COMPACT_THRESHOLD);
|
||||
});
|
||||
|
||||
it("resolves context window default — small local model", () => {
|
||||
expect(resolveContextWindowDefault("http://localhost:11434/v1", "llama3.2:latest")).toBe(DEFAULT_CONTEXT_WINDOW_LOCAL);
|
||||
});
|
||||
|
||||
it("resolves context window default — remote cloud API", () => {
|
||||
expect(resolveContextWindowDefault("https://api.openai.com/v1", "gpt-4")).toBe(DEFAULT_CONTEXT_WINDOW_CLOUD);
|
||||
});
|
||||
|
||||
it("resolves context window default — Ollama's cloud-routed models on a local endpoint", () => {
|
||||
// Regression: glm-5.2:cloud etc. run through the same localhost Ollama daemon as a genuinely
|
||||
// local model, so the base URL alone can't distinguish them — see isSmallLocalModel().
|
||||
expect(resolveContextWindowDefault("http://localhost:11434/v1", "glm-5.2:cloud")).toBe(DEFAULT_CONTEXT_WINDOW_CLOUD);
|
||||
expect(resolveContextWindowDefault("http://localhost:11434/v1", "qwen3.5:397b-cloud")).toBe(DEFAULT_CONTEXT_WINDOW_CLOUD);
|
||||
it("resolves context window default", () => {
|
||||
expect(resolveContextWindowDefault()).toBe(DEFAULT_CONTEXT_WINDOW);
|
||||
});
|
||||
|
||||
it("resolves max iterations default", () => {
|
||||
expect(resolveMaxIterations()).toBe(DEFAULT_MAX_ITERATIONS);
|
||||
});
|
||||
|
||||
it("resolves max output tokens default", () => {
|
||||
expect(resolveMaxOutputTokens()).toBe(DEFAULT_MAX_OUTPUT_TOKENS);
|
||||
});
|
||||
|
||||
it("reads max output tokens from env", () => {
|
||||
process.env.LOCODE_MAX_OUTPUT_TOKENS = "16384";
|
||||
expect(resolveMaxOutputTokens()).toBe(16_384);
|
||||
});
|
||||
|
||||
it("reads max output tokens from stored config", () => {
|
||||
saveStoredConfig({ maxOutputTokens: 4096 });
|
||||
expect(resolveMaxOutputTokens()).toBe(4096);
|
||||
});
|
||||
|
||||
it("rejects out-of-range max output tokens", () => {
|
||||
saveStoredConfig({ maxOutputTokens: 100 });
|
||||
expect(resolveMaxOutputTokens()).toBe(DEFAULT_MAX_OUTPUT_TOKENS);
|
||||
process.env.LOCODE_MAX_OUTPUT_TOKENS = "5000000";
|
||||
expect(resolveMaxOutputTokens()).toBe(DEFAULT_MAX_OUTPUT_TOKENS);
|
||||
});
|
||||
|
||||
it("resolves request timeout default", () => {
|
||||
expect(resolveRequestTimeoutMs()).toBe(DEFAULT_REQUEST_TIMEOUT_MS);
|
||||
});
|
||||
@@ -114,6 +81,52 @@ describe("config resolution", () => {
|
||||
expect(resolveRequestTimeoutMs()).toBe(DEFAULT_REQUEST_TIMEOUT_MS);
|
||||
});
|
||||
|
||||
it("resolves sub-agent max iterations default", () => {
|
||||
expect(resolveSubagentMaxIterations()).toBe(DEFAULT_SUBAGENT_MAX_ITERATIONS);
|
||||
});
|
||||
|
||||
it("reads sub-agent max iterations from env", () => {
|
||||
process.env.LOCODE_SUBAGENT_MAX_ITERATIONS = "30";
|
||||
expect(resolveSubagentMaxIterations()).toBe(30);
|
||||
});
|
||||
|
||||
it("reads sub-agent max iterations from stored config", () => {
|
||||
saveStoredConfig({ subagentMaxIterations: 25 });
|
||||
expect(resolveSubagentMaxIterations()).toBe(25);
|
||||
});
|
||||
|
||||
it("rejects out-of-range sub-agent max iterations", () => {
|
||||
process.env.LOCODE_SUBAGENT_MAX_ITERATIONS = "0"; // below floor
|
||||
expect(resolveSubagentMaxIterations()).toBe(DEFAULT_SUBAGENT_MAX_ITERATIONS);
|
||||
process.env.LOCODE_SUBAGENT_MAX_ITERATIONS = "5000"; // above cap
|
||||
expect(resolveSubagentMaxIterations()).toBe(DEFAULT_SUBAGENT_MAX_ITERATIONS);
|
||||
saveStoredConfig({ subagentMaxIterations: 0 }); // stored below floor
|
||||
expect(resolveSubagentMaxIterations()).toBe(DEFAULT_SUBAGENT_MAX_ITERATIONS);
|
||||
});
|
||||
|
||||
it("resolves sub-agent max depth default", () => {
|
||||
expect(resolveSubagentMaxDepth()).toBe(DEFAULT_SUBAGENT_MAX_DEPTH);
|
||||
});
|
||||
|
||||
it("reads sub-agent max depth from env", () => {
|
||||
process.env.LOCODE_SUBAGENT_MAX_DEPTH = "4";
|
||||
expect(resolveSubagentMaxDepth()).toBe(4);
|
||||
});
|
||||
|
||||
it("reads sub-agent max depth from stored config", () => {
|
||||
saveStoredConfig({ subagentMaxDepth: 3 });
|
||||
expect(resolveSubagentMaxDepth()).toBe(3);
|
||||
});
|
||||
|
||||
it("rejects out-of-range sub-agent max depth", () => {
|
||||
process.env.LOCODE_SUBAGENT_MAX_DEPTH = "-1"; // below floor
|
||||
expect(resolveSubagentMaxDepth()).toBe(DEFAULT_SUBAGENT_MAX_DEPTH);
|
||||
process.env.LOCODE_SUBAGENT_MAX_DEPTH = "99"; // above cap
|
||||
expect(resolveSubagentMaxDepth()).toBe(DEFAULT_SUBAGENT_MAX_DEPTH);
|
||||
saveStoredConfig({ subagentMaxDepth: -1 }); // stored below floor
|
||||
expect(resolveSubagentMaxDepth()).toBe(DEFAULT_SUBAGENT_MAX_DEPTH);
|
||||
});
|
||||
|
||||
it("resolves sub-agent timeout default", () => {
|
||||
expect(resolveSubagentTimeoutMs()).toBe(DEFAULT_SUBAGENT_TIMEOUT_MS);
|
||||
});
|
||||
@@ -136,25 +149,4 @@ describe("config resolution", () => {
|
||||
saveStoredConfig({ subagentTimeoutMs: 500 }); // stored below floor
|
||||
expect(resolveSubagentTimeoutMs()).toBe(DEFAULT_SUBAGENT_TIMEOUT_MS);
|
||||
});
|
||||
|
||||
it("resolves max retries default", () => {
|
||||
expect(resolveMaxRetries()).toBe(DEFAULT_MAX_RETRIES);
|
||||
});
|
||||
|
||||
it("reads max retries from env", () => {
|
||||
process.env.LOCODE_MAX_RETRIES = "3";
|
||||
expect(resolveMaxRetries()).toBe(3);
|
||||
});
|
||||
|
||||
it("reads max retries from stored config", () => {
|
||||
saveStoredConfig({ maxRetries: 5 });
|
||||
expect(resolveMaxRetries()).toBe(5);
|
||||
});
|
||||
|
||||
it("rejects out-of-range max retries", () => {
|
||||
saveStoredConfig({ maxRetries: 11 });
|
||||
expect(resolveMaxRetries()).toBe(DEFAULT_MAX_RETRIES);
|
||||
process.env.LOCODE_MAX_RETRIES = "-1";
|
||||
expect(resolveMaxRetries()).toBe(DEFAULT_MAX_RETRIES);
|
||||
});
|
||||
});
|
||||
+32
-46
@@ -1,15 +1,13 @@
|
||||
import {
|
||||
DEFAULT_AUTO_COMPACT_THRESHOLD,
|
||||
DEFAULT_CONTEXT_WINDOW_LOCAL,
|
||||
DEFAULT_CONTEXT_WINDOW_CLOUD,
|
||||
DEFAULT_CONTEXT_WINDOW,
|
||||
DEFAULT_MAX_ITERATIONS,
|
||||
DEFAULT_MAX_OUTPUT_TOKENS,
|
||||
DEFAULT_MAX_RETRIES,
|
||||
DEFAULT_REQUEST_TIMEOUT_MS,
|
||||
DEFAULT_SUBAGENT_MAX_DEPTH,
|
||||
DEFAULT_SUBAGENT_MAX_ITERATIONS,
|
||||
DEFAULT_SUBAGENT_TIMEOUT_MS,
|
||||
KNOWN_BACKENDS,
|
||||
type BackendName,
|
||||
isSmallLocalModel,
|
||||
} from "./defaults.js";
|
||||
import { loadStoredConfig } from "./store.js";
|
||||
|
||||
@@ -51,28 +49,13 @@ export function resolveModel(cliModel?: string): string | undefined {
|
||||
return cliModel ?? process.env.LOCODE_MODEL ?? stored.model;
|
||||
}
|
||||
|
||||
/** The fallback context window size to use when it can't be auto-detected from the backend. Picks a
|
||||
* small-local or cloud-scale default based on isSmallLocalModel() — the backend host alone isn't
|
||||
* enough, since Ollama's cloud-routed models (e.g. "glm-5.2:cloud") share a local host with
|
||||
* genuinely local ones. */
|
||||
export function resolveContextWindowDefault(baseURL: string, model: string): number {
|
||||
/** The fallback context window size to use when it can't be auto-detected from the backend. */
|
||||
export function resolveContextWindowDefault(): number {
|
||||
const stored = loadStoredConfig();
|
||||
const envValue = Number(process.env.LOCODE_CONTEXT_WINDOW);
|
||||
if (Number.isFinite(envValue) && envValue > 0) return envValue;
|
||||
if (typeof stored.contextWindow === "number" && stored.contextWindow > 0) return stored.contextWindow;
|
||||
return isSmallLocalModel(baseURL, model) ? DEFAULT_CONTEXT_WINDOW_LOCAL : DEFAULT_CONTEXT_WINDOW_CLOUD;
|
||||
}
|
||||
|
||||
/** Ceiling on a single response's max_tokens (see DEFAULT_MAX_OUTPUT_TOKENS), independent of the
|
||||
* context window. Bounded to 256–1,000,000 to reject pathological values. */
|
||||
export function resolveMaxOutputTokens(): number {
|
||||
const stored = loadStoredConfig();
|
||||
const envValue = Number(process.env.LOCODE_MAX_OUTPUT_TOKENS);
|
||||
if (Number.isFinite(envValue) && envValue >= 256 && envValue <= 1_000_000) return envValue;
|
||||
if (typeof stored.maxOutputTokens === "number" && stored.maxOutputTokens >= 256 && stored.maxOutputTokens <= 1_000_000) {
|
||||
return stored.maxOutputTokens;
|
||||
}
|
||||
return DEFAULT_MAX_OUTPUT_TOKENS;
|
||||
return DEFAULT_CONTEXT_WINDOW;
|
||||
}
|
||||
|
||||
/** Max tool calls allowed per turn before locode gives up. */
|
||||
@@ -84,6 +67,31 @@ export function resolveMaxIterations(): number {
|
||||
return DEFAULT_MAX_ITERATIONS;
|
||||
}
|
||||
|
||||
/** Max tool calls in a single sub-agent turn (see DEFAULT_SUBAGENT_MAX_ITERATIONS). Bounded to
|
||||
* 1–1000 to reject pathological values. */
|
||||
export function resolveSubagentMaxIterations(): number {
|
||||
const stored = loadStoredConfig();
|
||||
const envValue = Number(process.env.LOCODE_SUBAGENT_MAX_ITERATIONS);
|
||||
if (Number.isFinite(envValue) && envValue >= 1 && envValue <= 1000) return envValue;
|
||||
if (typeof stored.subagentMaxIterations === "number" && stored.subagentMaxIterations >= 1 && stored.subagentMaxIterations <= 1000) {
|
||||
return stored.subagentMaxIterations;
|
||||
}
|
||||
return DEFAULT_SUBAGENT_MAX_ITERATIONS;
|
||||
}
|
||||
|
||||
/** Max nesting depth for sub-agents (see DEFAULT_SUBAGENT_MAX_DEPTH). Bounded to 0–10: 0 lets the
|
||||
* main session still delegate (depth 1) but blocks that delegate from delegating further; values
|
||||
* above ~10 serve no real purpose and just invite runaway recursion. */
|
||||
export function resolveSubagentMaxDepth(): number {
|
||||
const stored = loadStoredConfig();
|
||||
const envValue = Number(process.env.LOCODE_SUBAGENT_MAX_DEPTH);
|
||||
if (Number.isFinite(envValue) && envValue >= 0 && envValue <= 10) return envValue;
|
||||
if (typeof stored.subagentMaxDepth === "number" && stored.subagentMaxDepth >= 0 && stored.subagentMaxDepth <= 10) {
|
||||
return stored.subagentMaxDepth;
|
||||
}
|
||||
return DEFAULT_SUBAGENT_MAX_DEPTH;
|
||||
}
|
||||
|
||||
/** Wall-clock budget for a single sub-agent turn (see DEFAULT_SUBAGENT_TIMEOUT_MS). Bounded to
|
||||
* 1s–1h to reject pathological values. */
|
||||
export function resolveSubagentTimeoutMs(): number {
|
||||
@@ -107,22 +115,8 @@ export function resolveAutoCompactThreshold(): number {
|
||||
return DEFAULT_AUTO_COMPACT_THRESHOLD;
|
||||
}
|
||||
|
||||
/** Max retry attempts the OpenAI SDK makes on transient failures (connection errors, 429, 5xx)
|
||||
* with exponential backoff. Bounded to 0–10 to reject pathological values. 0 = fail immediately,
|
||||
* matching locode's old behavior of never retrying (a slow local backend usually means the model
|
||||
* is genuinely stuck, not a transient blip — but some setups have occasional connection drops). */
|
||||
export function resolveMaxRetries(): number {
|
||||
const stored = loadStoredConfig();
|
||||
const envValue = Number(process.env.LOCODE_MAX_RETRIES);
|
||||
if (Number.isFinite(envValue) && envValue >= 0 && envValue <= 10) return envValue;
|
||||
if (typeof stored.maxRetries === "number" && stored.maxRetries >= 0 && stored.maxRetries <= 10) {
|
||||
return stored.maxRetries;
|
||||
}
|
||||
return DEFAULT_MAX_RETRIES;
|
||||
}
|
||||
|
||||
/** Milliseconds to wait on a single chat completion request before giving up (see backend/client.ts
|
||||
* for why locode defaults to no retries). Bounded to 10s–30min to reject pathological values. */
|
||||
* for why locode doesn't retry on top of this). Bounded to 10s–30min to reject pathological values. */
|
||||
export function resolveRequestTimeoutMs(): number {
|
||||
const stored = loadStoredConfig();
|
||||
const envValue = Number(process.env.LOCODE_REQUEST_TIMEOUT_MS);
|
||||
@@ -132,11 +126,3 @@ export function resolveRequestTimeoutMs(): number {
|
||||
}
|
||||
return DEFAULT_REQUEST_TIMEOUT_MS;
|
||||
}
|
||||
|
||||
/** User-configured LSP server overrides/additions (see StoredConfig.lspServers). An empty object
|
||||
* means "use the built-in language→server mappings only". Validated loosely: entries without a
|
||||
* command are dropped by configureLanguageSpecs, so we just pass them through. */
|
||||
export function resolveLspServers(): Record<string, { command: string; args?: string[]; extensions?: string[] }> {
|
||||
const stored = loadStoredConfig();
|
||||
return stored.lspServers ?? {};
|
||||
}
|
||||
|
||||
@@ -1,61 +0,0 @@
|
||||
import { describe, expect, it } from "vitest";
|
||||
import { DEFAULT_CONTEXT_WINDOW_CLOUD, DEFAULT_CONTEXT_WINDOW_LOCAL, isCloudRoutedModelName, isLocalBackendURL, isSmallLocalModel } from "./defaults.js";
|
||||
|
||||
describe("isLocalBackendURL", () => {
|
||||
it("recognizes localhost, 127.0.0.1, and ::1", () => {
|
||||
expect(isLocalBackendURL("http://localhost:11434/v1")).toBe(true);
|
||||
expect(isLocalBackendURL("http://127.0.0.1:1234/v1")).toBe(true);
|
||||
expect(isLocalBackendURL("http://[::1]:11434/v1")).toBe(true);
|
||||
});
|
||||
|
||||
it("rejects a remote host", () => {
|
||||
expect(isLocalBackendURL("https://api.openai.com/v1")).toBe(false);
|
||||
});
|
||||
|
||||
it("returns false for an unparseable URL instead of throwing", () => {
|
||||
expect(isLocalBackendURL("not a url")).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe("isCloudRoutedModelName", () => {
|
||||
it("recognizes Ollama's cloud tag conventions", () => {
|
||||
expect(isCloudRoutedModelName("glm-5.2:cloud")).toBe(true);
|
||||
expect(isCloudRoutedModelName("qwen3.5:397b-cloud")).toBe(true);
|
||||
expect(isCloudRoutedModelName("gpt-oss:120b-cloud")).toBe(true);
|
||||
});
|
||||
|
||||
it("does not match a genuinely local model", () => {
|
||||
expect(isCloudRoutedModelName("llama3.2:latest")).toBe(false);
|
||||
expect(isCloudRoutedModelName("gemma3:4b")).toBe(false);
|
||||
expect(isCloudRoutedModelName("phi4:14b")).toBe(false);
|
||||
});
|
||||
|
||||
it("does not false-positive on 'cloud' appearing outside the tag segment", () => {
|
||||
expect(isCloudRoutedModelName("cloudmodel:latest")).toBe(false);
|
||||
expect(isCloudRoutedModelName("cloud")).toBe(false); // no colon at all
|
||||
});
|
||||
});
|
||||
|
||||
describe("isSmallLocalModel", () => {
|
||||
it("is true for a genuinely local model on a local endpoint", () => {
|
||||
expect(isSmallLocalModel("http://localhost:11434/v1", "llama3.2:latest")).toBe(true);
|
||||
});
|
||||
|
||||
it("is false for a cloud-routed model even on a local endpoint", () => {
|
||||
expect(isSmallLocalModel("http://localhost:11434/v1", "glm-5.2:cloud")).toBe(false);
|
||||
});
|
||||
|
||||
it("is false for a remote backend regardless of model name", () => {
|
||||
expect(isSmallLocalModel("https://api.openai.com/v1", "gpt-4")).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe("DEFAULT_CONTEXT_WINDOW_LOCAL / CLOUD", () => {
|
||||
it("LOCAL is the small-local fallback (8 192)", () => {
|
||||
expect(DEFAULT_CONTEXT_WINDOW_LOCAL).toBe(8192);
|
||||
});
|
||||
|
||||
it("CLOUD is the cloud/large fallback (131 072)", () => {
|
||||
expect(DEFAULT_CONTEXT_WINDOW_CLOUD).toBe(131072);
|
||||
});
|
||||
});
|
||||
+29
-67
@@ -8,82 +8,44 @@ export const KNOWN_BACKENDS = {
|
||||
|
||||
export type BackendName = keyof typeof KNOWN_BACKENDS;
|
||||
|
||||
/** Returns true if the given base URL looks like a local backend (Ollama or LM Studio on localhost).
|
||||
* On its own this is NOT enough to tell a small local model from a large one — see
|
||||
* isSmallLocalModel() below, which is what callers should actually use. */
|
||||
export function isLocalBackendURL(baseURL: string): boolean {
|
||||
try {
|
||||
const url = new URL(baseURL);
|
||||
// WHATWG URL keeps the brackets on a literal IPv6 host in .hostname (e.g. "[::1]", not "::1") —
|
||||
// https://url.spec.whatwg.org/#concept-host-serializer.
|
||||
return url.hostname === "localhost" || url.hostname === "127.0.0.1" || url.hostname === "[::1]";
|
||||
} catch {
|
||||
return false;
|
||||
}
|
||||
}
|
||||
/** Used when the context window can't be auto-detected from the backend (see backend/contextWindow.ts)
|
||||
* and the user hasn't configured one — a conservative size common among smaller local models. */
|
||||
export const DEFAULT_CONTEXT_WINDOW = 8192;
|
||||
|
||||
/** Ollama's cloud-hosted models are proxied through the same local daemon as a genuinely local
|
||||
* model — same base URL, same host — so isLocalBackendURL() alone can't tell them apart. Ollama
|
||||
* names them with a "cloud" tag segment instead, e.g. "glm-5.2:cloud" or "qwen3.5:397b-cloud". */
|
||||
export function isCloudRoutedModelName(model: string): boolean {
|
||||
const tag = model.split(":")[1] ?? "";
|
||||
return tag === "cloud" || tag.endsWith("-cloud");
|
||||
}
|
||||
/** Max tool calls per *round* within a turn. Local models often issue one tool call per round, so
|
||||
* a multi-file edit + verify sequence can easily run past 50; 100 gives real tasks room to breathe
|
||||
* while still bounding a genuinely stuck model per round. runTurn automatically chains up to
|
||||
* MAX_CONTINUATION_ROUNDS rounds before pausing, so the effective per-turn budget is much larger. */
|
||||
export const DEFAULT_MAX_ITERATIONS = 100;
|
||||
|
||||
/** Whether locode should treat this model as a small/less-capable local model — extra system-prompt
|
||||
* guidance about unreliable tool calls and limited output, plus a conservative context-window
|
||||
* fallback — rather than a large, reliable one. True only when the backend is local AND the model
|
||||
* isn't one of Ollama's cloud-routed models under that same local endpoint; a backend hosted
|
||||
* elsewhere (a real remote/cloud API) is never treated as "small local" regardless of model name. */
|
||||
export function isSmallLocalModel(baseURL: string, model: string): boolean {
|
||||
return isLocalBackendURL(baseURL) && !isCloudRoutedModelName(model);
|
||||
}
|
||||
/** Number of internal rounds runTurn will chain automatically when a single round exhausts its
|
||||
* step budget without producing a final answer. This matches Claude Code's behavior of continuing
|
||||
* a long tool chain rather than stopping after every N steps and asking the user to continue.
|
||||
* Effective per-turn budget = DEFAULT_MAX_ITERATIONS * MAX_CONTINUATION_ROUNDS. */
|
||||
export const MAX_CONTINUATION_ROUNDS = 5;
|
||||
|
||||
/** Fallback context window when auto-detection fails and the user hasn't configured one.
|
||||
* Small local models (see isSmallLocalModel()) typically have 8k–32k context, so 8192 is a safe
|
||||
* conservative default. Everything else (cloud APIs, and Ollama's own cloud-routed models) typically
|
||||
* has 128k–1M+ context, so 131072 (128K) avoids severely underutilizing them. */
|
||||
export const DEFAULT_CONTEXT_WINDOW_LOCAL = 8192;
|
||||
export const DEFAULT_CONTEXT_WINDOW_CLOUD = 131072;
|
||||
/** Max tool calls in a single sub-agent turn. Sub-agents run headless (one round, no
|
||||
* auto-continue — see runSubAgentTurn) and are meant for one focused task, so this is intentionally
|
||||
* smaller than the parent's per-round budget (DEFAULT_MAX_ITERATIONS): a runaway sub-agent fails
|
||||
* fast and tells the parent to split the task rather than silently burning a large black-box
|
||||
* budget. Configurable via `LOCODE_SUBAGENT_MAX_ITERATIONS` / `locode config set subagentMaxIterations`. */
|
||||
export const DEFAULT_SUBAGENT_MAX_ITERATIONS = 50;
|
||||
|
||||
/** Ceiling on a single response's `max_tokens`, independent of the model's context window. Most
|
||||
* backends cap how much a single completion can generate well below the total context window they
|
||||
* advertise (e.g. Ollama's glm-5.1:cloud / glm-5.2:cloud report a 1,000,000-token context window
|
||||
* but error on a request above their real 131,072-token output cap: "max_tokens (500000) exceeds
|
||||
* model's maximum output tokens (131072)") — resolveMaxTokens (agent/loop.ts) used to request up to
|
||||
* the whole remaining window, which such backends rejected outright as a context/length error even
|
||||
* on the very first turn. 131072 (128K) matches that verified real ceiling exactly, so it's safe to
|
||||
* use as the default without erroring on the very first turn — going straight to a backend's real
|
||||
* cap only became reasonable once detectRepetitionLoop (agent/loop.ts) existed to abort a model
|
||||
* that gets stuck generating instead of relying on this value alone as the safety valve. Raise it
|
||||
* further via `locode config set maxOutputTokens` for backends known to allow more; lower it for
|
||||
* ones with a smaller real ceiling. */
|
||||
export const DEFAULT_MAX_OUTPUT_TOKENS = 131_072;
|
||||
|
||||
/** Max model requests per turn before locode pauses rather than looping forever. Each iteration
|
||||
* is one model generation request (one tool-call round-trip), and local models commonly issue a
|
||||
* single tool call per request — so a real multi-file task (read several files, edit each, grep
|
||||
* to verify, re-read) easily needs 40–60 requests. 50 was too tight and caused frequent
|
||||
* "Paused after 50 steps" soft-stops on legitimate work; 100 still caused frequent pauses on
|
||||
* larger tasks. 300 gives real tasks ample room to finish while still bounding a genuinely stuck
|
||||
* model. Hitting the cap is a soft pause, not a failure (the work so far is intact — send another
|
||||
* message to resume). Configurable via `maxIterations`, e.g. `locode config set maxIterations 500`
|
||||
* for very large batch jobs. */
|
||||
export const DEFAULT_MAX_ITERATIONS = 300;
|
||||
/** Max nesting depth for sub-agents. Depth 0 is the main session; a sub-agent it spawns is depth
|
||||
* 1, a sub-agent that one spawns is depth 2, and so on. A sub-agent at the cap can't spawn further
|
||||
* sub-agents (the toolset excludes `agent` for it, with an explicit depth-check backstop in
|
||||
* agent/loop.ts). The default of 2 allows one level of delegation plus a focused sub-task under
|
||||
* that, while keeping recursion shallow enough that a misbehaving model can't fan out
|
||||
* uncontrollably. Configurable via `LOCODE_SUBAGENT_MAX_DEPTH` / `locode config set subagentMaxDepth`. */
|
||||
export const DEFAULT_SUBAGENT_MAX_DEPTH = 2;
|
||||
|
||||
/** Fraction of the context window at which locode automatically summarizes the conversation.
|
||||
* User-configurable via `locode config set autoCompactThreshold`. */
|
||||
export const DEFAULT_AUTO_COMPACT_THRESHOLD = 0.85;
|
||||
|
||||
/** How long to wait on a single chat completion request before giving up. The OpenAI SDK retries
|
||||
* transient failures (connection errors, 429, 5xx) up to `maxRetries` times with exponential
|
||||
* backoff before surfacing the error; set to 0 to fail immediately like older locode versions.
|
||||
* Raise this via `maxRetries` if your backend has occasional transient blips. See backend/client.ts. */
|
||||
export const DEFAULT_MAX_RETRIES = 0;
|
||||
|
||||
/** How long to wait on a single chat completion request before giving up. Raise this via
|
||||
* `requestTimeoutMs` if your backend queues requests behind a concurrency limit (e.g. Ollama's
|
||||
* `OLLAMA_NUM_PARALLEL`) rather than serving them immediately. */
|
||||
/** How long to wait on a single chat completion request before giving up (no retries — see
|
||||
* backend/client.ts). Raise this via `requestTimeoutMs` if your backend queues requests behind a
|
||||
* concurrency limit (e.g. Ollama's `OLLAMA_NUM_PARALLEL`) rather than serving them immediately. */
|
||||
export const DEFAULT_REQUEST_TIMEOUT_MS = 180_000;
|
||||
|
||||
/** Wall-clock budget for a single sub-agent turn. Sub-agents make their own sequence of model
|
||||
|
||||
+19
-31
@@ -1,6 +1,7 @@
|
||||
import envPaths from "env-paths";
|
||||
import { existsSync, mkdirSync, readFileSync, writeFileSync } from "node:fs";
|
||||
import path from "node:path";
|
||||
import type { PermissionRule } from "../permissions/types.js";
|
||||
|
||||
export interface StoredConfig {
|
||||
backend?: string;
|
||||
@@ -10,24 +11,20 @@ export interface StoredConfig {
|
||||
contextWindow?: number;
|
||||
/** Max tool calls allowed per turn before locode gives up rather than looping forever. */
|
||||
maxIterations?: number;
|
||||
/** Max tool calls in a single sub-agent turn (smaller than maxIterations — sub-agents are bounded
|
||||
* to one focused task with no auto-continue). */
|
||||
subagentMaxIterations?: number;
|
||||
/** Max nesting depth for sub-agents (0 = main session, 1 = first delegation, …). */
|
||||
subagentMaxDepth?: number;
|
||||
/** Fraction of the context window (0.0–1.0) at which locode auto-compacts the conversation. */
|
||||
autoCompactThreshold?: number;
|
||||
/** Milliseconds to wait on a single chat completion request before giving up. */
|
||||
requestTimeoutMs?: number;
|
||||
/** Milliseconds of wall-clock budget for a single sub-agent turn. */
|
||||
subagentTimeoutMs?: number;
|
||||
/** Ceiling on a single response's max_tokens, independent of contextWindow. */
|
||||
maxOutputTokens?: number;
|
||||
/** Max retry attempts the OpenAI SDK makes on transient failures (connection errors, 429, 5xx)
|
||||
* before surfacing the error. 0 = fail immediately (old behavior); the SDK uses exponential
|
||||
* backoff between attempts. */
|
||||
maxRetries?: number;
|
||||
/** User-defined LSP server overrides/additions, keyed by language id (e.g. "java", "ruby",
|
||||
* "typescript"). Each entry is { command, args?, extensions? }. An entry for a built-in id
|
||||
* overrides its command/args; an entry with `extensions` also rewrites which file extensions
|
||||
* route to that language. Entries for new ids add support for languages locode doesn't ship a
|
||||
* server for. See `locode config set lspServers` (JSON value). */
|
||||
lspServers?: Record<string, { command: string; args?: string[]; extensions?: string[] }>;
|
||||
/** User-level permission rules (auto-approve/deny tool calls without prompting). Merged with
|
||||
* project-level rules from .locode/settings.json; deny wins across layers. See PermissionRule. */
|
||||
permissionRules?: PermissionRule[];
|
||||
}
|
||||
|
||||
const paths = envPaths("locode", { suffix: "" });
|
||||
@@ -41,39 +38,30 @@ export function configFilePath(): string {
|
||||
return configFileOverride ?? configFile;
|
||||
}
|
||||
|
||||
/** Directory holding the user-level config files (config.json, hooks.json, mcp.json, plugins.json,
|
||||
* and the freeform memory.md). Derived from configFilePath() so the test override
|
||||
* (_setConfigFilePathForTest) redirects this too. */
|
||||
export function configDirPath(): string {
|
||||
return path.dirname(configFilePath());
|
||||
}
|
||||
|
||||
/** @internal For tests only — redirect config persistence to `path` (pass undefined to reset). */
|
||||
export function _setConfigFilePathForTest(p: string | undefined): void {
|
||||
configFileOverride = p;
|
||||
invalidateConfigCache();
|
||||
}
|
||||
|
||||
// In-memory cache so every resolve*() call in the same process doesn't re-read and re-parse
|
||||
// the same small JSON file (8–10 calls during startup alone). Invalidated by saveStoredConfig
|
||||
// and by tests that change the config path.
|
||||
let cachedConfig: StoredConfig | undefined;
|
||||
|
||||
export function loadStoredConfig(): StoredConfig {
|
||||
if (cachedConfig !== undefined) return cachedConfig;
|
||||
const file = configFilePath();
|
||||
if (!existsSync(file)) { cachedConfig = {}; return cachedConfig; }
|
||||
if (!existsSync(file)) return {};
|
||||
try {
|
||||
cachedConfig = JSON.parse(readFileSync(file, "utf-8")) as StoredConfig;
|
||||
return cachedConfig;
|
||||
return JSON.parse(readFileSync(file, "utf-8")) as StoredConfig;
|
||||
} catch {
|
||||
cachedConfig = {};
|
||||
return cachedConfig;
|
||||
return {};
|
||||
}
|
||||
}
|
||||
|
||||
/** Invalidate the in-memory cache — called by saveStoredConfig and by tests that change
|
||||
* the config file path. */
|
||||
export function invalidateConfigCache(): void {
|
||||
cachedConfig = undefined;
|
||||
}
|
||||
|
||||
export function saveStoredConfig(cfg: StoredConfig): void {
|
||||
const file = configFilePath();
|
||||
mkdirSync(path.dirname(file), { recursive: true });
|
||||
writeFileSync(file, JSON.stringify(cfg, null, 2));
|
||||
cachedConfig = cfg;
|
||||
}
|
||||
@@ -91,22 +91,11 @@ describe("runHooksForEvent", () => {
|
||||
expect(result.additionalContext).toBeUndefined();
|
||||
});
|
||||
|
||||
it("injects a prompt hook's message as additional context", async () => {
|
||||
it("warns for unsupported prompt hooks", async () => {
|
||||
loadMergedHooks.loadMergedHooks.mockReturnValueOnce({
|
||||
SessionStart: [{ hooks: [{ type: "prompt", message: "Remember to check the changelog." }] }],
|
||||
SessionStart: [{ hooks: [{ type: "prompt", message: "ok?" } as any] }],
|
||||
});
|
||||
const result = await runHooksForEvent("SessionStart", ctx, {});
|
||||
expect(result.additionalContext).toBe("Remember to check the changelog.");
|
||||
expect(result.warnings).toEqual([]);
|
||||
});
|
||||
|
||||
it("combines a prompt hook's message with a command hook's stdout", async () => {
|
||||
loadMergedHooks.loadMergedHooks.mockReturnValueOnce({
|
||||
SessionStart: [{ hooks: [{ type: "prompt", message: "prompt message" }, { type: "command", command: "echo cmd" }] }],
|
||||
});
|
||||
execa.execa.mockResolvedValueOnce({ exitCode: 0, stdout: "cmd output", stderr: "", timedOut: false });
|
||||
const result = await runHooksForEvent("SessionStart", ctx, {});
|
||||
expect(result.additionalContext).toContain("prompt message");
|
||||
expect(result.additionalContext).toContain("cmd output");
|
||||
expect(result.warnings).toContain("1 prompt hook(s) skipped (not yet implemented).");
|
||||
});
|
||||
});
|
||||
|
||||
+11
-17
@@ -2,7 +2,7 @@ import { execa } from "execa";
|
||||
import path from "node:path";
|
||||
import { loadMergedHooks } from "./config.js";
|
||||
import { matcherMatches } from "./matcher.js";
|
||||
import type { Hook, HookCommand, HookEventName, HookHttp, HookPrompt } from "./types.js";
|
||||
import type { Hook, HookCommand, HookEventName, HookHttp } from "./types.js";
|
||||
|
||||
const DEFAULT_TIMEOUT_SECONDS = 30;
|
||||
|
||||
@@ -37,20 +37,6 @@ function isHttpHook(hook: Hook): hook is HookHttp {
|
||||
return hook.type === "http";
|
||||
}
|
||||
|
||||
function isPromptHook(hook: Hook): hook is HookPrompt {
|
||||
return hook.type === "prompt";
|
||||
}
|
||||
|
||||
/** A prompt hook has no process/response to run — it's just a static message that always
|
||||
* "succeeds" and folds into additionalContext the same way a command/http hook's plain-text
|
||||
* stdout does (see the outcome-handling loop in runHooksForEvent). Wrapped in a resolved promise
|
||||
* so it can share the same Promise.all as the other hook kinds. */
|
||||
async function runPromptHook(
|
||||
hook: HookPrompt,
|
||||
): Promise<{ exitCode: number; stdout: string; stderr: string; timedOut: boolean; json?: unknown; warning?: string }> {
|
||||
return { exitCode: 0, stdout: hook.message, stderr: "", timedOut: false };
|
||||
}
|
||||
|
||||
function parseOutput(output: string, schema?: "json"): { text?: string; json?: unknown; warning?: string } {
|
||||
if (!schema) return { text: output };
|
||||
if (schema !== "json") return { text: output };
|
||||
@@ -209,18 +195,26 @@ export async function runHooksForEvent(
|
||||
.flatMap((entry) => entry.hooks);
|
||||
if (hooks.length === 0) return EMPTY_RESULT;
|
||||
|
||||
// Prompt hooks are declared in the type system but not implemented in this pass — they need a
|
||||
// blocking UI flow the current runner doesn't have. Skip them with a warning so a config that
|
||||
// includes them doesn't silently do nothing.
|
||||
const skippedPrompts = hooks.filter((h) => h.type === "prompt").length;
|
||||
const runnableHooks = hooks.filter((hook) => hook.type !== "prompt");
|
||||
|
||||
const stdinPayload = JSON.stringify({ hook_event_name: event, session_id: ctx.sessionId, cwd: ctx.cwd, ...payload });
|
||||
|
||||
const outcomes = await Promise.all(
|
||||
hooks.map(async (hook) => {
|
||||
runnableHooks.map(async (hook) => {
|
||||
if (isCommandHook(hook)) return runCommandHook(hook, stdinPayload, ctx);
|
||||
if (isHttpHook(hook)) return runHttpHook(hook, stdinPayload, ctx);
|
||||
if (isPromptHook(hook)) return runPromptHook(hook);
|
||||
return { exitCode: 1, stdout: "", stderr: `Unsupported hook type: ${(hook as Hook).type}`, timedOut: false };
|
||||
}),
|
||||
);
|
||||
|
||||
const result: HookResult = { blocked: false, warnings: [] };
|
||||
if (skippedPrompts > 0) {
|
||||
result.warnings.push(`${skippedPrompts} prompt hook(s) skipped (not yet implemented).`);
|
||||
}
|
||||
for (const outcome of outcomes) {
|
||||
if (outcome.warning) {
|
||||
result.warnings.push(outcome.warning);
|
||||
|
||||
+1
-3
@@ -40,10 +40,8 @@ export interface HookHttp extends HookBase {
|
||||
|
||||
export interface HookPrompt extends HookBase {
|
||||
type: "prompt";
|
||||
/** Injected verbatim as additional context, the same way a command/http hook's stdout is —
|
||||
* see runner.ts. Always "succeeds" (there's no process/response to fail); a prompt hook can't
|
||||
* block an event the way a command hook's exit code 2 can. */
|
||||
message: string;
|
||||
/** Not yet implemented — prompt hooks require a UI blocking flow the current runner doesn't support. */
|
||||
}
|
||||
|
||||
export type Hook = HookCommand | HookHttp | HookPrompt;
|
||||
|
||||
+23
-47
@@ -3,7 +3,6 @@ import { StdioClientTransport } from "@modelcontextprotocol/sdk/client/stdio.js"
|
||||
import { StreamableHTTPClientTransport } from "@modelcontextprotocol/sdk/client/streamableHttp.js";
|
||||
import type { Transport } from "@modelcontextprotocol/sdk/shared/transport.js";
|
||||
import { isHttpServerConfig, type McpServerConfig } from "./types.js";
|
||||
import pkg from "../../package.json" with { type: "json" };
|
||||
|
||||
export interface McpTextContentBlock {
|
||||
type: "text";
|
||||
@@ -55,53 +54,30 @@ export interface ConnectedMcpServer {
|
||||
tools: McpToolInfo[];
|
||||
}
|
||||
|
||||
/** Max connection attempts for an MCP server that fails transiently (process slow to start,
|
||||
* HTTP 503, etc.). A stdio server whose command genuinely doesn't exist fails immediately every
|
||||
* time, so retries only help the transient case — kept small to avoid stalling startup. */
|
||||
const MCP_CONNECT_MAX_ATTEMPTS = 3;
|
||||
const MCP_CONNECT_BASE_DELAY_MS = 500;
|
||||
|
||||
async function sleep(ms: number, signal?: AbortSignal): Promise<void> {
|
||||
return new Promise((resolve, reject) => {
|
||||
const t = setTimeout(resolve, ms);
|
||||
signal?.addEventListener("abort", () => { clearTimeout(t); reject(new Error("aborted")); }, { once: true });
|
||||
});
|
||||
}
|
||||
|
||||
export async function connectMcpServer(name: string, config: McpServerConfig): Promise<ConnectedMcpServer> {
|
||||
let lastErr: unknown;
|
||||
for (let attempt = 1; attempt <= MCP_CONNECT_MAX_ATTEMPTS; attempt++) {
|
||||
const transport: Transport = isHttpServerConfig(config)
|
||||
? new StreamableHTTPClientTransport(new URL(config.url), {
|
||||
requestInit: config.headers ? { headers: config.headers } : undefined,
|
||||
})
|
||||
: new StdioClientTransport({
|
||||
command: config.command,
|
||||
args: config.args,
|
||||
env: config.env,
|
||||
// Default is "inherit", which would leak the child's stderr straight into the terminal
|
||||
// and corrupt Ink's alternate-screen UI. Pipe it instead so it's just discarded.
|
||||
stderr: "pipe",
|
||||
});
|
||||
const transport: Transport = isHttpServerConfig(config)
|
||||
? new StreamableHTTPClientTransport(new URL(config.url), {
|
||||
requestInit: config.headers ? { headers: config.headers } : undefined,
|
||||
})
|
||||
: new StdioClientTransport({
|
||||
command: config.command,
|
||||
args: config.args,
|
||||
env: config.env,
|
||||
// Default is "inherit", which would leak the child's stderr straight into the terminal
|
||||
// and corrupt Ink's alternate-screen UI. Pipe it instead so it's just discarded.
|
||||
stderr: "pipe",
|
||||
});
|
||||
|
||||
const client = new Client({ name: "locode", version: pkg.version });
|
||||
try {
|
||||
await client.connect(transport);
|
||||
const { tools } = await client.listTools();
|
||||
return { name, client, transport, tools: tools as McpToolInfo[] };
|
||||
} catch (err) {
|
||||
// listTools failed or connect failed — close the transport so the stdio subprocess / HTTP
|
||||
// connection isn't orphaned. manager.ts catches the rejection as an error status, but
|
||||
// without this the child process keeps running.
|
||||
await transport.close().catch(() => {});
|
||||
lastErr = err;
|
||||
// Retry with exponential backoff for transient failures; the last attempt's error is what
|
||||
// the caller sees. A genuinely broken config (missing binary, bad URL) fails fast every time,
|
||||
// so the retries just add a small delay — acceptable for the rare transient-startup case.
|
||||
if (attempt < MCP_CONNECT_MAX_ATTEMPTS) {
|
||||
await sleep(MCP_CONNECT_BASE_DELAY_MS * Math.pow(2, attempt - 1));
|
||||
}
|
||||
}
|
||||
const client = new Client({ name: "locode", version: "0.3.1" });
|
||||
await client.connect(transport);
|
||||
try {
|
||||
const { tools } = await client.listTools();
|
||||
return { name, client, transport, tools: tools as McpToolInfo[] };
|
||||
} catch (err) {
|
||||
// listTools failed (server connected but never responded to the listing) — close the
|
||||
// transport so the stdio subprocess / HTTP connection isn't orphaned. manager.ts catches
|
||||
// the rejection as an error status, but without this the child process keeps running.
|
||||
await transport.close().catch(() => {});
|
||||
throw err;
|
||||
}
|
||||
throw lastErr;
|
||||
}
|
||||
|
||||
@@ -76,12 +76,3 @@ export async function disconnectAllMcpServers(): Promise<void> {
|
||||
await Promise.allSettled(connections.map((c) => c.transport.close()));
|
||||
connections = [];
|
||||
}
|
||||
|
||||
/** Re-connects to every configured MCP server (e.g. after the user edits `.mcp.json` or restarts a
|
||||
* server process), swapping in the new connections and returning the fresh tool list so the caller
|
||||
* can rebuild the session's toolset. Equivalent to a fresh `connectConfiguredMcpServers` call,
|
||||
* exposed separately so callers can name the intent ("reconnect") without implying the first-run
|
||||
* setup path. */
|
||||
export async function reconnectMcpServers(cwd: string): Promise<ToolDef[]> {
|
||||
return connectConfiguredMcpServers(cwd);
|
||||
}
|
||||
|
||||
@@ -0,0 +1,130 @@
|
||||
import { describe, expect, it } from "vitest";
|
||||
import { PermissionManager } from "./permissionManager.js";
|
||||
import type { PermissionRule } from "./types.js";
|
||||
|
||||
describe("PermissionManager rules", () => {
|
||||
describe("checkRules", () => {
|
||||
it("returns null when no rule matches the tool", () => {
|
||||
const pm = new PermissionManager([{ tool: "bash", allow: true }]);
|
||||
expect(pm.checkRules("edit_file", { command: "ls" })).toBeNull();
|
||||
});
|
||||
|
||||
it("returns allow for a matching allow rule without argPattern", () => {
|
||||
const pm = new PermissionManager([{ tool: "bash", allow: true }]);
|
||||
expect(pm.checkRules("bash", { command: "rm -rf /" })).toBe("allow");
|
||||
});
|
||||
|
||||
it("returns deny for a matching deny rule without argPattern", () => {
|
||||
const pm = new PermissionManager([{ tool: "bash", allow: false }]);
|
||||
expect(pm.checkRules("bash", { command: "ls" })).toBe("deny");
|
||||
});
|
||||
|
||||
it("deny wins over allow regardless of order", () => {
|
||||
const allowFirst: PermissionRule[] = [
|
||||
{ tool: "bash", allow: true },
|
||||
{ tool: "bash", allow: false },
|
||||
];
|
||||
const denyFirst: PermissionRule[] = [
|
||||
{ tool: "bash", allow: false },
|
||||
{ tool: "bash", allow: true },
|
||||
];
|
||||
expect(new PermissionManager(allowFirst).checkRules("bash", {})).toBe("deny");
|
||||
expect(new PermissionManager(denyFirst).checkRules("bash", {})).toBe("deny");
|
||||
});
|
||||
|
||||
it("deny with argPattern only denies matching args, falling through otherwise", () => {
|
||||
const pm = new PermissionManager([{ tool: "bash", argPattern: "rm\\s+-rf", allow: false }]);
|
||||
expect(pm.checkRules("bash", { command: "rm -rf /" })).toBe("deny");
|
||||
// Non-matching args: no rule fires → null.
|
||||
expect(pm.checkRules("bash", { command: "ls" })).toBeNull();
|
||||
});
|
||||
|
||||
it("allow with argPattern auto-approves only matching args", () => {
|
||||
// argPattern is a regex tested against JSON.stringify(args), so for {command:"npm test"} the
|
||||
// serialized text is {"command":"npm test"} — match the substring (no ^ anchor, which would
|
||||
// bind to the leading brace).
|
||||
const pm = new PermissionManager([{ tool: "bash", argPattern: "npm (test|run)", allow: true }]);
|
||||
expect(pm.checkRules("bash", { command: "npm test" })).toBe("allow");
|
||||
expect(pm.checkRules("bash", { command: "npm install" })).toBeNull();
|
||||
});
|
||||
|
||||
it("argPattern is tested against JSON.stringify(args), so nested fields match", () => {
|
||||
const pm = new PermissionManager([{ tool: "edit_file", argPattern: "secret", allow: false }]);
|
||||
expect(pm.checkRules("edit_file", { path: "/safe", old_string: "secret" })).toBe("deny");
|
||||
expect(pm.checkRules("edit_file", { path: "/safe", old_string: "public" })).toBeNull();
|
||||
});
|
||||
|
||||
it("an invalid regex in argPattern is skipped, not thrown", () => {
|
||||
const pm = new PermissionManager([
|
||||
{ tool: "bash", argPattern: "(", allow: false }, // invalid regex
|
||||
{ tool: "bash", allow: true },
|
||||
]);
|
||||
// The invalid deny rule is skipped, so the allow rule fires.
|
||||
expect(pm.checkRules("bash", { command: "ls" })).toBe("allow");
|
||||
});
|
||||
|
||||
it("treats undefined args as an empty string for pattern matching", () => {
|
||||
const pm = new PermissionManager([{ tool: "bash", argPattern: "^$", allow: true }]);
|
||||
expect(pm.checkRules("bash", undefined)).toBe("allow");
|
||||
});
|
||||
});
|
||||
|
||||
describe("isAutoApproved", () => {
|
||||
it("auto-approves when an allow rule matches", () => {
|
||||
const pm = new PermissionManager([{ tool: "bash", argPattern: "npm test", allow: true }]);
|
||||
expect(pm.isAutoApproved("bash", { command: "npm test" })).toBe(true);
|
||||
expect(pm.isAutoApproved("bash", { command: "rm -rf" })).toBe(false);
|
||||
});
|
||||
|
||||
it("does not auto-approve when a deny rule matches (deny handled in the gate, not here)", () => {
|
||||
const pm = new PermissionManager([{ tool: "bash", allow: false }]);
|
||||
expect(pm.isAutoApproved("bash", {})).toBe(false);
|
||||
});
|
||||
|
||||
it("auto-edit mode auto-approves write_file/edit_file only", () => {
|
||||
const pm = new PermissionManager([]);
|
||||
pm.setMode("auto-edit");
|
||||
expect(pm.isAutoApproved("write_file", { path: "x" })).toBe(true);
|
||||
expect(pm.isAutoApproved("edit_file", { path: "x" })).toBe(true);
|
||||
// bash is mutating but not a file-edit tool — still prompts even in auto-edit.
|
||||
expect(pm.isAutoApproved("bash", { command: "ls" })).toBe(false);
|
||||
});
|
||||
|
||||
it("auto-accept mode also only covers file-edit tools, not bash/git", () => {
|
||||
const pm = new PermissionManager([]);
|
||||
pm.setMode("auto-accept");
|
||||
expect(pm.isAutoApproved("write_file", {})).toBe(true);
|
||||
expect(pm.isAutoApproved("bash", {})).toBe(false);
|
||||
});
|
||||
|
||||
it("default mode falls back to the session-allowed list", () => {
|
||||
const pm = new PermissionManager([]);
|
||||
pm.allowForSession("edit_file");
|
||||
expect(pm.isAutoApproved("edit_file")).toBe(true);
|
||||
expect(pm.isAutoApproved("write_file")).toBe(false);
|
||||
});
|
||||
|
||||
it("an allow rule overrides even default mode (no session-allow needed)", () => {
|
||||
const pm = new PermissionManager([{ tool: "bash", allow: true }]);
|
||||
expect(pm.isAutoApproved("bash", {})).toBe(true);
|
||||
});
|
||||
});
|
||||
|
||||
describe("setRules / listRules", () => {
|
||||
it("setRules replaces the active rule set", () => {
|
||||
const pm = new PermissionManager([{ tool: "bash", allow: true }]);
|
||||
pm.setRules([{ tool: "edit_file", allow: false }]);
|
||||
expect(pm.checkRules("bash", {})).toBeNull();
|
||||
expect(pm.checkRules("edit_file", {})).toBe("deny");
|
||||
});
|
||||
|
||||
it("listRules returns a defensive copy", () => {
|
||||
const pm = new PermissionManager([{ tool: "bash", allow: true }]);
|
||||
const out = pm.listRules();
|
||||
expect(out).toEqual([{ tool: "bash", allow: true }]);
|
||||
out.push({ tool: "edit_file", allow: false });
|
||||
// Mutating the returned array must not affect the manager.
|
||||
expect(pm.listRules()).toEqual([{ tool: "bash", allow: true }]);
|
||||
});
|
||||
});
|
||||
});
|
||||
@@ -1,9 +1,14 @@
|
||||
import type { PermissionMode } from "./types.js";
|
||||
import type { PermissionMode, PermissionRule } from "./types.js";
|
||||
import { AUTO_EDIT_TOOLS } from "./types.js";
|
||||
|
||||
export class PermissionManager {
|
||||
private allowedForSession = new Set<string>();
|
||||
private mode: PermissionMode = "default";
|
||||
private rules: PermissionRule[] = [];
|
||||
|
||||
constructor(rules: PermissionRule[] = []) {
|
||||
this.rules = rules;
|
||||
}
|
||||
|
||||
getMode(): PermissionMode {
|
||||
return this.mode;
|
||||
@@ -13,14 +18,48 @@ export class PermissionManager {
|
||||
this.mode = mode;
|
||||
}
|
||||
|
||||
/** Replace the active permission-rule set (used after loading/merging user + project rules). */
|
||||
setRules(rules: PermissionRule[]): void {
|
||||
this.rules = rules;
|
||||
}
|
||||
|
||||
listRules(): PermissionRule[] {
|
||||
return this.rules.map((r) => ({ ...r }));
|
||||
}
|
||||
|
||||
/** Evaluate the persistent rules against a tool call. Returns `"deny"` if any deny rule matches
|
||||
* (deny wins regardless of order or layer), `"allow"` if an allow rule matches and no deny does,
|
||||
* or `null` when no rule matches (the caller falls through to mode/session logic and prompting).
|
||||
* `argPattern` (when present) is a regex tested against JSON.stringify(args); an invalid regex is
|
||||
* skipped rather than crashing the turn. */
|
||||
checkRules(toolName: string, args: unknown): "deny" | "allow" | null {
|
||||
const serialized = args === undefined ? "" : JSON.stringify(args);
|
||||
let allowMatch = false;
|
||||
for (const r of this.rules) {
|
||||
if (r.tool !== toolName) continue;
|
||||
if (r.argPattern !== undefined) {
|
||||
try {
|
||||
if (!new RegExp(r.argPattern).test(serialized)) continue;
|
||||
} catch {
|
||||
// Invalid regex in a rule — skip it rather than blocking every call to this tool.
|
||||
continue;
|
||||
}
|
||||
}
|
||||
if (!r.allow) return "deny";
|
||||
allowMatch = true;
|
||||
}
|
||||
return allowMatch ? "allow" : null;
|
||||
}
|
||||
|
||||
/** Check whether a mutating tool should be auto-approved (no confirmation needed). */
|
||||
isAutoApproved(toolName: string): boolean {
|
||||
// auto-accept approves EVERY mutating tool (bash, git_commit, write_file, edit_file, ...) —
|
||||
// the "I trust everything, don't ask" mode. auto-edit is the narrower "approve file edits only"
|
||||
// mode, auto-approving just the file-edit tools so a user can batch-edit without approving each
|
||||
// one but still gate dangerous shell/git operations. default falls back to the session-allowed
|
||||
// list (tools the user approved "for this session" in a prior prompt).
|
||||
if (this.mode === "auto-accept") return true;
|
||||
isAutoApproved(toolName: string, args?: unknown): boolean {
|
||||
// An explicit allow rule auto-approves (a deny rule is handled separately, in the gate, before
|
||||
// this is reached — so we only need to check for "allow" here).
|
||||
if (this.checkRules(toolName, args) === "allow") return true;
|
||||
// auto-accept only covers the same file-edit tools as auto-edit, not arbitrary mutating tools
|
||||
// such as bash or git_commit. This prevents a user who intended "approve edits" from silently
|
||||
// approving every dangerous operation.
|
||||
if (this.mode === "auto-accept" && AUTO_EDIT_TOOLS.has(toolName)) return true;
|
||||
if (this.mode === "auto-edit" && AUTO_EDIT_TOOLS.has(toolName)) return true;
|
||||
// "default" — check session-allowed list
|
||||
return this.allowedForSession.has(toolName);
|
||||
|
||||
@@ -5,6 +5,19 @@ export type PermissionMode = "default" | "auto-edit" | "auto-accept" | "plan";
|
||||
/** Tools that are auto-accepted in "auto-edit" mode. */
|
||||
export const AUTO_EDIT_TOOLS = new Set(["write_file", "edit_file"]);
|
||||
|
||||
/** A persistent permission rule (from user config or project .locode/settings.json) that either
|
||||
* auto-approves or outright blocks a tool call, without prompting. When `argPattern` is omitted the
|
||||
* rule matches any call to `tool`; when present it's a regex tested against JSON.stringify(args), so
|
||||
* e.g. `{ tool: "bash", argPattern: "npm (test|run)", allow: true }` auto-approves test/run
|
||||
* commands while still prompting for others (the regex matches a substring of the serialized args,
|
||||
* so don't anchor with `^` — that would bind to the leading `{` of the JSON). Deny rules take
|
||||
* precedence over allow rules. */
|
||||
export interface PermissionRule {
|
||||
tool: string;
|
||||
argPattern?: string;
|
||||
allow: boolean;
|
||||
}
|
||||
|
||||
export type ConfirmFn = (opts: {
|
||||
toolName: string;
|
||||
args: unknown;
|
||||
|
||||
@@ -1,87 +0,0 @@
|
||||
import { mkdtempSync, readFileSync, rmSync } from "node:fs";
|
||||
import os from "node:os";
|
||||
import path from "node:path";
|
||||
import { describe, it, expect, afterEach, beforeEach } from "vitest";
|
||||
import type { ChatCompletionMessageParam } from "openai/resources/chat/completions";
|
||||
import { exportSession, sessionToJson, sessionToMarkdown, defaultExportFilename } from "./exportSession.js";
|
||||
|
||||
const meta = { model: "test-model", createdAt: "2026-01-01T00:00:00.000Z" };
|
||||
|
||||
const messages: ChatCompletionMessageParam[] = [
|
||||
{ role: "user", content: "hello" },
|
||||
{ role: "assistant", content: "let me check", tool_calls: [{ id: "call_1", type: "function", function: { name: "read_file", arguments: '{"path":"a.ts"}' } }] },
|
||||
{ role: "tool", tool_call_id: "call_1", content: "file contents" },
|
||||
{ role: "assistant", content: "done" },
|
||||
];
|
||||
|
||||
describe("sessionToMarkdown (full transcript)", () => {
|
||||
it("includes user text, assistant text, tool calls, and tool results", () => {
|
||||
const md = sessionToMarkdown(messages, meta);
|
||||
expect(md).toContain("### You");
|
||||
expect(md).toContain("hello");
|
||||
expect(md).toContain("let me check");
|
||||
expect(md).toContain("#### Tool calls");
|
||||
expect(md).toContain('"name": "read_file"');
|
||||
expect(md).toContain("#### Tool result");
|
||||
expect(md).toContain("file contents");
|
||||
expect(md).toContain("done");
|
||||
});
|
||||
|
||||
it("does not drop a tool-only assistant turn (no text)", () => {
|
||||
const md = sessionToMarkdown(
|
||||
[{ role: "assistant", content: null, tool_calls: [{ id: "c", type: "function", function: { name: "grep", arguments: "{}" } }] } as ChatCompletionMessageParam],
|
||||
meta,
|
||||
);
|
||||
expect(md).toContain("#### Tool calls");
|
||||
expect(md).toContain('"name": "grep"');
|
||||
});
|
||||
});
|
||||
|
||||
describe("sessionToJson", () => {
|
||||
it("produces a JSON object with meta, exportedAt, and the verbatim messages", () => {
|
||||
const json = sessionToJson(messages, meta);
|
||||
const parsed = JSON.parse(json);
|
||||
expect(parsed.model).toBe("test-model");
|
||||
expect(parsed.exportedAt).toBeTruthy();
|
||||
expect(parsed.messages).toHaveLength(4);
|
||||
expect(parsed.messages[1].tool_calls[0].function.name).toBe("read_file");
|
||||
});
|
||||
});
|
||||
|
||||
describe("defaultExportFilename", () => {
|
||||
it("defaults to a .md extension", () => {
|
||||
expect(defaultExportFilename()).toMatch(/\.md$/);
|
||||
});
|
||||
it("uses .json for the json format", () => {
|
||||
expect(defaultExportFilename("json")).toMatch(/\.json$/);
|
||||
});
|
||||
});
|
||||
|
||||
describe("exportSession", () => {
|
||||
let cwd: string;
|
||||
beforeEach(() => {
|
||||
cwd = mkdtempSync(path.join(os.tmpdir(), "locode-export-"));
|
||||
});
|
||||
afterEach(() => {
|
||||
rmSync(cwd, { recursive: true, force: true });
|
||||
});
|
||||
|
||||
it("writes a markdown file by default", async () => {
|
||||
const resolved = await exportSession(messages, meta, cwd, "out.md");
|
||||
const content = readFileSync(resolved, "utf-8");
|
||||
expect(content).toContain("# locode conversation");
|
||||
expect(content).toContain("hello");
|
||||
});
|
||||
|
||||
it("writes a JSON file when format is json", async () => {
|
||||
const resolved = await exportSession(messages, meta, cwd, "out.json", "json");
|
||||
const content = readFileSync(resolved, "utf-8");
|
||||
const parsed = JSON.parse(content);
|
||||
expect(parsed.messages).toHaveLength(4);
|
||||
});
|
||||
|
||||
it("auto-generates a filename with the right extension when none given", async () => {
|
||||
const resolved = await exportSession(messages, meta, cwd, undefined, "json");
|
||||
expect(resolved).toMatch(/\.json$/);
|
||||
});
|
||||
});
|
||||
@@ -3,41 +3,15 @@ import path from "node:path";
|
||||
import type { ChatCompletionMessageParam } from "openai/resources/chat/completions";
|
||||
import { writeFileAtomic } from "../utils/writeFileAtomic.js";
|
||||
|
||||
/** Renders one message as a markdown section for a FULL transcript export — including tool
|
||||
* calls and their results, which the old prose-only export dropped. A tool-call assistant turn
|
||||
* lists each call as a fenced JSON block; a tool-result message is rendered as a fenced result.
|
||||
* Multimodal user content (text + image parts) is reduced to its text parts plus an
|
||||
* `[image attached]` placeholder. Returns null only for genuinely empty turns. */
|
||||
/** Mirrors the filtering used when replaying a resumed session (see App.tsx initSessionFromRecord):
|
||||
* only plain user/assistant text turns are human-readable — raw tool-call/tool-result payloads
|
||||
* and fallback-mode `tool_result` blocks are internal bookkeeping, not conversation content. */
|
||||
function messageSection(m: ChatCompletionMessageParam): string | null {
|
||||
if (m.role === "user") {
|
||||
if (typeof m.content === "string") {
|
||||
if (m.content.startsWith("```tool_result")) {
|
||||
// A fallback-mode tool result block — render it verbatim under a Tool result heading.
|
||||
return `#### Tool result\n\n${m.content}`;
|
||||
}
|
||||
return `### You\n\n${m.content}`;
|
||||
}
|
||||
if (Array.isArray(m.content)) {
|
||||
const parts = m.content.map((p) => (p.type === "text" ? p.text : "[image attached]")).join("\n");
|
||||
return parts.trim() ? `### You\n\n${parts}` : null;
|
||||
}
|
||||
return null;
|
||||
if (m.role === "user" && typeof m.content === "string" && !m.content.startsWith("```tool_result")) {
|
||||
return `### You\n\n${m.content}`;
|
||||
}
|
||||
if (m.role === "assistant") {
|
||||
const text = typeof m.content === "string" ? m.content : "";
|
||||
const calls = (m as { tool_calls?: { id: string; function: { name: string; arguments: string } }[] }).tool_calls;
|
||||
const parts: string[] = [];
|
||||
if (text.trim()) parts.push(`### Assistant\n\n${text}`);
|
||||
if (calls && calls.length) {
|
||||
const block = calls.map((c) => `{"name": "${c.function.name}", "arguments": ${c.function.arguments}}`).join("\n");
|
||||
parts.push(`#### Tool calls\n\n` + "```json\n" + block + "\n```");
|
||||
}
|
||||
return parts.length ? parts.join("\n\n") : null;
|
||||
}
|
||||
if (m.role === "tool") {
|
||||
const tm = m as { content?: string; tool_call_id?: string };
|
||||
const body = typeof tm.content === "string" ? tm.content : JSON.stringify(tm.content);
|
||||
return `#### Tool result${tm.tool_call_id ? ` (${tm.tool_call_id})` : ""}\n\n` + "```\n" + body + "\n```";
|
||||
if (m.role === "assistant" && typeof m.content === "string" && m.content) {
|
||||
return `### Assistant\n\n${m.content}`;
|
||||
}
|
||||
return null;
|
||||
}
|
||||
@@ -47,8 +21,6 @@ export interface ExportMeta {
|
||||
createdAt: string;
|
||||
}
|
||||
|
||||
export type ExportFormat = "markdown" | "json";
|
||||
|
||||
export function sessionToMarkdown(messages: ChatCompletionMessageParam[], meta: ExportMeta): string {
|
||||
const header = [
|
||||
"# locode conversation",
|
||||
@@ -61,35 +33,22 @@ export function sessionToMarkdown(messages: ChatCompletionMessageParam[], meta:
|
||||
return [header, ...sections].join("\n\n");
|
||||
}
|
||||
|
||||
/** A JSON export is the full record (messages verbatim + metadata), suitable for cross-machine
|
||||
* replay/sharing or feeding into another tool. The markdown export is for humans. */
|
||||
export function sessionToJson(messages: ChatCompletionMessageParam[], meta: ExportMeta): string {
|
||||
return JSON.stringify({ ...meta, exportedAt: new Date().toISOString(), messages }, null, 2);
|
||||
}
|
||||
|
||||
export function defaultExportFilename(format: ExportFormat = "markdown"): string {
|
||||
export function defaultExportFilename(): string {
|
||||
const stamp = new Date().toISOString().replace(/[:.]/g, "-");
|
||||
return `locode-export-${stamp}.${format === "json" ? "json" : "md"}`;
|
||||
return `locode-export-${stamp}.md`;
|
||||
}
|
||||
|
||||
/** Writes the conversation to a file and returns the resolved absolute path. `target` may be a bare
|
||||
* filename, a relative path, or an absolute path; a bare directory (or nothing at all) falls back
|
||||
* to an auto-generated filename inside `cwd`. `format` selects a human markdown transcript
|
||||
* (default, now including tool calls/results) or a machine-readable JSON dump. Written atomically
|
||||
/** Writes the conversation to a markdown file and returns the resolved absolute path.
|
||||
* `target` may be a bare filename, a relative path, or an absolute path; a bare directory
|
||||
* (or nothing at all) falls back to an auto-generated filename inside `cwd`. Written atomically
|
||||
* (temp file + rename), matching sessionStore's saves, so a crash mid-export can't leave a
|
||||
* truncated file. */
|
||||
export async function exportSession(
|
||||
messages: ChatCompletionMessageParam[],
|
||||
meta: ExportMeta,
|
||||
cwd: string,
|
||||
target?: string,
|
||||
format: ExportFormat = "markdown",
|
||||
): Promise<string> {
|
||||
const filename = target?.trim() || defaultExportFilename(format);
|
||||
export async function exportSession(messages: ChatCompletionMessageParam[], meta: ExportMeta, cwd: string, target?: string): Promise<string> {
|
||||
const filename = target?.trim() || defaultExportFilename();
|
||||
let resolved = path.isAbsolute(filename) ? filename : path.resolve(cwd, filename);
|
||||
if (existsSync(resolved) && statSync(resolved).isDirectory()) {
|
||||
resolved = path.join(resolved, defaultExportFilename(format));
|
||||
resolved = path.join(resolved, defaultExportFilename());
|
||||
}
|
||||
await writeFileAtomic(resolved, format === "json" ? sessionToJson(messages, meta) : sessionToMarkdown(messages, meta));
|
||||
await writeFileAtomic(resolved, sessionToMarkdown(messages, meta));
|
||||
return resolved;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,81 +0,0 @@
|
||||
import { describe, expect, it } from "vitest";
|
||||
import type { ChatCompletionMessageParam } from "openai/resources/chat/completions";
|
||||
import { buildReplayHistory } from "./replayHistory.js";
|
||||
|
||||
describe("buildReplayHistory", () => {
|
||||
it("replays plain user/assistant text turns", () => {
|
||||
const messages: ChatCompletionMessageParam[] = [
|
||||
{ role: "user", content: "hi" },
|
||||
{ role: "assistant", content: "hello there" },
|
||||
];
|
||||
expect(buildReplayHistory(messages)).toMatchObject([
|
||||
{ kind: "user", text: "hi" },
|
||||
{ kind: "assistant", text: "hello there" },
|
||||
]);
|
||||
});
|
||||
|
||||
it("replays native-mode tool calls and results, correlating the result's name via tool_call_id", () => {
|
||||
const messages: ChatCompletionMessageParam[] = [
|
||||
{ role: "user", content: "read foo.txt" },
|
||||
{
|
||||
role: "assistant",
|
||||
content: null,
|
||||
tool_calls: [
|
||||
{ id: "call_1", type: "function", function: { name: "read_file", arguments: JSON.stringify({ path: "foo.txt" }) } },
|
||||
],
|
||||
},
|
||||
{ role: "tool", tool_call_id: "call_1", content: JSON.stringify({ totalLines: 3 }) },
|
||||
{ role: "assistant", content: "It has 3 lines." },
|
||||
];
|
||||
expect(buildReplayHistory(messages)).toMatchObject([
|
||||
{ kind: "user", text: "read foo.txt" },
|
||||
{ kind: "tool_call", label: "Read(foo.txt)" },
|
||||
{ kind: "tool_result", summary: "Read 3 lines", isError: false },
|
||||
{ kind: "assistant", text: "It has 3 lines." },
|
||||
]);
|
||||
});
|
||||
|
||||
it("marks a native-mode tool error result", () => {
|
||||
const messages: ChatCompletionMessageParam[] = [
|
||||
{
|
||||
role: "assistant",
|
||||
content: null,
|
||||
tool_calls: [{ id: "call_1", type: "function", function: { name: "bash", arguments: "{}" } }],
|
||||
},
|
||||
{ role: "tool", tool_call_id: "call_1", content: JSON.stringify({ error: "command not found" }) },
|
||||
];
|
||||
expect(buildReplayHistory(messages)).toMatchObject([
|
||||
{ kind: "tool_call", label: "Bash()" },
|
||||
{ kind: "tool_result", summary: "command not found", isError: true },
|
||||
]);
|
||||
});
|
||||
|
||||
it("replays fallback-mode tool_call blocks embedded in assistant text", () => {
|
||||
const messages: ChatCompletionMessageParam[] = [
|
||||
{
|
||||
role: "assistant",
|
||||
content: 'Let me check.\n```tool_call\n{"name":"read_file","arguments":{"path":"foo.txt"}}\n```',
|
||||
},
|
||||
{
|
||||
role: "user",
|
||||
content: '```tool_result\n{"name":"read_file","result":{"totalLines":3}}\n```',
|
||||
},
|
||||
];
|
||||
expect(buildReplayHistory(messages)).toMatchObject([
|
||||
{ kind: "assistant", text: "Let me check." },
|
||||
{ kind: "tool_call", label: "Read(foo.txt)" },
|
||||
{ kind: "tool_result", summary: "Read 3 lines", isError: false },
|
||||
]);
|
||||
});
|
||||
|
||||
it("drops the empty assistant text bubble when a native tool call has no accompanying prose", () => {
|
||||
const messages: ChatCompletionMessageParam[] = [
|
||||
{
|
||||
role: "assistant",
|
||||
content: null,
|
||||
tool_calls: [{ id: "call_1", type: "function", function: { name: "grep", arguments: "{}" } }],
|
||||
},
|
||||
];
|
||||
expect(buildReplayHistory(messages)).toMatchObject([{ kind: "tool_call", label: "Grep()" }]);
|
||||
});
|
||||
});
|
||||
@@ -1,108 +0,0 @@
|
||||
import type { ChatCompletionMessageParam } from "openai/resources/chat/completions";
|
||||
import { parseFallbackToolCalls } from "../toolcalling/fallbackParser.js";
|
||||
import { nextId, type HistoryItem, type NewHistoryItem } from "../ui/ink/types.js";
|
||||
import { formatCallLabel, summarizeToolResult } from "../ui/toolSummary.js";
|
||||
|
||||
const TOOL_CALL_BLOCK_RE = /```tool_call\s*[\s\S]*?```/g;
|
||||
const TOOL_RESULT_BLOCK_RE = /^```tool_result\n([\s\S]*?)\n```$/;
|
||||
|
||||
/** A message's `content` can be a plain string or an array of content parts (text/image_url) —
|
||||
* see pushToolResultMessage in agent/loop.ts. Only the text part matters for replay display. */
|
||||
function textContent(content: ChatCompletionMessageParam["content"]): string | null {
|
||||
if (typeof content === "string") return content;
|
||||
if (Array.isArray(content)) {
|
||||
const part = content.find((p): p is { type: "text"; text: string } => (p as { type?: string }).type === "text");
|
||||
return part?.text ?? null;
|
||||
}
|
||||
return null;
|
||||
}
|
||||
|
||||
function isErrorResult(result: unknown): boolean {
|
||||
return !!(result && typeof result === "object" && "error" in (result as object));
|
||||
}
|
||||
|
||||
/** Rebuilds the tool_call/tool_result HistoryItems a resumed session's transcript is otherwise
|
||||
* missing (see App.tsx initSessionFromRecord) — the persisted record (SessionRecord.messages) is
|
||||
* the raw OpenAI-shape history, which carries everything needed (tool name, arguments, result)
|
||||
* even though it was never saved as a pre-rendered display label. Handles both native mode
|
||||
* (assistant `tool_calls` + matching `tool`-role messages, correlated by `tool_call_id`) and
|
||||
* fallback mode (` ```tool_call``` ` blocks embedded in assistant text, ` ```tool_result``` `
|
||||
* blocks embedded in user text — see fallbackParser.ts and pushToolResultMessage). */
|
||||
export function buildReplayHistory(messages: ChatCompletionMessageParam[]): HistoryItem[] {
|
||||
const items: HistoryItem[] = [];
|
||||
// Native mode: `tool` messages only carry a tool_call_id, not the tool's name — remember the
|
||||
// name from the assistant message that made the call so its later result can be labeled.
|
||||
const pendingCallNames = new Map<string, string>();
|
||||
|
||||
const push = (item: NewHistoryItem) => items.push({ id: nextId(), ...item } as HistoryItem);
|
||||
|
||||
for (const m of messages) {
|
||||
if (m.role === "user") {
|
||||
const text = textContent(m.content);
|
||||
if (text === null) continue;
|
||||
const fallbackResult = TOOL_RESULT_BLOCK_RE.exec(text.trim());
|
||||
if (fallbackResult) {
|
||||
try {
|
||||
const { name, result } = JSON.parse(fallbackResult[1]!) as { name: string; result: unknown };
|
||||
push({ kind: "tool_result", summary: summarizeToolResult(name, result), isError: isErrorResult(result) });
|
||||
} catch {
|
||||
// Malformed persisted block (shouldn't happen since we wrote it) — drop rather than
|
||||
// show the raw JSON to the user.
|
||||
}
|
||||
continue;
|
||||
}
|
||||
if (text) push({ kind: "user", text });
|
||||
continue;
|
||||
}
|
||||
|
||||
if (m.role === "assistant") {
|
||||
const toolCalls = m.tool_calls;
|
||||
if (toolCalls?.length) {
|
||||
const text = textContent(m.content);
|
||||
if (text) push({ kind: "assistant", text });
|
||||
for (const call of toolCalls) {
|
||||
if (call.type !== "function") continue;
|
||||
let args: unknown = {};
|
||||
try {
|
||||
args = JSON.parse(call.function.arguments);
|
||||
} catch {
|
||||
// Leave args as {} — formatCallLabel degrades gracefully for a missing field.
|
||||
}
|
||||
pendingCallNames.set(call.id, call.function.name);
|
||||
push({ kind: "tool_call", label: formatCallLabel(call.function.name, args) });
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
const text = textContent(m.content);
|
||||
if (!text) continue;
|
||||
const parsed = parseFallbackToolCalls(text);
|
||||
if (parsed.calls.length) {
|
||||
const stripped = text.replace(TOOL_CALL_BLOCK_RE, "").trim();
|
||||
if (stripped) push({ kind: "assistant", text: stripped });
|
||||
for (const call of parsed.calls) {
|
||||
push({ kind: "tool_call", label: formatCallLabel(call.name, call.arguments) });
|
||||
}
|
||||
} else {
|
||||
push({ kind: "assistant", text });
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
if (m.role === "tool") {
|
||||
const name = pendingCallNames.get(m.tool_call_id) ?? "unknown";
|
||||
const text = textContent(m.content);
|
||||
let result: unknown = text;
|
||||
if (text) {
|
||||
try {
|
||||
result = JSON.parse(text);
|
||||
} catch {
|
||||
result = text;
|
||||
}
|
||||
}
|
||||
push({ kind: "tool_result", summary: summarizeToolResult(name, result), isError: isErrorResult(result) });
|
||||
}
|
||||
}
|
||||
|
||||
return items;
|
||||
}
|
||||
@@ -1,14 +1,6 @@
|
||||
import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
|
||||
import { existsSync, mkdirSync, rmSync, unlinkSync, writeFileSync } from "node:fs";
|
||||
import { afterEach, beforeEach, describe, expect, it } from "vitest";
|
||||
import { existsSync, mkdirSync, rmSync, writeFileSync } from "node:fs";
|
||||
import path from "node:path";
|
||||
|
||||
// unlinkSync is mocked (default: pass-through to the real implementation) only so the "can't
|
||||
// actually delete the file" test below can make a single call fail — every other test's calls to
|
||||
// unlinkSync still hit the real filesystem via this same mock.
|
||||
vi.mock("node:fs", async (importOriginal) => {
|
||||
const actual = await importOriginal<typeof import("node:fs")>();
|
||||
return { ...actual, unlinkSync: vi.fn(actual.unlinkSync) };
|
||||
});
|
||||
import envPaths from "env-paths";
|
||||
import {
|
||||
deleteSession,
|
||||
@@ -78,32 +70,6 @@ describe("sessionStore", () => {
|
||||
expect(listSessions()).toHaveLength(0);
|
||||
});
|
||||
|
||||
it("rebuilds the summary index from disk when no index file exists yet", () => {
|
||||
writeFileSync(path.join(dir, "manual-session.json"), JSON.stringify(makeRecord("manual-session")));
|
||||
expect(listSessions()).toHaveLength(1);
|
||||
expect(listSessions()[0]?.id).toBe("manual-session");
|
||||
});
|
||||
|
||||
it("reports failure (not success) when the file can't actually be deleted", async () => {
|
||||
// Regression: a real unlink failure (e.g. Windows EPERM/EBUSY from a file lock) used to still
|
||||
// return true and drop the entry from the index — reporting success while orphaning the file
|
||||
// on disk with no way to reference it again.
|
||||
await saveSession(makeRecord("locked-session"));
|
||||
vi.mocked(unlinkSync).mockImplementationOnce(() => {
|
||||
throw Object.assign(new Error("EBUSY: resource busy or locked"), { code: "EBUSY" });
|
||||
});
|
||||
expect(deleteSession("locked-session")).toBe(false);
|
||||
expect(loadSession("locked-session")?.id).toBe("locked-session");
|
||||
expect(listSessions()).toHaveLength(1);
|
||||
});
|
||||
|
||||
it("self-heals when a session file is deleted outside of deleteSession()", async () => {
|
||||
await saveSession(makeRecord("will-vanish"));
|
||||
expect(listSessions()).toHaveLength(1);
|
||||
unlinkSync(path.join(dir, "will-vanish.json"));
|
||||
expect(listSessions()).toHaveLength(0);
|
||||
});
|
||||
|
||||
it("deriveTitle extracts the first user message", () => {
|
||||
const title = deriveTitle([
|
||||
{ role: "system", content: "sys" },
|
||||
|
||||
+26
-118
@@ -1,9 +1,9 @@
|
||||
import envPaths from "env-paths";
|
||||
import { existsSync, readdirSync, readFileSync, renameSync, unlinkSync, writeFileSync } from "node:fs";
|
||||
import { existsSync, readdirSync, readFileSync, unlinkSync } from "node:fs";
|
||||
import path from "node:path";
|
||||
import type { ChatCompletionMessageParam } from "openai/resources/chat/completions";
|
||||
import type { ToolCallMode } from "../backend/capabilityProbe.js";
|
||||
import type { TaskStoreSnapshot } from "../tools/task.js";
|
||||
import type { PermissionMode } from "../permissions/types.js";
|
||||
import { writeFileAtomic } from "../utils/writeFileAtomic.js";
|
||||
|
||||
/** A saved conversation. `messages` excludes the system prompt — it's rebuilt fresh from the
|
||||
@@ -16,11 +16,10 @@ export interface SessionRecord {
|
||||
baseURL: string;
|
||||
model: string;
|
||||
mode: ToolCallMode;
|
||||
/** Permission mode at save time, so plan/auto-edit/auto-accept survives a resume instead of
|
||||
* always resetting to default. Omitted by older saved sessions — treated as "default". */
|
||||
permissionMode?: PermissionMode;
|
||||
messages: ChatCompletionMessageParam[];
|
||||
/** Tool names the user approved "for this session" — preserved across resume. */
|
||||
allowedTools?: string[];
|
||||
/** Persisted task store snapshot so tasks survive session resume. */
|
||||
tasks?: TaskStoreSnapshot;
|
||||
}
|
||||
|
||||
export interface SessionSummary {
|
||||
@@ -50,92 +49,6 @@ function filePath(id: string): string {
|
||||
return path.join(dir, `${safeSessionId(id)}.json`);
|
||||
}
|
||||
|
||||
const INDEX_FILENAME = "_index.json";
|
||||
|
||||
function indexFilePath(): string {
|
||||
return path.join(dir, INDEX_FILENAME);
|
||||
}
|
||||
|
||||
function summarize(record: SessionRecord): SessionSummary {
|
||||
return {
|
||||
id: record.id,
|
||||
updatedAt: record.updatedAt,
|
||||
title: deriveTitle(record.messages),
|
||||
model: record.model,
|
||||
baseURL: record.baseURL,
|
||||
messageCount: record.messages.length,
|
||||
};
|
||||
}
|
||||
|
||||
/** Reads the on-disk summary index and reconciles it against the actual session files, so a stale or
|
||||
* missing index self-heals instead of ever going wrong: entries whose file was deleted (by this or
|
||||
* another locode process) are dropped, and files present on disk but missing from the index (a fresh
|
||||
* install, a crash before the last index write, another process's save racing this one) are parsed
|
||||
* individually. This keeps the common case to O(session count) stat calls instead of O(total
|
||||
* transcript bytes) — listSessions() used to JSON.parse every saved session in full just to read 6
|
||||
* summary fields off each one. */
|
||||
function readIndex(): Map<string, SessionSummary> {
|
||||
let index = new Map<string, SessionSummary>();
|
||||
if (existsSync(indexFilePath())) {
|
||||
try {
|
||||
const entries = JSON.parse(readFileSync(indexFilePath(), "utf-8")) as SessionSummary[];
|
||||
index = new Map(entries.map((e) => [e.id, e]));
|
||||
} catch (err) {
|
||||
// eslint-disable-next-line no-console
|
||||
console.warn("[sessionStore] failed to parse index file, rebuilding:", err);
|
||||
index = new Map();
|
||||
}
|
||||
}
|
||||
for (const id of index.keys()) {
|
||||
if (!existsSync(filePath(id))) index.delete(id);
|
||||
}
|
||||
const known = new Set([...index.keys()].map((id) => safeSessionId(id)));
|
||||
if (existsSync(dir)) {
|
||||
for (const entry of readdirSync(dir)) {
|
||||
if (!entry.endsWith(".json") || entry === INDEX_FILENAME) continue;
|
||||
const stem = entry.slice(0, -".json".length);
|
||||
if (known.has(stem)) continue;
|
||||
try {
|
||||
const record = JSON.parse(readFileSync(path.join(dir, entry), "utf-8")) as SessionRecord;
|
||||
index.set(record.id, summarize(record));
|
||||
} catch (err) {
|
||||
// Skip corrupt/partial session files — but log so disk issues aren't silent.
|
||||
// eslint-disable-next-line no-console
|
||||
console.warn(`[sessionStore] skipping corrupt session file ${entry}:`, err);
|
||||
}
|
||||
}
|
||||
}
|
||||
return index;
|
||||
}
|
||||
|
||||
/** Best-effort atomic write of the index. A failed write just means the next readIndex() call
|
||||
* re-parses whatever files it doesn't recognize yet — never incorrect data, only a missed
|
||||
* optimization. */
|
||||
function writeIndex(index: Map<string, SessionSummary>): void {
|
||||
const tmp = `${indexFilePath()}.${process.pid}.${Date.now()}.tmp`;
|
||||
try {
|
||||
writeFileSync(tmp, JSON.stringify([...index.values()]));
|
||||
renameSync(tmp, indexFilePath());
|
||||
} catch (err) {
|
||||
try {
|
||||
unlinkSync(tmp);
|
||||
} catch {
|
||||
// tmp may not have been created if the write itself failed
|
||||
}
|
||||
// eslint-disable-next-line no-console
|
||||
console.warn("[sessionStore] failed to write session index:", err);
|
||||
}
|
||||
}
|
||||
|
||||
/** Updates one entry in the persisted index. Two saves for different sessions racing this can lose
|
||||
* one's index write, but never lose data: readIndex() picks up any on-disk session file it doesn't
|
||||
* recognize, so the loser just costs the next listSessions() call one extra parse. */
|
||||
function updateIndexEntry(record: SessionRecord): void {
|
||||
const index = readIndex();
|
||||
index.set(record.id, summarize(record));
|
||||
writeIndex(index);
|
||||
}
|
||||
|
||||
export function deriveTitle(messages: ChatCompletionMessageParam[]): string {
|
||||
const first = messages.find((m) => m.role === "user");
|
||||
const text = first && typeof first.content === "string" ? first.content.trim() : "";
|
||||
@@ -154,8 +67,7 @@ const saveQueues = new Map<string, Promise<void>>();
|
||||
* otherwise load as "no session found" (see loadSession, which treats an unparseable file as absent). */
|
||||
export async function saveSession(record: SessionRecord): Promise<void> {
|
||||
const file = filePath(record.id);
|
||||
const write = () =>
|
||||
writeFileAtomic(file, JSON.stringify(record, null, 2)).then(() => updateIndexEntry(record));
|
||||
const write = () => writeFileAtomic(file, JSON.stringify(record, null, 2));
|
||||
// Run whether or not the previous save rejected, so one failure can't stall the chain.
|
||||
const prev = saveQueues.get(record.id);
|
||||
const next = (prev ?? Promise.resolve()).then(write, write);
|
||||
@@ -177,19 +89,31 @@ export function loadSession(id: string): SessionRecord | undefined {
|
||||
if (!existsSync(file)) return undefined;
|
||||
try {
|
||||
return JSON.parse(readFileSync(file, "utf-8")) as SessionRecord;
|
||||
} catch (err) {
|
||||
// Corrupt or unreadable session file — treat as absent, but log so disk issues
|
||||
// aren't completely silent. Callers can't distinguish "no file" from "corrupt file",
|
||||
// but at least the log preserves the reason.
|
||||
// eslint-disable-next-line no-console
|
||||
console.warn(`[sessionStore] failed to load session ${id}, treating as absent:`, err);
|
||||
} catch {
|
||||
return undefined;
|
||||
}
|
||||
}
|
||||
|
||||
export function listSessions(): SessionSummary[] {
|
||||
if (!existsSync(dir)) return [];
|
||||
return [...readIndex().values()].sort((a, b) => b.updatedAt.localeCompare(a.updatedAt));
|
||||
const summaries: SessionSummary[] = [];
|
||||
for (const entry of readdirSync(dir)) {
|
||||
if (!entry.endsWith(".json")) continue;
|
||||
try {
|
||||
const record = JSON.parse(readFileSync(path.join(dir, entry), "utf-8")) as SessionRecord;
|
||||
summaries.push({
|
||||
id: record.id,
|
||||
updatedAt: record.updatedAt,
|
||||
title: deriveTitle(record.messages),
|
||||
model: record.model,
|
||||
baseURL: record.baseURL,
|
||||
messageCount: record.messages.length,
|
||||
});
|
||||
} catch {
|
||||
// Skip corrupt/partial session files
|
||||
}
|
||||
}
|
||||
return summaries.sort((a, b) => b.updatedAt.localeCompare(a.updatedAt));
|
||||
}
|
||||
|
||||
export function mostRecentSessionId(): string | undefined {
|
||||
@@ -199,22 +123,6 @@ export function mostRecentSessionId(): string | undefined {
|
||||
export function deleteSession(id: string): boolean {
|
||||
const file = filePath(id);
|
||||
if (!existsSync(file)) return false;
|
||||
try {
|
||||
unlinkSync(file);
|
||||
} catch (err) {
|
||||
// ENOENT is harmless (race with another process) — the end state (file gone) is what we
|
||||
// wanted anyway, so fall through and report success. Any other error (e.g. EPERM/EBUSY from
|
||||
// a file lock, common on Windows) means the file is still on disk — report failure and leave
|
||||
// the index entry alone, or listSessions()/resume would silently orphan a file no one could
|
||||
// reference again (removed from the index, but never actually deleted).
|
||||
if ((err as NodeJS.ErrnoException).code !== "ENOENT") {
|
||||
// eslint-disable-next-line no-console
|
||||
console.warn(`[sessionStore] failed to delete session file ${file}:`, err);
|
||||
return false;
|
||||
}
|
||||
}
|
||||
const index = readIndex();
|
||||
index.delete(id);
|
||||
writeIndex(index);
|
||||
unlinkSync(file);
|
||||
return true;
|
||||
}
|
||||
|
||||
@@ -36,7 +36,7 @@ export function buildPluginAgentTool(agent: PluginAgentDef): ToolDef<{ prompt: s
|
||||
{ description: agent.name, prompt },
|
||||
{ systemPrompt: agent.systemPrompt, toolNames: agent.tools },
|
||||
);
|
||||
return { agent: agent.name, result };
|
||||
return { agent: agent.name, result: result.result, ...(result.resumable ? { agentId: result.agentId } : {}) };
|
||||
},
|
||||
};
|
||||
}
|
||||
|
||||
@@ -46,21 +46,4 @@ describe("loadPlugin", () => {
|
||||
expect(plugin.agents).toHaveLength(1);
|
||||
expect(plugin.agents[0]).toMatchObject({ name: "review", description: "Code reviewer" });
|
||||
});
|
||||
|
||||
it("parses a command's allowed-tools frontmatter, translating Claude Code tool names", () => {
|
||||
mkdirSync(path.join(tempDir, ".claude-plugin"), { recursive: true });
|
||||
writeFileSync(path.join(tempDir, ".claude-plugin", "plugin.json"), JSON.stringify({ name: "test-plugin" }));
|
||||
mkdirSync(path.join(tempDir, "commands"), { recursive: true });
|
||||
writeFileSync(
|
||||
path.join(tempDir, "commands", "readonly.md"),
|
||||
"---\ndescription: Look but don't touch\nallowed-tools: Read, Grep\n---\nInvestigate $ARGUMENTS",
|
||||
);
|
||||
writeFileSync(path.join(tempDir, "commands", "unrestricted.md"), "---\ndescription: No restriction\n---\nDo $ARGUMENTS");
|
||||
|
||||
const plugin = loadPlugin(tempDir);
|
||||
const readonly = plugin.commands.find((c) => c.name === "readonly");
|
||||
const unrestricted = plugin.commands.find((c) => c.name === "unrestricted");
|
||||
expect(readonly?.allowedTools).toEqual(["read_file", "grep"]);
|
||||
expect(unrestricted?.allowedTools).toBeUndefined();
|
||||
});
|
||||
});
|
||||
|
||||
+5
-13
@@ -20,17 +20,6 @@ function listMarkdownFiles(dir: string): string[] {
|
||||
return readdirSync(dir).filter((f) => f.endsWith(".md"));
|
||||
}
|
||||
|
||||
/** Parses a comma-separated tool-name list from frontmatter (agents' `tools`, commands'
|
||||
* `allowed-tools`) and translates each from Claude Code's built-in names to locode's. Returns
|
||||
* undefined for an absent/empty field so callers can treat that as "no restriction". */
|
||||
function parseToolList(field: string | undefined): string[] | undefined {
|
||||
const tools = field
|
||||
?.split(",")
|
||||
.map((t) => resolveToolName(t.trim()))
|
||||
.filter(Boolean);
|
||||
return tools?.length ? tools : undefined;
|
||||
}
|
||||
|
||||
function loadCommands(pluginRoot: string, pluginName: string): PluginCommand[] {
|
||||
const dir = path.join(pluginRoot, "commands");
|
||||
return listMarkdownFiles(dir).map((entry) => {
|
||||
@@ -41,7 +30,6 @@ function loadCommands(pluginRoot: string, pluginName: string): PluginCommand[] {
|
||||
description: frontmatter.description,
|
||||
argumentHint: frontmatter["argument-hint"],
|
||||
template: body,
|
||||
allowedTools: parseToolList(frontmatter["allowed-tools"]),
|
||||
};
|
||||
});
|
||||
}
|
||||
@@ -51,11 +39,15 @@ function loadAgents(pluginRoot: string, pluginName: string): PluginAgentDef[] {
|
||||
return listMarkdownFiles(dir).map((entry) => {
|
||||
const { frontmatter, body } = parseFrontmatter(readFileSync(path.join(dir, entry), "utf-8"));
|
||||
const name = frontmatter.name || entry.replace(/\.md$/, "");
|
||||
const tools = frontmatter.tools
|
||||
?.split(",")
|
||||
.map((t) => resolveToolName(t.trim()))
|
||||
.filter(Boolean);
|
||||
return {
|
||||
pluginName,
|
||||
name,
|
||||
description: frontmatter.description || name,
|
||||
tools: parseToolList(frontmatter.tools),
|
||||
tools: tools?.length ? tools : undefined,
|
||||
systemPrompt: body,
|
||||
};
|
||||
});
|
||||
|
||||
@@ -13,7 +13,10 @@ const CLAUDE_TOOL_NAME_MAP: Record<string, string> = {
|
||||
webfetch: "web_fetch",
|
||||
websearch: "web_search",
|
||||
task: "agent",
|
||||
todowrite: "todo_write",
|
||||
taskcreate: "task_create",
|
||||
tasklist: "task_list",
|
||||
taskget: "task_get",
|
||||
taskupdate: "task_update",
|
||||
};
|
||||
|
||||
export function resolveToolName(name: string): string {
|
||||
|
||||
@@ -16,10 +16,6 @@ export interface PluginCommand {
|
||||
/** The markdown body — expanded via $ARGUMENTS/$1../$9 (see expandTemplate.ts) and submitted
|
||||
* as the turn's input when the command is invoked. */
|
||||
template: string;
|
||||
/** Tool names this command's turn is restricted to (already translated from Claude Code's
|
||||
* built-in tool names — see toolNameMap.ts), from frontmatter `allowed-tools`; undefined means
|
||||
* no restriction (the command runs with the session's full toolset). */
|
||||
allowedTools?: string[];
|
||||
}
|
||||
|
||||
export interface PluginAgentDef {
|
||||
|
||||
@@ -0,0 +1,281 @@
|
||||
import { describe, expect, it, vi, beforeEach, afterEach } from "vitest";
|
||||
import { mkdtempSync, rmSync, existsSync, readFileSync } from "node:fs";
|
||||
import { tmpdir } from "node:os";
|
||||
import path from "node:path";
|
||||
import { CronStore, parseCron } from "./cron.js";
|
||||
import { cronCreateTool, cronDeleteTool, cronListTool, scheduleWakeupTool } from "../tools/cron.js";
|
||||
import type { ToolContext } from "../tools/types.js";
|
||||
|
||||
function ctxWith(store?: CronStore): ToolContext {
|
||||
return { cwd: "/x", ...(store ? { cronStore: store } : {}) };
|
||||
}
|
||||
|
||||
describe("parseCron", () => {
|
||||
it("parses a 5-field expression into allowed values", () => {
|
||||
const spec = parseCron("0 9 * * 1-5");
|
||||
expect(spec.minute).toEqual([0]);
|
||||
expect(spec.hour).toEqual([9]);
|
||||
expect(spec.dom).toEqual(Array.from({ length: 31 }, (_, i) => i + 1));
|
||||
expect(spec.month).toEqual([1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12]);
|
||||
expect(spec.dow).toEqual([1, 2, 3, 4, 5]);
|
||||
expect(spec.domStar).toBe(true);
|
||||
expect(spec.dowStar).toBe(false);
|
||||
});
|
||||
|
||||
it("supports */N step and comma-lists", () => {
|
||||
const spec = parseCron("*/15 8-17 * * 0,6");
|
||||
expect(spec.minute).toEqual([0, 15, 30, 45]);
|
||||
expect(spec.hour).toEqual([8, 9, 10, 11, 12, 13, 14, 15, 16, 17]);
|
||||
expect(spec.dow).toEqual([0, 6]);
|
||||
});
|
||||
|
||||
it("throws on a wrong field count", () => {
|
||||
expect(() => parseCron("0 9 * *")).toThrow(/5 fields/);
|
||||
expect(() => parseCron("0 9 * * * *")).toThrow(/5 fields/);
|
||||
});
|
||||
|
||||
it("throws on out-of-range values", () => {
|
||||
expect(() => parseCron("60 9 * * *")).toThrow(/out of range/);
|
||||
expect(() => parseCron("0 24 * * *")).toThrow(/out of range/);
|
||||
expect(() => parseCron("0 9 32 * *")).toThrow(/out of range/);
|
||||
});
|
||||
|
||||
it("matches a Date correctly (weekday cron with dom=* → dow governs)", () => {
|
||||
// Every Monday: dom=* (star), dow=1. 2026-08-17 is a Monday; 2026-08-16 is a Sunday.
|
||||
const s = parseCron("0 0 * * 1");
|
||||
const mon = new Date(2026, 7, 17, 0, 0);
|
||||
const sun = new Date(2026, 7, 16, 0, 0);
|
||||
// Inline the Vixie-cron matcher logic (the store uses it internally).
|
||||
const domMatch = (d: Date) => s.dom.includes(d.getDate());
|
||||
const dowMatch = (d: Date) => s.dow.includes(d.getDay());
|
||||
const ok = (d: Date) => (s.domStar ? dowMatch(d) : s.dowStar ? domMatch(d) : domMatch(d) || dowMatch(d));
|
||||
expect(ok(mon)).toBe(true);
|
||||
expect(ok(sun)).toBe(false);
|
||||
});
|
||||
|
||||
it("fires on EITHER dom OR dow when both are restricted (Vixie semantics)", () => {
|
||||
// dom=15, dow=0 (Sunday): fires on the 15th of any month OR any Sunday.
|
||||
const s = parseCron("0 0 15 * 0");
|
||||
const domMatch = (d: Date) => s.dom.includes(d.getDate());
|
||||
const dowMatch = (d: Date) => s.dow.includes(d.getDay());
|
||||
const ok = (d: Date) => (s.domStar ? dowMatch(d) : s.dowStar ? domMatch(d) : domMatch(d) || dowMatch(d));
|
||||
// 2026-08-16 is Sunday the 16th (not the 15th) → matches via dow.
|
||||
expect(ok(new Date(2026, 7, 16, 0, 0))).toBe(true);
|
||||
// 2026-08-15 is Saturday the 15th → matches via dom.
|
||||
expect(ok(new Date(2026, 7, 15, 0, 0))).toBe(true);
|
||||
// 2026-08-14 is Friday the 14th → no match.
|
||||
expect(ok(new Date(2026, 7, 14, 0, 0))).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe("CronStore", () => {
|
||||
it("create validates the cron expression up front", () => {
|
||||
const s = new CronStore();
|
||||
expect(() => s.create({ cron: "bad expr", prompt: "x" })).toThrow();
|
||||
});
|
||||
|
||||
it("create/list/delete round-trips jobs with sequential ids", () => {
|
||||
const s = new CronStore();
|
||||
const a = s.create({ cron: "0 9 * * *", prompt: "morning" });
|
||||
const b = s.create({ cron: "0 10 * * *", prompt: "late", recurring: false });
|
||||
expect(a.id).toBe("c1");
|
||||
expect(b.id).toBe("c2");
|
||||
expect(s.list()).toHaveLength(2);
|
||||
expect(s.delete("c1")).toBe(true);
|
||||
expect(s.list()).toHaveLength(1);
|
||||
expect(s.delete("nope")).toBe(false);
|
||||
});
|
||||
|
||||
it("tick fires due recurring jobs (deduped within the minute) and enqueues the prompt", () => {
|
||||
const s = new CronStore();
|
||||
const enqueued: string[] = [];
|
||||
s.start({ enqueue: (p) => enqueued.push(p), isIdle: () => true });
|
||||
// A cron that matches every minute, so it's definitely due now.
|
||||
s.create({ cron: "* * * * *", prompt: "tick" });
|
||||
// Call the private tick via a cast (the real loop uses setInterval).
|
||||
(s as unknown as { tick: () => void }).tick();
|
||||
expect(enqueued).toEqual(["tick"]);
|
||||
// A second tick in the same minute must NOT re-fire (deduped by lastFiredMinute).
|
||||
(s as unknown as { tick: () => void }).tick();
|
||||
expect(enqueued).toEqual(["tick"]);
|
||||
s.stop();
|
||||
});
|
||||
|
||||
it("tick does not fire while isIdle is false", () => {
|
||||
const s = new CronStore();
|
||||
const enqueued: string[] = [];
|
||||
s.start({ enqueue: (p) => enqueued.push(p), isIdle: () => false });
|
||||
s.create({ cron: "* * * * *", prompt: "tick" });
|
||||
(s as unknown as { tick: () => void }).tick();
|
||||
expect(enqueued).toEqual([]);
|
||||
s.stop();
|
||||
});
|
||||
|
||||
it("one-shot jobs (recurring:false) are deleted after firing once", () => {
|
||||
const s = new CronStore();
|
||||
const enqueued: string[] = [];
|
||||
s.start({ enqueue: (p) => enqueued.push(p), isIdle: () => true });
|
||||
s.create({ cron: "* * * * *", prompt: "once", recurring: false });
|
||||
(s as unknown as { tick: () => void }).tick();
|
||||
expect(enqueued).toEqual(["once"]);
|
||||
expect(s.list()).toHaveLength(0); // removed after the one fire
|
||||
s.stop();
|
||||
});
|
||||
|
||||
it("scheduleWakeup clamps delay to [60,3600] and fires once then is removed", () => {
|
||||
const s = new CronStore();
|
||||
const enqueued: string[] = [];
|
||||
s.start({ enqueue: (p) => enqueued.push(p), isIdle: () => true });
|
||||
const res = s.scheduleWakeup({ delaySeconds: 5, prompt: "wake" }) as { id: string; fireAt: number };
|
||||
expect(res.id).toMatch(/^w/);
|
||||
// Clamped to 60s, so a tick right now (well before fireAt) must not enqueue.
|
||||
(s as unknown as { tick: () => void }).tick();
|
||||
expect(enqueued).toEqual([]);
|
||||
// Force the wakeup into the past (mutate the STORE's internal entry, not the listWakeups copy)
|
||||
// and tick again so the wakeup is now due.
|
||||
const internal = (s as unknown as { wakeups: Map<string, { fireAt: number }> }).wakeups;
|
||||
const id = [...internal.keys()][0]!;
|
||||
internal.get(id)!.fireAt = Date.now() - 1000;
|
||||
(s as unknown as { tick: () => void }).tick();
|
||||
expect(enqueued).toEqual(["wake"]);
|
||||
expect(s.listWakeups()).toHaveLength(0);
|
||||
s.stop();
|
||||
});
|
||||
|
||||
it("scheduleWakeup with stop:true clears all wakeups", () => {
|
||||
const s = new CronStore();
|
||||
s.scheduleWakeup({ delaySeconds: 60, prompt: "a" });
|
||||
s.scheduleWakeup({ delaySeconds: 120, prompt: "b" });
|
||||
expect(s.listWakeups()).toHaveLength(2);
|
||||
const res = s.scheduleWakeup({ delaySeconds: 60, prompt: "", stop: true });
|
||||
expect(res).toEqual({ stopped: true });
|
||||
expect(s.listWakeups()).toHaveLength(0);
|
||||
});
|
||||
|
||||
describe("durable persistence", () => {
|
||||
let dir: string;
|
||||
beforeEach(() => {
|
||||
dir = mkdtempSync(path.join(tmpdir(), "locode-cron-test-"));
|
||||
});
|
||||
afterEach(() => {
|
||||
try {
|
||||
rmSync(dir, { recursive: true, force: true });
|
||||
} catch {
|
||||
/* leave for OS temp sweep */
|
||||
}
|
||||
});
|
||||
|
||||
it("persists durable jobs to scheduled_tasks.json and reloads them on construction", () => {
|
||||
const s1 = new CronStore(dir);
|
||||
s1.create({ cron: "0 9 * * *", prompt: "morning", durable: true });
|
||||
s1.create({ cron: "0 10 * * *", prompt: "ephemeral" }); // not durable
|
||||
expect(existsSync(path.join(dir, "scheduled_tasks.json"))).toBe(true);
|
||||
const raw = JSON.parse(readFileSync(path.join(dir, "scheduled_tasks.json"), "utf-8")) as {
|
||||
jobs: { prompt: string; durable: boolean }[];
|
||||
};
|
||||
expect(raw.jobs).toHaveLength(1);
|
||||
expect(raw.jobs[0]!.prompt).toBe("morning");
|
||||
|
||||
// A fresh store pointed at the same dir reloads the durable job only.
|
||||
const s2 = new CronStore(dir);
|
||||
expect(s2.list()).toHaveLength(1);
|
||||
expect(s2.list()[0]!.prompt).toBe("morning");
|
||||
// The reloaded id shouldn't collide with a new one (seq was bumped past it).
|
||||
const next = s2.create({ cron: "0 11 * * *", prompt: "next" });
|
||||
expect(next.id).not.toBe("c1");
|
||||
});
|
||||
|
||||
it("non-durable jobs are NOT persisted", () => {
|
||||
const s1 = new CronStore(dir);
|
||||
s1.create({ cron: "0 9 * * *", prompt: "ephemeral", durable: false });
|
||||
const s2 = new CronStore(dir);
|
||||
expect(s2.list()).toHaveLength(0);
|
||||
});
|
||||
});
|
||||
});
|
||||
|
||||
describe("cron/schedule tools", () => {
|
||||
it("cron_create returns the job id and snapshot, validating cron syntax", async () => {
|
||||
const store = new CronStore();
|
||||
const result = (await cronCreateTool.handler(
|
||||
{ cron: "0 9 * * 1-5", prompt: "weekday standup", recurring: true },
|
||||
ctxWith(store),
|
||||
)) as { id: string; job: { cron: string; recurring: boolean } };
|
||||
expect(result.id).toBe("c1");
|
||||
expect(result.job.cron).toBe("0 9 * * 1-5");
|
||||
});
|
||||
|
||||
it("cron_create returns an error for invalid cron syntax", async () => {
|
||||
const store = new CronStore();
|
||||
const result = (await cronCreateTool.handler(
|
||||
{ cron: "not cron", prompt: "x" },
|
||||
ctxWith(store),
|
||||
)) as { error: string };
|
||||
expect(result.error).toMatch(/field|5 fields|range/i);
|
||||
});
|
||||
|
||||
it("cron_create returns an error when no store is available", async () => {
|
||||
const result = (await cronCreateTool.handler({ cron: "0 9 * * *", prompt: "x" }, ctxWith(undefined))) as {
|
||||
error: string;
|
||||
};
|
||||
expect(result.error).toMatch(/not available/i);
|
||||
});
|
||||
|
||||
it("cron_list returns the jobs; empty (not error) when no store", async () => {
|
||||
const store = new CronStore();
|
||||
store.create({ cron: "0 9 * * *", prompt: "x" });
|
||||
const result = (await cronListTool.handler({}, ctxWith(store))) as { jobs: { id: string }[] };
|
||||
expect(result.jobs).toHaveLength(1);
|
||||
const empty = (await cronListTool.handler({}, ctxWith(undefined))) as { jobs: unknown[] };
|
||||
expect(empty.jobs).toEqual([]);
|
||||
});
|
||||
|
||||
it("cron_delete removes a job and returns { deleted }", async () => {
|
||||
const store = new CronStore();
|
||||
store.create({ cron: "0 9 * * *", prompt: "x" });
|
||||
const result = (await cronDeleteTool.handler({ id: "c1" }, ctxWith(store))) as { deleted: string };
|
||||
expect(result.deleted).toBe("c1");
|
||||
expect(store.list()).toHaveLength(0);
|
||||
});
|
||||
|
||||
it("cron_delete returns an error for an unknown id", async () => {
|
||||
const store = new CronStore();
|
||||
const result = (await cronDeleteTool.handler({ id: "c99" }, ctxWith(store))) as { error: string };
|
||||
expect(result.error).toMatch(/not found/i);
|
||||
});
|
||||
|
||||
it("schedule_wakeup returns a wakeup id, or { stopped } when stop is true", async () => {
|
||||
const store = new CronStore();
|
||||
const result = (await scheduleWakeupTool.handler(
|
||||
{ delaySeconds: 120, prompt: "check back" },
|
||||
ctxWith(store),
|
||||
)) as { id: string; fireAt: number };
|
||||
expect(result.id).toMatch(/^w/);
|
||||
const stop = (await scheduleWakeupTool.handler(
|
||||
{ delaySeconds: 60, prompt: "", stop: true },
|
||||
ctxWith(store),
|
||||
)) as { stopped: boolean };
|
||||
expect(stop.stopped).toBe(true);
|
||||
});
|
||||
|
||||
it("all cron/schedule tools are non-mutating (no confirmation prompt)", () => {
|
||||
expect(cronCreateTool.mutating).toBe(false);
|
||||
expect(cronListTool.mutating).toBe(false);
|
||||
expect(cronDeleteTool.mutating).toBe(false);
|
||||
expect(scheduleWakeupTool.mutating).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe("cron tool schema validation", () => {
|
||||
it("requires a cron expression and prompt on cron_create", () => {
|
||||
expect(() => cronCreateTool.schema.parse({ cron: "", prompt: "x" })).toThrow();
|
||||
expect(() => cronCreateTool.schema.parse({ cron: "0 9 * * *", prompt: "" })).toThrow();
|
||||
});
|
||||
it("requires an id on cron_delete", () => {
|
||||
expect(() => cronDeleteTool.schema.parse({ id: "" })).toThrow();
|
||||
});
|
||||
it("requires a positive integer delaySeconds on schedule_wakeup", () => {
|
||||
expect(() => scheduleWakeupTool.schema.parse({ delaySeconds: 0, prompt: "x" })).toThrow();
|
||||
expect(() => scheduleWakeupTool.schema.parse({ delaySeconds: 1.5, prompt: "x" })).toThrow();
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,296 @@
|
||||
import { readFileSync, writeFileSync, existsSync, mkdirSync } from "node:fs";
|
||||
import path from "node:path";
|
||||
|
||||
/** A scheduled, recurring cron job (5-field cron expression in the user's LOCAL timezone, matching
|
||||
* Claude Code). `recurring: false` is a one-shot that fires once then auto-deletes. `durable` jobs
|
||||
* are persisted to disk so they survive a restart; session-only jobs die with the process. */
|
||||
export interface CronJob {
|
||||
id: string;
|
||||
cron: string;
|
||||
prompt: string;
|
||||
recurring: boolean;
|
||||
durable: boolean;
|
||||
/** Epoch ms the job was created — used for the 7-day auto-expiry on recurring jobs. */
|
||||
createdAt: number;
|
||||
/** Epoch-minute of the most recent fire, so a job doesn't re-fire within the same minute. */
|
||||
lastFiredMinute?: number;
|
||||
/** Set true once the 7-day expiry has fired its final run, so the tick deletes it after enqueue. */
|
||||
expired?: boolean;
|
||||
}
|
||||
|
||||
/** A one-shot delayed prompt (the ScheduleWakeup primitive), used for self-paced loops. Fires once
|
||||
* at fireAt (epoch ms) then is removed. Session-only — never persisted. */
|
||||
export interface Wakeup {
|
||||
id: string;
|
||||
fireAt: number;
|
||||
prompt: string;
|
||||
}
|
||||
|
||||
/** Parsed 5-field cron spec. `domStar`/`dowStar` record whether the day-of-month / day-of-week fields
|
||||
* were `*` — needed for Vixie-cron semantics (when both are restricted, fire on EITHER match). */
|
||||
interface CronSpec {
|
||||
minute: number[];
|
||||
hour: number[];
|
||||
dom: number[];
|
||||
month: number[];
|
||||
dow: number[];
|
||||
domStar: boolean;
|
||||
dowStar: boolean;
|
||||
}
|
||||
|
||||
const FIELD_RANGES: Record<string, [number, number]> = {
|
||||
minute: [0, 59],
|
||||
hour: [0, 23],
|
||||
dom: [1, 31],
|
||||
month: [1, 12],
|
||||
dow: [0, 6],
|
||||
};
|
||||
|
||||
// Parses a single cron field into the sorted list of allowed values. Supports `*`, the `*/N` step
|
||||
// form, `N`, `N-M`, `N-M/S`, and comma-lists of any of these. Throws on out-of-range or unparseable
|
||||
// input. (Line comments, not JSDoc, because the `*/N` step syntax contains a `*/` that would close a
|
||||
// block comment prematurely.)
|
||||
function parseField(field: string, range: [number, number]): { values: number[]; isStar: boolean } {
|
||||
const min = range[0];
|
||||
const max = range[1];
|
||||
const out = new Set<number>();
|
||||
const isStar = field === "*" || field === "*/1";
|
||||
for (const part of field.split(",")) {
|
||||
const slashIdx = part.indexOf("/");
|
||||
let rangePart = part;
|
||||
let step = 1;
|
||||
if (slashIdx !== -1) {
|
||||
rangePart = part.slice(0, slashIdx);
|
||||
step = Number(part.slice(slashIdx + 1));
|
||||
}
|
||||
let lo: number;
|
||||
let hi: number;
|
||||
if (rangePart === "*") {
|
||||
lo = min;
|
||||
hi = max;
|
||||
} else if (rangePart.includes("-")) {
|
||||
const [a, b] = rangePart.split("-");
|
||||
lo = Number(a);
|
||||
hi = Number(b);
|
||||
} else {
|
||||
lo = hi = Number(rangePart);
|
||||
}
|
||||
if (!Number.isFinite(lo) || !Number.isFinite(hi) || !Number.isFinite(step) || step < 1) {
|
||||
throw new Error(`invalid cron field "${field}"`);
|
||||
}
|
||||
if (lo < min || hi > max || lo > hi) {
|
||||
throw new Error(`cron field "${field}" out of range [${min}-${max}]`);
|
||||
}
|
||||
for (let v = lo; v <= hi; v += step) out.add(v);
|
||||
}
|
||||
return { values: [...out].sort((a, b) => a - b), isStar };
|
||||
}
|
||||
|
||||
/** Parses a 5-field cron expression into a matcher spec. Throws on malformed input. */
|
||||
export function parseCron(expr: string): CronSpec {
|
||||
const fields = expr.trim().split(/\s+/);
|
||||
if (fields.length !== 5) throw new Error(`cron expression must have 5 fields, got ${fields.length}`);
|
||||
const minute = fields[0]!;
|
||||
const hour = fields[1]!;
|
||||
const dom = fields[2]!;
|
||||
const month = fields[3]!;
|
||||
const dow = fields[4]!;
|
||||
const m = parseField(minute, FIELD_RANGES.minute!);
|
||||
const h = parseField(hour, FIELD_RANGES.hour!);
|
||||
const dm = parseField(dom, FIELD_RANGES.dom!);
|
||||
const mo = parseField(month, FIELD_RANGES.month!);
|
||||
const dw = parseField(dow, FIELD_RANGES.dow!);
|
||||
return { minute: m.values, hour: h.values, dom: dm.values, month: mo.values, dow: dw.values, domStar: dm.isStar, dowStar: dw.isStar };
|
||||
}
|
||||
|
||||
/** Whether a cron spec matches a given local Date. */
|
||||
function cronMatches(spec: CronSpec, d: Date): boolean {
|
||||
if (!spec.minute.includes(d.getMinutes())) return false;
|
||||
if (!spec.hour.includes(d.getHours())) return false;
|
||||
if (!spec.month.includes(d.getMonth() + 1)) return false;
|
||||
const domMatch = spec.dom.includes(d.getDate());
|
||||
const dowMatch = spec.dow.includes(d.getDay());
|
||||
// Vixie-cron: when both day fields are restricted, fire on EITHER; when one is *, the other governs.
|
||||
if (spec.domStar && spec.dowStar) return true;
|
||||
if (spec.domStar) return dowMatch;
|
||||
if (spec.dowStar) return domMatch;
|
||||
return domMatch || dowMatch;
|
||||
}
|
||||
|
||||
const SEVEN_DAYS_MS = 7 * 24 * 60 * 60 * 1000;
|
||||
const TICK_INTERVAL_MS = 15_000;
|
||||
|
||||
/** In-memory scheduler for cron jobs and one-shot wakeups. Held by the Session; the App starts it
|
||||
* with an `enqueue` callback (submit a turn) and an `isIdle` predicate (true when no turn is
|
||||
* running) so jobs only fire while the REPL is idle, matching Claude Code. Durable jobs persist
|
||||
* to `<configDir>/scheduled_tasks.json`; session-only jobs don't. */
|
||||
export class CronStore {
|
||||
private jobs = new Map<string, CronJob>();
|
||||
private wakeups = new Map<string, Wakeup>();
|
||||
private seq = 0;
|
||||
private wakeSeq = 0;
|
||||
private timer: ReturnType<typeof setInterval> | undefined;
|
||||
private enqueue: ((prompt: string) => void) | undefined;
|
||||
private isIdle: (() => boolean) | undefined;
|
||||
private readonly file: string | undefined;
|
||||
|
||||
constructor(configDir?: string) {
|
||||
if (configDir) {
|
||||
this.file = path.join(configDir, "scheduled_tasks.json");
|
||||
this.loadDurable();
|
||||
}
|
||||
}
|
||||
|
||||
private nextId(): string {
|
||||
this.seq += 1;
|
||||
return `c${this.seq}`;
|
||||
}
|
||||
|
||||
private nextWakeupId(): string {
|
||||
this.wakeSeq += 1;
|
||||
return `w${this.wakeSeq}`;
|
||||
}
|
||||
|
||||
/** Begins the tick loop. Called once by the App after the session is wired. */
|
||||
start(opts: { enqueue: (prompt: string) => void; isIdle: () => boolean }): void {
|
||||
this.enqueue = opts.enqueue;
|
||||
this.isIdle = opts.isIdle;
|
||||
if (this.timer) return;
|
||||
this.timer = setInterval(() => this.tick(), TICK_INTERVAL_MS);
|
||||
// setInterval keeps the event loop alive; unref so the process can still exit naturally when the
|
||||
// UI closes (the App stops the store on unmount anyway, but this is a backstop).
|
||||
this.timer.unref?.();
|
||||
}
|
||||
|
||||
/** Stops the tick loop (e.g. on session/App teardown). */
|
||||
stop(): void {
|
||||
if (this.timer) {
|
||||
clearInterval(this.timer);
|
||||
this.timer = undefined;
|
||||
}
|
||||
}
|
||||
|
||||
private tick(): void {
|
||||
if (!this.enqueue || !this.isIdle) return;
|
||||
if (!this.isIdle()) return; // only fire while the REPL is idle
|
||||
const now = Date.now();
|
||||
const nowMinute = Math.floor(now / 60_000);
|
||||
const nowDate = new Date(now);
|
||||
const toDelete: string[] = [];
|
||||
for (const job of this.jobs.values()) {
|
||||
const age = now - job.createdAt;
|
||||
// 7-day auto-expiry: recurring jobs fire one final time then are deleted.
|
||||
if (job.recurring && age >= SEVEN_DAYS_MS) {
|
||||
this.enqueue(job.prompt);
|
||||
toDelete.push(job.id);
|
||||
continue;
|
||||
}
|
||||
if (job.lastFiredMinute === nowMinute) continue;
|
||||
const spec = parseCron(job.cron);
|
||||
if (cronMatches(spec, nowDate)) {
|
||||
this.enqueue(job.prompt);
|
||||
job.lastFiredMinute = nowMinute;
|
||||
if (!job.recurring) toDelete.push(job.id); // one-shot fires once then is removed
|
||||
}
|
||||
}
|
||||
for (const id of toDelete) this.delete(id);
|
||||
|
||||
// Wakeups: fire when their time has come.
|
||||
const wakeDelete: string[] = [];
|
||||
for (const w of this.wakeups.values()) {
|
||||
if (now >= w.fireAt) {
|
||||
this.enqueue(w.prompt);
|
||||
wakeDelete.push(w.id);
|
||||
}
|
||||
}
|
||||
for (const id of wakeDelete) this.wakeups.delete(id);
|
||||
}
|
||||
|
||||
create(input: { cron: string; prompt: string; recurring?: boolean; durable?: boolean }): CronJob {
|
||||
parseCron(input.cron); // validate syntax up front
|
||||
const recurring = input.recurring ?? true;
|
||||
const durable = input.durable ?? false;
|
||||
const job: CronJob = {
|
||||
id: this.nextId(),
|
||||
cron: input.cron,
|
||||
prompt: input.prompt,
|
||||
recurring,
|
||||
durable,
|
||||
createdAt: Date.now(),
|
||||
};
|
||||
this.jobs.set(job.id, job);
|
||||
if (durable) this.persist();
|
||||
return job;
|
||||
}
|
||||
|
||||
list(): CronJob[] {
|
||||
return [...this.jobs.values()].map((j) => ({ ...j }));
|
||||
}
|
||||
|
||||
get(id: string): CronJob | undefined {
|
||||
const j = this.jobs.get(id);
|
||||
return j ? { ...j } : undefined;
|
||||
}
|
||||
|
||||
delete(id: string): boolean {
|
||||
const existed = this.jobs.delete(id);
|
||||
if (existed) this.persist();
|
||||
return existed;
|
||||
}
|
||||
|
||||
/** Schedules a one-shot wakeup `delaySeconds` from now. If `stop` is true, cancels ALL wakeups
|
||||
* instead (used to end a self-paced loop). Returns the wakeup id, or `{ stopped: true }`. */
|
||||
scheduleWakeup(input: { delaySeconds: number; prompt: string; stop?: boolean }):
|
||||
| { id: string; fireAt: number }
|
||||
| { stopped: true } {
|
||||
if (input.stop) {
|
||||
this.wakeups.clear();
|
||||
return { stopped: true };
|
||||
}
|
||||
const delay = Math.max(60, Math.min(3600, input.delaySeconds));
|
||||
const id = this.nextWakeupId();
|
||||
const fireAt = Date.now() + delay * 1000;
|
||||
this.wakeups.set(id, { id, fireAt, prompt: input.prompt });
|
||||
return { id, fireAt };
|
||||
}
|
||||
|
||||
listWakeups(): Wakeup[] {
|
||||
return [...this.wakeups.values()].map((w) => ({ ...w }));
|
||||
}
|
||||
|
||||
private loadDurable(): void {
|
||||
if (!this.file || !existsSync(this.file)) return;
|
||||
try {
|
||||
const data = JSON.parse(readFileSync(this.file, "utf-8")) as { jobs?: CronJob[] };
|
||||
for (const j of data.jobs ?? []) {
|
||||
// Only durable jobs are persisted; skip any that slipped in without the flag.
|
||||
if (!j.durable) continue;
|
||||
this.jobs.set(j.id, { ...j });
|
||||
// Bump the seq past any restored ids so new ids don't collide.
|
||||
const n = Number(j.id.replace(/^c/, ""));
|
||||
if (Number.isFinite(n) && n > this.seq) this.seq = n;
|
||||
}
|
||||
} catch {
|
||||
// A corrupt persistence file shouldn't block startup — just start with no durable jobs.
|
||||
}
|
||||
}
|
||||
|
||||
private persist(): void {
|
||||
if (!this.file) return;
|
||||
try {
|
||||
const dir = path.dirname(this.file);
|
||||
if (!existsSync(dir)) mkdirSync(dir, { recursive: true });
|
||||
const durable = [...this.jobs.values()].filter((j) => j.durable).map((j) => ({
|
||||
id: j.id,
|
||||
cron: j.cron,
|
||||
prompt: j.prompt,
|
||||
recurring: j.recurring,
|
||||
durable: j.durable,
|
||||
createdAt: j.createdAt,
|
||||
}));
|
||||
writeFileSync(this.file, JSON.stringify({ jobs: durable }, null, 2), "utf-8");
|
||||
} catch {
|
||||
// Persistence is best-effort; a write failure must not crash the scheduler.
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1,78 +1,31 @@
|
||||
import { describe, it, expect } from "vitest";
|
||||
import { describe, expect, it } from "vitest";
|
||||
import { parseFallbackToolCalls } from "./fallbackParser.js";
|
||||
|
||||
describe("parseFallbackToolCalls", () => {
|
||||
it("parses the instructed ```tool_call fenced format", () => {
|
||||
const out = parseFallbackToolCalls(
|
||||
'```tool_call\n{"name":"read_file","arguments":{"path":"src/a.ts"}}\n```',
|
||||
it("parses a plain tool_call block", () => {
|
||||
const result = parseFallbackToolCalls("```tool_call\n{\"name\": \"read_file\", \"arguments\": {\"path\": \"src/x.ts\"}}\n```");
|
||||
expect(result.malformed).toBe(false);
|
||||
expect(result.calls).toHaveLength(1);
|
||||
expect(result.calls[0]).toEqual({ name: "read_file", arguments: { path: "src/x.ts" } });
|
||||
});
|
||||
|
||||
it("strips inner ```json fences", () => {
|
||||
const result = parseFallbackToolCalls(
|
||||
"```tool_call\n```json\n{\"name\": \"read_file\", \"arguments\": {\"path\": \"src/x.ts\"}}\n```\n```",
|
||||
);
|
||||
expect(out.calls).toEqual([{ name: "read_file", arguments: { path: "src/a.ts" } }]);
|
||||
expect(out.malformed).toBe(false);
|
||||
expect(result.calls).toHaveLength(1);
|
||||
expect(result.calls[0]).toEqual({ name: "read_file", arguments: { path: "src/x.ts" } });
|
||||
});
|
||||
|
||||
it("returns no calls and not-malformed for plain prose with no tool attempt", () => {
|
||||
const out = parseFallbackToolCalls("Here is the answer: use foo().");
|
||||
expect(out.calls).toEqual([]);
|
||||
expect(out.malformed).toBe(false);
|
||||
it("marks top-level arguments (not nested in 'arguments') as malformed", () => {
|
||||
const result = parseFallbackToolCalls("```tool_call\n{\"name\": \"read_file\", \"path\": \"src/x.ts\"}\n```");
|
||||
expect(result.malformed).toBe(true);
|
||||
expect(result.calls).toHaveLength(0);
|
||||
});
|
||||
|
||||
it("flags a tool_call block whose JSON is unparseable as malformed", () => {
|
||||
const out = parseFallbackToolCalls("```tool_call\n{not valid json\n```");
|
||||
expect(out.calls).toEqual([]);
|
||||
expect(out.malformed).toBe(true);
|
||||
it("marks invalid JSON as malformed", () => {
|
||||
const result = parseFallbackToolCalls("```tool_call\nnot json\n```");
|
||||
expect(result.malformed).toBe(true);
|
||||
expect(result.calls).toHaveLength(0);
|
||||
});
|
||||
|
||||
it("accepts a ```json fenced block when no tool_call block is present", () => {
|
||||
const out = parseFallbackToolCalls(
|
||||
'```json\n{"name":"grep","arguments":{"pattern":"foo"}}\n```',
|
||||
);
|
||||
expect(out.calls).toEqual([{ name: "grep", arguments: { pattern: "foo" } }]);
|
||||
expect(out.malformed).toBe(false);
|
||||
});
|
||||
|
||||
it("prefers a ```tool_call block over a ```json block when both appear", () => {
|
||||
const out = parseFallbackToolCalls(
|
||||
'```tool_call\n{"name":"read_file","arguments":{"path":"a"}}\n```\n' +
|
||||
'```json\n{"name":"grep","arguments":{"pattern":"x"}}\n```',
|
||||
);
|
||||
expect(out.calls).toHaveLength(1);
|
||||
expect(out.calls[0]!.name).toBe("read_file");
|
||||
});
|
||||
|
||||
it("does not treat a ```json fence inside a tool_call block as the terminator", () => {
|
||||
const out = parseFallbackToolCalls(
|
||||
"```tool_call\n```json\n{\"name\":\"read_file\",\"arguments\":{\"path\":\"a\"}}\n```\n```",
|
||||
);
|
||||
expect(out.calls).toEqual([{ name: "read_file", arguments: { path: "a" } }]);
|
||||
});
|
||||
|
||||
it("extracts a bare (unfenced) tool-call object from surrounding prose", () => {
|
||||
const out = parseFallbackToolCalls(
|
||||
'Let me read that file.\n{"name":"read_file","arguments":{"path":"src/loop.ts"}}\nThat should help.',
|
||||
);
|
||||
expect(out.calls).toEqual([{ name: "read_file", arguments: { path: "src/loop.ts" } }]);
|
||||
});
|
||||
|
||||
it("ignores bare braces in prose that don't look like a tool call", () => {
|
||||
const out = parseFallbackToolCalls("The config is { key: value } and that's it.");
|
||||
expect(out.calls).toEqual([]);
|
||||
expect(out.malformed).toBe(false);
|
||||
});
|
||||
|
||||
it("repairs a truncated JSON tool_call block via partial-JSON repair", () => {
|
||||
// max_tokens clipped the closing brace and quote
|
||||
const out = parseFallbackToolCalls('```tool_call\n{"name":"read_file","arguments":{"path":"src/lo');
|
||||
expect(out.calls).toEqual([{ name: "read_file", arguments: { path: "src/lo" } }]);
|
||||
});
|
||||
|
||||
it("accepts a tool call with omitted arguments as empty arguments", () => {
|
||||
const out = parseFallbackToolCalls('```tool_call\n{"name":"git_status"}\n```');
|
||||
expect(out.calls).toEqual([{ name: "git_status", arguments: {} }]);
|
||||
});
|
||||
|
||||
it("flags a tool call whose arguments is not an object", () => {
|
||||
const out = parseFallbackToolCalls('```tool_call\n{"name":"x","arguments":"foo"}\n```');
|
||||
expect(out.calls).toEqual([]);
|
||||
expect(out.malformed).toBe(true);
|
||||
});
|
||||
});
|
||||
});
|
||||
|
||||
@@ -1,5 +1,3 @@
|
||||
import { repairPartialJson } from "./partialJson.js";
|
||||
|
||||
export interface FallbackToolCall {
|
||||
name: string;
|
||||
arguments: Record<string, unknown>;
|
||||
@@ -10,155 +8,30 @@ export interface FallbackParseResult {
|
||||
malformed: boolean;
|
||||
}
|
||||
|
||||
// Local models in fallback mode (no native function calling) are asked to emit tool calls as a
|
||||
// fenced ```tool_call block. In practice they frequently deviate, so the parser is lenient about
|
||||
// FORMAT but strict about CONTENT: anything we extract must still parse to { name, arguments }.
|
||||
// Accepted shapes, in priority order:
|
||||
// 1. A ```tool_call fenced block (the instructed format). The opening fence is ```tool_call on
|
||||
// its own line; the closing fence is ``` on its own line — anchored so a ```json block INSIDE
|
||||
// isn't mistaken for the terminator. Some models wrap the JSON in an inner ```json fence; we
|
||||
// strip that inner fence before parsing.
|
||||
// 2. A ```json fenced block whose content is a tool-call object (name + arguments). Models that
|
||||
// ignore the custom "tool_call" fence name but reach for the familiar "json" one.
|
||||
// 3. A bare tool-call object appearing in the response with no fence at all. We scan for the
|
||||
// first balanced {...} that contains a string "name" and an object "arguments". To avoid
|
||||
// matching arbitrary prose-embedded JSON, we require the recognizable keys.
|
||||
//
|
||||
// Every extracted candidate goes through `coerce`, which parses (with partial-JSON repair as a
|
||||
// last resort for truncated streaming output) and validates the name/arguments shape. A candidate
|
||||
// that doesn't yield a valid call sets `malformed` — the loop nudges the model to retry rather than
|
||||
// silently ending the task with no tool executed.
|
||||
|
||||
// Opening fence ```tool_call on its own line; closing ``` on its own line. Multiline-anchored so
|
||||
// an inner ```json fence can't be read as the terminator.
|
||||
const TOOL_CALL_FENCE_RE = /^```tool_call\s*\n([\s\S]*?)\n```(?:\n|$)/gm;
|
||||
// A ```json block — only used if no ```tool_call block matched, since a json fence may carry prose.
|
||||
const JSON_FENCE_RE = /^```json\s*\n([\s\S]*?)\n```(?:\n|$)/gm;
|
||||
|
||||
/** Pull the textual content out of a fenced block, stripping any inner ```json fence a model may
|
||||
* have nested inside it. Returns the cleaned, trimmed body. */
|
||||
function cleanFencedBody(raw: string): string {
|
||||
return raw.replace(/^```(?:json)?\s*|\s*```$/g, "").trim();
|
||||
}
|
||||
|
||||
/** Parse a candidate string into a tool call, or null if it isn't one. Tries a direct parse, then a
|
||||
* partial-JSON repair (truncation / trailing comma / unbalanced braces) for clipped streaming. */
|
||||
function coerce(candidate: string): FallbackToolCall | null {
|
||||
const trimmed = candidate.trim();
|
||||
if (!trimmed) return null;
|
||||
// Direct parse first.
|
||||
let parsed: unknown = null;
|
||||
try {
|
||||
parsed = JSON.parse(trimmed);
|
||||
} catch {
|
||||
parsed = null;
|
||||
}
|
||||
// Repair pass for truncated / sloppy JSON from local streaming.
|
||||
if (parsed === null) parsed = repairPartialJson(trimmed);
|
||||
if (!parsed || typeof parsed !== "object") return null;
|
||||
const obj = parsed as Record<string, unknown>;
|
||||
if (typeof obj.name !== "string") return null;
|
||||
if (obj.arguments === undefined || obj.arguments === null) {
|
||||
// Some models omit arguments entirely when the tool takes none — treat as empty.
|
||||
return { name: obj.name, arguments: {} };
|
||||
}
|
||||
if (typeof obj.arguments !== "object" || Array.isArray(obj.arguments)) return null;
|
||||
return { name: obj.name, arguments: obj.arguments as Record<string, unknown> };
|
||||
}
|
||||
|
||||
/** Scan `content` for the first balanced {...} object containing a `"name"` string and an
|
||||
* `"arguments"` object, with no fence at all. Tracks string literals so braces inside strings
|
||||
* don't affect nesting, and cuts to the first complete top-level object. */
|
||||
function findBareToolCall(content: string): string | null {
|
||||
const start = content.indexOf("{");
|
||||
if (start < 0) return null;
|
||||
let depth = 0;
|
||||
let inStr = false;
|
||||
for (let i = start; i < content.length; i++) {
|
||||
const c = content[i]!;
|
||||
if (inStr) {
|
||||
if (c === "\\") {
|
||||
i++;
|
||||
continue;
|
||||
}
|
||||
if (c === '"') inStr = false;
|
||||
continue;
|
||||
}
|
||||
if (c === '"') {
|
||||
inStr = true;
|
||||
continue;
|
||||
}
|
||||
if (c === "{") depth++;
|
||||
else if (c === "}") {
|
||||
depth--;
|
||||
if (depth === 0) {
|
||||
const candidate = content.slice(start, i + 1);
|
||||
// Only accept it if it actually looks like a tool call — otherwise keep scanning.
|
||||
if (/"name"\s*:/.test(candidate) && /"arguments"\s*:/.test(candidate)) {
|
||||
return candidate;
|
||||
}
|
||||
// Reset to the next brace past this point to keep looking.
|
||||
const next = content.indexOf("{", i + 1);
|
||||
if (next < 0) return null;
|
||||
i = next - 1;
|
||||
depth = 0;
|
||||
}
|
||||
}
|
||||
}
|
||||
// Ran off the end with an unclosed object — a truncated tool call (max_tokens clipped the
|
||||
// closing brace, and possibly the closing fence too). Hand the unbalanced substring back;
|
||||
// `coerce` will run it through partial-JSON repair to close what's open.
|
||||
if (depth > 0) {
|
||||
const candidate = content.slice(start);
|
||||
if (/"name"\s*:/.test(candidate) && /"arguments"\s*:/.test(candidate)) {
|
||||
return candidate;
|
||||
}
|
||||
}
|
||||
return null;
|
||||
}
|
||||
// Opening fence is ```tool_call on its own line; closing fence is ``` on its own line.
|
||||
// This prevents ```json inside the block from being mistaken for the terminator.
|
||||
const BLOCK_RE = /^```tool_call\s*\n([\s\S]*?)\n```(?:\n|$)/gm;
|
||||
|
||||
export function parseFallbackToolCalls(content: string): FallbackParseResult {
|
||||
const calls: FallbackToolCall[] = [];
|
||||
let malformed = false;
|
||||
let sawAnyCandidate = false;
|
||||
|
||||
// (1) ```tool_call fenced blocks (the instructed format) — there may be several.
|
||||
for (const match of content.matchAll(TOOL_CALL_FENCE_RE)) {
|
||||
sawAnyCandidate = true;
|
||||
const body = cleanFencedBody(match[1] ?? "");
|
||||
const call = coerce(body);
|
||||
if (call) calls.push(call);
|
||||
else malformed = true;
|
||||
}
|
||||
if (calls.length > 0) return { calls, malformed };
|
||||
|
||||
// (2) ```json fenced blocks — only if no tool_call block matched. Take the first that coerces.
|
||||
for (const match of content.matchAll(JSON_FENCE_RE)) {
|
||||
const call = coerce(cleanFencedBody(match[1] ?? ""));
|
||||
if (call) {
|
||||
calls.push(call);
|
||||
return { calls, malformed: false };
|
||||
for (const match of content.matchAll(BLOCK_RE)) {
|
||||
const raw = match[1]?.trim() ?? "";
|
||||
// Some local models emit markdown fences inside the tool_call block (e.g. ```json ... ```).
|
||||
// Strip them so the inner JSON can be parsed.
|
||||
const cleaned = raw.replace(/^```(?:json)?\s*|\s*```$/g, "").trim();
|
||||
try {
|
||||
const parsed = JSON.parse(cleaned || raw);
|
||||
if (parsed && typeof parsed.name === "string" && typeof parsed.arguments === "object" && parsed.arguments !== null) {
|
||||
calls.push({ name: parsed.name, arguments: parsed.arguments });
|
||||
} else {
|
||||
malformed = true;
|
||||
}
|
||||
} catch {
|
||||
malformed = true;
|
||||
}
|
||||
sawAnyCandidate = true;
|
||||
malformed = true;
|
||||
}
|
||||
if (calls.length > 0) return { calls, malformed };
|
||||
|
||||
// (3) Bare (unfenced) tool-call object — last resort. At most one: the prompt says one call per
|
||||
// response, and extracting multiple bare objects from prose is too error-prone.
|
||||
const bare = findBareToolCall(content);
|
||||
if (bare !== null) {
|
||||
sawAnyCandidate = true;
|
||||
const call = coerce(bare);
|
||||
if (call) {
|
||||
calls.push(call);
|
||||
return { calls, malformed: false };
|
||||
}
|
||||
malformed = true;
|
||||
}
|
||||
|
||||
// If we never saw anything that even looked like a tool-call attempt, that's not malformed —
|
||||
// the model simply answered in prose (no tool needed). `malformed` stays false.
|
||||
void sawAnyCandidate;
|
||||
return { calls, malformed };
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,14 +1,24 @@
|
||||
export const FALLBACK_TOOL_INSTRUCTIONS = `This model does not support native function calling. To call a tool, write a fenced code block:
|
||||
export const FALLBACK_TOOL_INSTRUCTIONS = `This model does not support native function calling. To call a tool, write exactly one fenced code block of the form:
|
||||
|
||||
\`\`\`tool_call
|
||||
{"name": "read_file", "arguments": {"path": "src/index.ts"}}
|
||||
{"name": "TOOL_NAME", "arguments": {"arg1": "value1", "arg2": "value2"}}
|
||||
\`\`\`
|
||||
|
||||
Rules:
|
||||
- One tool call per response. Wait for the result before calling another.
|
||||
- Must contain valid JSON with "name" and "arguments" keys.
|
||||
- If no tool is needed, answer normally without a fenced block.
|
||||
- One tool call per response. Wait for the \`\`\`tool_result\`\`\` before calling another.
|
||||
- The fenced block must contain a single JSON object with exactly two keys: "name" and "arguments".
|
||||
- "arguments" must be an object matching the tool's schema. Do not put the arguments at the top level.
|
||||
- If no tool is needed, answer normally without any \`\`\`tool_call\`\`\` block.
|
||||
|
||||
The result is returned in a \`\`\`tool_result\`\`\` block. Then answer normally or call another tool.`;
|
||||
Example:
|
||||
\`\`\`tool_call
|
||||
{"name": "read_file", "arguments": {"path": "src/index.ts", "limit": 50}}
|
||||
\`\`\`
|
||||
|
||||
export const FALLBACK_RETRY_NUDGE = `Your last \`tool_call\` block wasn't valid JSON with "name" and "arguments" keys. Try again using the correct format, or answer without a tool call.`;
|
||||
When a tool result shows an error or empty output, do not repeat the exact same call. Adjust your arguments or ask the user.`;
|
||||
|
||||
export const FALLBACK_RETRY_NUDGE = `Your last \`tool_call\` block was invalid. Check:
|
||||
- It must be a single JSON object inside the fence, not plain text or multiple objects.
|
||||
- It must have "name" (string) and "arguments" (object) keys.
|
||||
- Argument values must match the tool's expected types.
|
||||
Try again with the correct format, or answer without a tool call.`;
|
||||
|
||||
@@ -1,45 +0,0 @@
|
||||
import { describe, expect, it } from "vitest";
|
||||
import { z } from "zod";
|
||||
import { resolveToolCall } from "./nativeAdapter.js";
|
||||
import type { ToolDef } from "../tools/types.js";
|
||||
|
||||
function makeRegistry(): Map<string, ToolDef> {
|
||||
const tool: ToolDef = {
|
||||
name: "write_file",
|
||||
description: "d",
|
||||
schema: z.object({ path: z.string(), content: z.string() }),
|
||||
mutating: true,
|
||||
handler: async () => ({}),
|
||||
};
|
||||
return new Map([["write_file", tool]]);
|
||||
}
|
||||
|
||||
function call(args: string) {
|
||||
return {
|
||||
id: "1",
|
||||
type: "function" as const,
|
||||
function: { name: "write_file", arguments: args },
|
||||
};
|
||||
}
|
||||
|
||||
describe("resolveToolCall — argument repair", () => {
|
||||
it("resolves normally when arguments are already valid JSON", () => {
|
||||
const r = resolveToolCall(call('{"path":"a.ts","content":"hi"}'), makeRegistry());
|
||||
expect("tool" in r).toBe(true);
|
||||
});
|
||||
|
||||
it("repairs arguments containing a raw (unescaped) newline instead of failing outright", () => {
|
||||
// A local model echoing multi-line file content into a non-streaming completion often pastes
|
||||
// a real newline byte into the JSON string rather than escaping it — this path (unlike the
|
||||
// streaming path in loop.ts) previously had no repair attempt at all.
|
||||
const args = '{"path":"a.ts","content":"line1\nline2"}';
|
||||
const r = resolveToolCall(call(args), makeRegistry());
|
||||
expect("tool" in r).toBe(true);
|
||||
if ("tool" in r) expect((r.args as { content: string }).content).toBe("line1\nline2");
|
||||
});
|
||||
|
||||
it("still errors when the arguments are unsalvageable", () => {
|
||||
const r = resolveToolCall(call("not json at all"), makeRegistry());
|
||||
expect(r).toEqual({ error: "arguments were not valid JSON" });
|
||||
});
|
||||
});
|
||||
@@ -1,7 +1,6 @@
|
||||
import type { ChatCompletionMessageToolCall, ChatCompletionTool } from "openai/resources/chat/completions";
|
||||
import { z } from "zod";
|
||||
import type { ToolDef } from "../tools/types.js";
|
||||
import { repairPartialJson } from "./partialJson.js";
|
||||
import { resolveToolInvocation, type ResolvedToolCall } from "./resolve.js";
|
||||
|
||||
export function toOpenAITools(tools: ToolDef[]): ChatCompletionTool[] {
|
||||
@@ -26,13 +25,7 @@ export function resolveToolCall(
|
||||
try {
|
||||
parsedArgs = JSON.parse(call.function.arguments || "{}");
|
||||
} catch {
|
||||
// Non-streaming completions land here directly (unlike the streaming path in loop.ts, which
|
||||
// already repairs before this point) — without a repair attempt here too, a call whose
|
||||
// arguments contain e.g. an unescaped literal newline (a local model echoing multi-line file
|
||||
// content raw) fails outright instead of being salvaged.
|
||||
const repaired = repairPartialJson(call.function.arguments || "");
|
||||
if (repaired === null) return { error: "arguments were not valid JSON" };
|
||||
parsedArgs = repaired;
|
||||
return { error: "arguments were not valid JSON" };
|
||||
}
|
||||
return resolveToolInvocation(call.function.name, parsedArgs, registry);
|
||||
}
|
||||
|
||||
@@ -1,83 +0,0 @@
|
||||
import { describe, it, expect } from "vitest";
|
||||
import { repairPartialJson } from "./partialJson.js";
|
||||
|
||||
describe("repairPartialJson", () => {
|
||||
it("parses already-valid JSON unchanged", () => {
|
||||
expect(repairPartialJson('{"path":"src/a.ts"}')).toEqual({ path: "src/a.ts" });
|
||||
expect(repairPartialJson("[]")).toEqual([]);
|
||||
expect(repairPartialJson(" 42 ")).toBe(42);
|
||||
expect(repairPartialJson('{"a":1}\n')).toEqual({ a: 1 });
|
||||
});
|
||||
|
||||
it("returns null for empty / whitespace input", () => {
|
||||
expect(repairPartialJson("")).toBeNull();
|
||||
expect(repairPartialJson(" ")).toBeNull();
|
||||
});
|
||||
|
||||
it("strips a trailing comma before a closing bracket", () => {
|
||||
expect(repairPartialJson('{"a":1,}')).toEqual({ a: 1 });
|
||||
expect(repairPartialJson('[1,2,]')).toEqual([1, 2]);
|
||||
expect(repairPartialJson('{"a":{"b":2,},}')).toEqual({ a: { b: 2 } });
|
||||
});
|
||||
|
||||
it("strips stray trailing content after a complete value", () => {
|
||||
expect(repairPartialJson('{"a":1}\n```')).toEqual({ a: 1 });
|
||||
expect(repairPartialJson('{"a":1}garbage')).toEqual({ a: 1 });
|
||||
expect(repairPartialJson('[1,2] }')).toEqual([1, 2]);
|
||||
});
|
||||
|
||||
it("closes a truncated string value", () => {
|
||||
// max_tokens clipped mid-value: {"path":"src/lo → needs closing quote + brace
|
||||
expect(repairPartialJson('{"path":"src/lo')).toEqual({ path: "src/lo" });
|
||||
expect(repairPartialJson('{"a":"hello wor')).toEqual({ a: "hello wor" });
|
||||
});
|
||||
|
||||
it("balances unclosed braces and brackets from truncation", () => {
|
||||
expect(repairPartialJson('{"a":1')).toEqual({ a: 1 });
|
||||
expect(repairPartialJson('{"a":{"b":2')).toEqual({ a: { b: 2 } });
|
||||
expect(repairPartialJson("[1,2")).toEqual([1, 2]);
|
||||
expect(repairPartialJson('{"items":[1,2')).toEqual({ items: [1, 2] });
|
||||
});
|
||||
|
||||
it("does not count braces inside string literals", () => {
|
||||
// The braces/brackets inside the string are content, not nesting.
|
||||
expect(repairPartialJson('{"code":"func() { return [1] "')).toEqual({
|
||||
code: "func() { return [1] ",
|
||||
});
|
||||
expect(repairPartialJson('{"s":"\\\"escaped\\\""}')).toEqual({ s: '"escaped"' });
|
||||
});
|
||||
|
||||
it("handles escaped quotes inside strings during truncation repair", () => {
|
||||
// Unterminated string with an escaped quote inside: {"s":"a\"b
|
||||
expect(repairPartialJson('{"s":"a\\"b')).toEqual({ s: 'a"b' });
|
||||
});
|
||||
|
||||
it("combined: trailing comma exposed after balancing", () => {
|
||||
// {"a":1,"b":2, (truncated with trailing comma) → close brace, then strip comma
|
||||
expect(repairPartialJson('{"a":1,"b":2,')).toEqual({ a: 1, b: 2 });
|
||||
});
|
||||
|
||||
it("returns null when input is not salvageable as object/array/scalar", () => {
|
||||
expect(repairPartialJson("just prose with no json")).toBeNull();
|
||||
expect(repairPartialJson("{:}")).toBeNull();
|
||||
});
|
||||
|
||||
it("escapes a raw (unescaped) literal newline inside a string value", () => {
|
||||
// A local model echoing multi-line file content often pastes real \n bytes into the JSON
|
||||
// string instead of writing the two-char `\n` escape — JSON.parse rejects that outright.
|
||||
expect(repairPartialJson('{"content":"line1\nline2"}')).toEqual({ content: "line1\nline2" });
|
||||
});
|
||||
|
||||
it("escapes a raw carriage return inside a string value (CRLF source content)", () => {
|
||||
expect(repairPartialJson('{"content":"line1\r\nline2"}')).toEqual({ content: "line1\r\nline2" });
|
||||
});
|
||||
|
||||
it("leaves an already-escaped \\n sequence untouched", () => {
|
||||
expect(repairPartialJson('{"content":"line1\\nline2"}')).toEqual({ content: "line1\nline2" });
|
||||
});
|
||||
|
||||
it("combines raw-newline escaping with truncation repair", () => {
|
||||
// Truncated mid-value AND containing a raw newline earlier in the string.
|
||||
expect(repairPartialJson('{"content":"line1\nline2')).toEqual({ content: "line1\nline2" });
|
||||
});
|
||||
});
|
||||
@@ -1,252 +0,0 @@
|
||||
// Partial/truncated JSON repair for native streaming tool-call arguments.
|
||||
//
|
||||
// Local-model backends (Ollama, LM Studio) streaming tool calls accumulate the `arguments` string
|
||||
// across deltas. Two common pathologies produce a string that JSON.parse rejects but that contains
|
||||
// all the semantic content the model intended:
|
||||
//
|
||||
// 1. TRUNCATION — max_tokens clipped the JSON mid-value. The string ends inside a string value,
|
||||
// an array, or an object: `{"path": "src/lo`, `{"items": [1, 2`, `{"a": {"b": 1`.
|
||||
// 2. LOCAL-MODEL SLOPPINESS — a trailing comma, an unbalanced brace/bracket, or a trailing
|
||||
// garbage token after the closing brace: `{"path": "x.ts",}`, `{"a": 1 `.
|
||||
//
|
||||
// This module attempts a cheap, conservative repair BEFORE the caller falls back to a full
|
||||
// non-streaming regeneration (which is expensive on a local backend and often fails identically
|
||||
// when the cause was max_tokens). It only closes what's open and trims what's stray — it never
|
||||
// invents keys or values, so a genuinely malformed call still fails downstream at schema validation.
|
||||
//
|
||||
// The repair is best-effort: if it can't produce parseable JSON, it returns null and the caller
|
||||
// keeps its existing retry path. It is deliberately string-based (no AST) so it's trivially fast and
|
||||
// has no dependencies, and so it handles truncated input that a strict parser can't even build an
|
||||
// AST from.
|
||||
|
||||
/** Parse `s` as JSON; on success return the value, on failure return null (never throws). */
|
||||
function tryParse(s: string): unknown {
|
||||
try {
|
||||
return JSON.parse(s);
|
||||
} catch {
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
/** JSON disallows raw control characters (0x00-0x1F — notably literal newline, CR, tab) inside
|
||||
* string literals; they must be written as `\n`/`\r`/`\t`/`\u00XX`. Local models echoing
|
||||
* multi-line file content (very common for write_file/edit_file on this project's CRLF-heavy
|
||||
* source) routinely paste it in unescaped, which makes an otherwise complete, semantically
|
||||
* correct tool call fail JSON.parse with "Bad control character in string literal". Escaping
|
||||
* only touches raw bytes found *inside* a string (tracked the same way `skipString` does, so an
|
||||
* existing `\\n` escape sequence is left alone) — it never changes where a string starts/ends or
|
||||
* where a brace/bracket falls outside one, so it's safe to run before the other repair steps. */
|
||||
function escapeRawControlCharsInStrings(s: string): string {
|
||||
let out = "";
|
||||
let inStr = false;
|
||||
let changed = false;
|
||||
for (let i = 0; i < s.length; i++) {
|
||||
const c = s[i]!;
|
||||
if (!inStr) {
|
||||
if (c === '"') inStr = true;
|
||||
out += c;
|
||||
continue;
|
||||
}
|
||||
if (c === "\\") {
|
||||
// Preserve an existing escape sequence verbatim — don't touch the char after the backslash.
|
||||
out += c + (s[i + 1] ?? "");
|
||||
i++;
|
||||
continue;
|
||||
}
|
||||
if (c === '"') {
|
||||
inStr = false;
|
||||
out += c;
|
||||
continue;
|
||||
}
|
||||
const code = c.charCodeAt(0);
|
||||
if (code < 0x20) {
|
||||
changed = true;
|
||||
if (c === "\n") out += "\\n";
|
||||
else if (c === "\r") out += "\\r";
|
||||
else if (c === "\t") out += "\\t";
|
||||
else out += "\\u" + code.toString(16).padStart(4, "0");
|
||||
continue;
|
||||
}
|
||||
out += c;
|
||||
}
|
||||
return changed ? out : s;
|
||||
}
|
||||
|
||||
/** Skip past the next JSON string literal starting at `i` (the opening quote). Returns the index
|
||||
* just past the closing quote. Strings are the only place braces/brackets can appear without
|
||||
* affecting nesting, so we must not count them while inside one. Handles `\"` and other escapes. */
|
||||
function skipString(s: string, i: number): number {
|
||||
let j = i + 1; // past opening quote
|
||||
for (; j < s.length; j++) {
|
||||
const c = s[j]!;
|
||||
if (c === "\\") {
|
||||
j++; // skip the escaped char (covers \", \\, etc.)
|
||||
continue;
|
||||
}
|
||||
if (c === '"') return j + 1; // past closing quote
|
||||
}
|
||||
return j; // ran off the end — unterminated string
|
||||
}
|
||||
|
||||
/** Attempts to repair `raw` into parseable JSON. Returns the parsed value on success, or null if no
|
||||
* repair produced valid JSON. Steps, applied in order of how cheap and safe they are:
|
||||
*
|
||||
* 1. Maybe it already parses (trailing whitespace/newlines are fine for JSON.parse) — return as-is.
|
||||
* 2. Escape raw control characters (literal newline/CR/tab) found inside string literals — a
|
||||
* model echoing multi-line file content unescaped, common on this project's CRLF sources.
|
||||
* 3. Strip a trailing comma before an expected-but-absent `}` or `]` (common local-model slip).
|
||||
* 4. Strip stray non-JSON tokens after the first complete top-level value (`{"a":1}\n` → `{"a":1}`,
|
||||
* and `{"a":1}garbage` → `{"a":1}` — JSON.parse rejects trailing content, so trim to the first
|
||||
* complete value).
|
||||
* 5. Close unterminated strings, then balance still-open braces/brackets (truncation repair).
|
||||
*
|
||||
* Each step re-attempts a parse, so the cheapest fix that works wins. */
|
||||
export function repairPartialJson(raw: string): unknown | null {
|
||||
if (!raw) return null;
|
||||
const trimmed = raw.trim();
|
||||
if (!trimmed) return null;
|
||||
|
||||
// (1) Already valid?
|
||||
const direct = tryParse(trimmed);
|
||||
if (direct !== null) return direct;
|
||||
|
||||
// (2) Raw control characters (literal newlines/CR/tab) inside string literals — see
|
||||
// escapeRawControlCharsInStrings for why this is common. Escaping never moves a quote or
|
||||
// brace, so every remaining step below runs against this version instead of the original.
|
||||
const working = escapeRawControlCharsInStrings(trimmed);
|
||||
if (working !== trimmed) {
|
||||
const v = tryParse(working);
|
||||
if (v !== null) return v;
|
||||
}
|
||||
|
||||
// (3) Trailing comma before end-of-object/array: `{"a":1,}` or `[1,2,]`. Repeat until none
|
||||
// left so a nested shape like `{"a":{"b":2,},}` clears both commas (innermost-first).
|
||||
let noTrailing = working;
|
||||
let prev: string;
|
||||
do {
|
||||
prev = noTrailing;
|
||||
noTrailing = noTrailing.replace(/,\s*([\]}]+\s*$)/, "$1");
|
||||
} while (noTrailing !== prev);
|
||||
if (noTrailing !== working) {
|
||||
const v = tryParse(noTrailing);
|
||||
if (v !== null) return v;
|
||||
}
|
||||
|
||||
// (4) Stray trailing content after the first complete value. JSON.parse refuses trailing tokens,
|
||||
// but a model often emits a closing brace then a stray newline, a repeated token, or prose.
|
||||
// Find the end of the first balanced top-level value and cut there.
|
||||
const cut = cutToFirstCompleteValue(working);
|
||||
if (cut !== null && cut !== working) {
|
||||
const v = tryParse(cut);
|
||||
if (v !== null) return v;
|
||||
}
|
||||
|
||||
// (5) Truncation repair: close an unterminated string, then balance open braces/brackets.
|
||||
const balanced = balanceAndClose(working);
|
||||
if (balanced !== null && balanced !== working) {
|
||||
// Re-run the earlier cheap fixes on the balanced result (a trailing comma may now be exposed).
|
||||
const v = tryParse(balanced);
|
||||
if (v !== null) return v;
|
||||
const v2 = tryParse(balanced.replace(/,\s*([\]}]\s*$)/, "$1"));
|
||||
if (v2 !== null) return v2;
|
||||
}
|
||||
|
||||
return null;
|
||||
}
|
||||
|
||||
/** If `s` starts with a complete top-level JSON value followed by stray content, return just that
|
||||
* value (as a substring). Returns null if we can't find a clean boundary (e.g. the value is itself
|
||||
* truncated). Walks the string tracking string literals and nesting depth. */
|
||||
function cutToFirstCompleteValue(s: string): string | null {
|
||||
let i = 0;
|
||||
// Skip leading whitespace.
|
||||
while (i < s.length && /\s/.test(s[i]!)) i++;
|
||||
if (i >= s.length) return null;
|
||||
|
||||
const start = i;
|
||||
const stack: string[] = [];
|
||||
let inStr = false;
|
||||
|
||||
while (i < s.length) {
|
||||
const c = s[i]!;
|
||||
if (inStr) {
|
||||
if (c === "\\") {
|
||||
i += 2;
|
||||
continue;
|
||||
}
|
||||
if (c === '"') inStr = false;
|
||||
i++;
|
||||
continue;
|
||||
}
|
||||
if (c === '"') {
|
||||
inStr = true;
|
||||
i++;
|
||||
continue;
|
||||
}
|
||||
if (c === "{" || c === "[") {
|
||||
stack.push(c);
|
||||
i++;
|
||||
continue;
|
||||
}
|
||||
if (c === "}" || c === "]") {
|
||||
stack.pop();
|
||||
i++;
|
||||
// If the stack is empty, this was the end of the top-level value — cut here.
|
||||
if (stack.length === 0) return s.slice(start, i);
|
||||
continue;
|
||||
}
|
||||
// A bare scalar (number/true/false/null) ends at the next delimiter/comma/whitespace.
|
||||
if (stack.length === 0 && (c === "," || c === "}" || c === "]" || /\s/.test(c))) {
|
||||
return s.slice(start, i);
|
||||
}
|
||||
i++;
|
||||
}
|
||||
|
||||
// Ran off the end without closing the top-level value → it's truncated, not "complete + stray".
|
||||
if (stack.length > 0) return null;
|
||||
return null;
|
||||
}
|
||||
|
||||
/** Closes an unterminated trailing string and balances any open braces/brackets. Returns the
|
||||
* repaired string, or null if nothing needed closing (caller can compare to skip a no-op parse). */
|
||||
function balanceAndClose(s: string): string | null {
|
||||
let out = s;
|
||||
const stack: string[] = [];
|
||||
let inStr = false;
|
||||
let i = 0;
|
||||
|
||||
for (; i < out.length; i++) {
|
||||
const c = out[i]!;
|
||||
if (inStr) {
|
||||
if (c === "\\") {
|
||||
i++;
|
||||
continue;
|
||||
}
|
||||
if (c === '"') inStr = false;
|
||||
continue;
|
||||
}
|
||||
if (c === '"') {
|
||||
inStr = true;
|
||||
continue;
|
||||
}
|
||||
if (c === "{") stack.push("}");
|
||||
else if (c === "[") stack.push("]");
|
||||
else if (c === "}" || c === "]") stack.pop();
|
||||
}
|
||||
|
||||
// If we ended inside a string, close it. A truncated value like `{"path":"src/lo` needs a closing
|
||||
// quote before we can balance the outer braces.
|
||||
if (inStr) {
|
||||
out += '"';
|
||||
}
|
||||
|
||||
// Close anything still open, innermost-first. Truncation mid-array/object → append the closers.
|
||||
if (stack.length === 0 && !inStr) {
|
||||
// Nothing to close — but a trailing comma may have been the only issue; let the caller handle it.
|
||||
return inStr ? out : null;
|
||||
}
|
||||
while (stack.length) {
|
||||
out += stack.pop();
|
||||
}
|
||||
return out;
|
||||
}
|
||||
+64
-48
@@ -1,5 +1,6 @@
|
||||
import { z } from "zod";
|
||||
import type { SubAgentResult } from "./types.js";
|
||||
import { AGENT_TYPE_NAMES, getAgentType } from "./agentTypes.js";
|
||||
import type { SubAgentOverrides } from "./types.js";
|
||||
import type { ToolDef } from "./types.js";
|
||||
|
||||
const schema = z.object({
|
||||
@@ -8,64 +9,79 @@ const schema = z.object({
|
||||
.string()
|
||||
.describe(
|
||||
"Full, self-contained task description. The sub-agent has no conversation memory and cannot ask follow-ups.",
|
||||
)
|
||||
.optional(),
|
||||
tasks: z
|
||||
.array(
|
||||
z.object({
|
||||
description: z.string().describe("Short label for this sub-task."),
|
||||
prompt: z
|
||||
.string()
|
||||
.describe(
|
||||
"Full, self-contained task description for this sub-task. The sub-agent has no conversation memory.",
|
||||
),
|
||||
}),
|
||||
)
|
||||
.min(2)
|
||||
),
|
||||
agentType: z
|
||||
.enum(AGENT_TYPE_NAMES)
|
||||
.optional()
|
||||
.describe(
|
||||
"Run several sub-agents IN PARALLEL (read-heavy research/audit tasks). Each gets its own " +
|
||||
"isolated context and returns independently. Use this for many files, e.g. one sub-agent " +
|
||||
"per directory or per concern (security, performance, tests). Prefer `prompt` for a single task.",
|
||||
)
|
||||
.optional(),
|
||||
}).refine((v) => v.prompt || v.tasks, {
|
||||
message: "Provide either `prompt` (single sub-agent) or `tasks` (parallel batch).",
|
||||
"Specialist sub-agent type. 'general-purpose' (default) has full tool access and may edit files. " +
|
||||
"'explore' is read-only research (locate code, map structure). 'code-reviewer' is read-only review " +
|
||||
"(find bugs, verify claims, report findings). 'planner' is read-only planning (design an implementation " +
|
||||
"plan with steps and tradeoffs). 'debugger' reproduces and fixes a bug with a minimal, verified fix. " +
|
||||
"'test-writer' writes focused tests, runs them, and iterates until they pass. Read-only types " +
|
||||
"(explore, code-reviewer, planner) cannot modify files.",
|
||||
),
|
||||
name: z
|
||||
.string()
|
||||
.optional()
|
||||
.describe(
|
||||
"Optional name for a shared-cwd (single, non-parallel) delegation. Naming it creates a named TEAMMATE: " +
|
||||
"you can then continue it with send_message by name and see it via list_teammates, without tracking its " +
|
||||
"agentId. Use this when you'll send a teammate several messages (e.g. a 'researcher' you'll re-query). " +
|
||||
"Names must be unique within your session — re-using an existing name returns an error instead of " +
|
||||
"clobbering the teammate. Parallel (worktree-isolated) delegations are fire-and-forget and ignore the name.",
|
||||
),
|
||||
});
|
||||
|
||||
export const agentTool: ToolDef<z.infer<typeof schema>> = {
|
||||
name: "agent",
|
||||
description:
|
||||
"Delegate a task to a sub-agent with its own tool loop (no nested agents). Only the final answer is returned. " +
|
||||
"For many files, split into multiple sub-agents — pass a `tasks` array to run several in PARALLEL " +
|
||||
"(each returns independently; one failure doesn't discard the others). Sub-agents have a smaller " +
|
||||
"step budget; if one runs out, narrow the task rather than retrying. Mutating tool calls from any " +
|
||||
"sub-agent still go through the same permission prompts (serialized, so parallel sub-agents that " +
|
||||
"edit don't race).",
|
||||
"For many files, split into multiple sub-agents. Sub-agents have a smaller step budget; if one runs out, " +
|
||||
"narrow the task rather than retrying. Pick an agentType: general-purpose (full access, may edit), explore " +
|
||||
"(read-only research), code-reviewer (read-only review), planner (read-only design of an implementation plan), " +
|
||||
"debugger (reproduce and fix a bug with a minimal verified fix), or test-writer (write and run focused tests). " +
|
||||
"A single delegation runs in your working directory " +
|
||||
"and persists its edits; when you emit several 'agent'/'agent__*' calls in one response they run in parallel, " +
|
||||
"each in an isolated throwaway git worktree whose file changes are discarded — use parallel batches for " +
|
||||
"research/review/analysis (the returned answer is the deliverable), and a single call for implementation. " +
|
||||
"A single (non-parallel) delegation is resumable: it returns an agentId you can pass to send_message to " +
|
||||
"continue it. Pass `name` to give it a stable teammate name you can address by name instead of the agentId.",
|
||||
schema,
|
||||
mutating: false,
|
||||
handler: async (args, ctx) => {
|
||||
const runSubAgent = ctx.runSubAgent;
|
||||
if (!runSubAgent) {
|
||||
if (!ctx.runSubAgent) {
|
||||
throw new Error("Sub-agents are not available in this context.");
|
||||
}
|
||||
// Parallel batch: run each task as its own sub-agent, concurrently. A single failure surfaces
|
||||
// as that task's `error` rather than rejecting the whole batch — the model gets every sibling's
|
||||
// result and can retry just the one that failed instead of paying for all of them again.
|
||||
if (args.tasks && args.tasks.length > 0) {
|
||||
const results = await Promise.all(
|
||||
args.tasks.map(async (t): Promise<SubAgentResult> => {
|
||||
try {
|
||||
const result = await runSubAgent({ description: t.description, prompt: t.prompt });
|
||||
return { description: t.description, result };
|
||||
} catch (err) {
|
||||
return { description: t.description, error: (err as Error).message ?? String(err) };
|
||||
}
|
||||
}),
|
||||
);
|
||||
return { description: args.description, results };
|
||||
// Pre-flight a name collision so we never clobber an existing teammate (and never waste a
|
||||
// delegation whose name we'd refuse to register). If resolveTeammate is absent (a non-session
|
||||
// context), skip the check — there's no roster to clobber and nothing to register later either.
|
||||
if (args.name && ctx.resolveTeammate?.(args.name)) {
|
||||
return {
|
||||
error: `A teammate named "${args.name}" already exists. Use send_message with name "${args.name}" to continue it, or pick a different name for this new delegation.`,
|
||||
};
|
||||
}
|
||||
// Single sub-agent (the original path).
|
||||
const result = await runSubAgent({ description: args.description, prompt: args.prompt ?? "" });
|
||||
return { description: args.description, result };
|
||||
const spec = getAgentType(args.agentType);
|
||||
// general-purpose inherits the full toolset and uses the generic prompt (no overrides). Specialist
|
||||
// types restrict tools and add their identity as a prompt addendum on top of the generic prompt.
|
||||
const overrides: SubAgentOverrides | undefined =
|
||||
spec.name === "general-purpose" ? undefined : { toolNames: spec.toolNames, systemPromptAddendum: spec.systemPromptAddendum };
|
||||
const result = await ctx.runSubAgent({ description: args.description, prompt: args.prompt }, overrides);
|
||||
// Register the name on the session's roster only when the sub-agent is resumable (shared-cwd, not
|
||||
// worktree-isolated) AND a roster is wired. Isolated parallel agents are fire-and-forget, so a
|
||||
// name wouldn't be addressable — the result simply omits agentId and name, signalling
|
||||
// non-resumability. We only echo `name` when it was actually registered, so the model never sees a
|
||||
// name that send_message can't resolve (e.g. a non-session context with no roster).
|
||||
let registered = false;
|
||||
if (args.name && result.resumable && ctx.registerTeammate) {
|
||||
ctx.registerTeammate(args.name, result.agentId);
|
||||
registered = true;
|
||||
}
|
||||
return {
|
||||
description: args.description,
|
||||
result: result.result,
|
||||
...(result.resumable ? { agentId: result.agentId } : {}),
|
||||
...(registered ? { name: args.name } : {}),
|
||||
};
|
||||
},
|
||||
};
|
||||
};
|
||||
@@ -0,0 +1,57 @@
|
||||
import { describe, expect, it } from "vitest";
|
||||
import { AGENT_TYPES, AGENT_TYPE_NAMES, getAgentType, isReadOnlyAgentType } from "./agentTypes.js";
|
||||
|
||||
describe("agentTypes registry", () => {
|
||||
it("exposes the six built-in types", () => {
|
||||
expect(AGENT_TYPE_NAMES).toEqual(["general-purpose", "explore", "code-reviewer", "planner", "debugger", "test-writer"]);
|
||||
expect(AGENT_TYPES).toHaveLength(6);
|
||||
});
|
||||
|
||||
it("every type has a non-empty addendum", () => {
|
||||
for (const t of AGENT_TYPES) expect(t.systemPromptAddendum.length).toBeGreaterThan(0);
|
||||
});
|
||||
|
||||
it("getAgentType resolves known names and falls back to general-purpose", () => {
|
||||
expect(getAgentType("explore").name).toBe("explore");
|
||||
expect(getAgentType("code-reviewer").name).toBe("code-reviewer");
|
||||
expect(getAgentType("planner").name).toBe("planner");
|
||||
expect(getAgentType("debugger").name).toBe("debugger");
|
||||
expect(getAgentType("test-writer").name).toBe("test-writer");
|
||||
expect(getAgentType("general-purpose").name).toBe("general-purpose");
|
||||
expect(getAgentType(undefined).name).toBe("general-purpose");
|
||||
expect(getAgentType("bogus").name).toBe("general-purpose");
|
||||
});
|
||||
|
||||
it("read-only types (explore, code-reviewer, planner) vs mutating types (general-purpose, debugger, test-writer)", () => {
|
||||
expect(isReadOnlyAgentType("explore")).toBe(true);
|
||||
expect(isReadOnlyAgentType("code-reviewer")).toBe(true);
|
||||
expect(isReadOnlyAgentType("planner")).toBe(true);
|
||||
expect(isReadOnlyAgentType("general-purpose")).toBe(false);
|
||||
expect(isReadOnlyAgentType("debugger")).toBe(false);
|
||||
expect(isReadOnlyAgentType("test-writer")).toBe(false);
|
||||
expect(isReadOnlyAgentType(undefined)).toBe(false);
|
||||
});
|
||||
|
||||
it("read-only toolsets contain only known non-mutating tools", () => {
|
||||
const explore = getAgentType("explore");
|
||||
const reviewer = getAgentType("code-reviewer");
|
||||
const planner = getAgentType("planner");
|
||||
expect(explore.toolNames).not.toContain("write_file");
|
||||
expect(explore.toolNames).not.toContain("bash");
|
||||
expect(explore.toolNames).toContain("read_file");
|
||||
expect(reviewer.toolNames).toContain("git_status");
|
||||
expect(reviewer.toolNames).not.toContain("edit_file");
|
||||
expect(planner.toolNames).toContain("git_status");
|
||||
expect(planner.toolNames).not.toContain("edit_file");
|
||||
expect(planner.toolNames).not.toContain("bash");
|
||||
});
|
||||
|
||||
it("debugger and test-writer can mutate (reproduce/fix and write/run tests)", () => {
|
||||
const debugger_ = getAgentType("debugger");
|
||||
const testWriter = getAgentType("test-writer");
|
||||
expect(debugger_.toolNames).toContain("bash");
|
||||
expect(debugger_.toolNames).toContain("edit_file");
|
||||
expect(testWriter.toolNames).toContain("write_file");
|
||||
expect(testWriter.toolNames).toContain("bash");
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,110 @@
|
||||
import type { ToolDef } from "./types.js";
|
||||
|
||||
/** A built-in specialist sub-agent type the model can request via the `agent` tool's `agentType`
|
||||
* field. Each restricts the sub-agent's toolset (read-only types can't mutate) and prepends a
|
||||
* specialist addendum to the generic sub-agent system prompt. Plugin agents (agent__*) are a
|
||||
* separate, full-replacement mechanism; these built-in types layer on top of the generic prompt so
|
||||
* the standard tool discipline still applies. */
|
||||
export interface AgentTypeSpec {
|
||||
name: string;
|
||||
/** One-line summary surfaced to the model via the schema enum description. */
|
||||
description: string;
|
||||
/** Tools the sub-agent may use (by ToolDef name). Omit to inherit the parent's full toolset minus
|
||||
* further-nesting agent tools (same as a generic sub-agent). */
|
||||
toolNames?: string[];
|
||||
/** Prepended (as an addendum) to the generic sub-agent system prompt — NOT a full replacement, so
|
||||
* the standard tool-use discipline survives. */
|
||||
systemPromptAddendum: string;
|
||||
}
|
||||
|
||||
export const AGENT_TYPES: AgentTypeSpec[] = [
|
||||
{
|
||||
name: "general-purpose",
|
||||
description: "Full tool access (default). Use for implementation and any task that may edit files.",
|
||||
// toolNames omitted → inherit the parent's full toolset.
|
||||
systemPromptAddendum:
|
||||
"You are a general-purpose sub-agent. You may read, search, and edit files to complete the delegated task.",
|
||||
},
|
||||
{
|
||||
name: "explore",
|
||||
description: "Read-only research: locate code and map structure across many files. Cannot modify anything.",
|
||||
toolNames: ["read_file", "list_files", "grep", "web_search", "web_fetch"],
|
||||
systemPromptAddendum:
|
||||
"You are an Explore agent — a read-only research specialist. Your job is to locate code, map structure, and gather " +
|
||||
"facts across the codebase to answer a specific question. You have only read/search tools and must not modify anything. " +
|
||||
"Read excerpts rather than whole files; report conclusions with the file:line references that back them, not file dumps. " +
|
||||
"If the answer isn't findable, say so plainly.",
|
||||
},
|
||||
{
|
||||
name: "code-reviewer",
|
||||
description: "Read-only code review: find bugs, verify claims against code, report ranked findings. Cannot modify.",
|
||||
toolNames: ["read_file", "list_files", "grep", "git_status", "web_search", "web_fetch"],
|
||||
systemPromptAddendum:
|
||||
"You are a code-review specialist. Review the relevant code for correctness, edge cases, and likely bugs. You have only " +
|
||||
"read/search tools. Verify every claim against the actual code rather than assuming. Report concrete findings with " +
|
||||
"file:line anchors, ranked most-severe first; if you find nothing wrong, say so rather than inventing issues. Do not " +
|
||||
"modify code — report only.",
|
||||
},
|
||||
{
|
||||
name: "planner",
|
||||
description: "Read-only planning: design an implementation plan with files to change, steps, and tradeoffs. Cannot modify.",
|
||||
toolNames: ["read_file", "list_files", "grep", "git_status", "web_search", "web_fetch"],
|
||||
systemPromptAddendum:
|
||||
"You are a planning specialist. Investigate the codebase enough to design a concrete implementation plan — which files to " +
|
||||
"change, in what order, and how, with the key code anchors (file:line) that justify each step. Surface tradeoffs and " +
|
||||
"risks between approaches, and call out anything you'd need to verify before implementing. You have only read/search " +
|
||||
"tools and must not modify anything. Return a step-by-step plan, not code dumps.",
|
||||
},
|
||||
{
|
||||
name: "debugger",
|
||||
description: "Reproduce and fix a bug: form hypotheses, read code, run commands to reproduce, apply a minimal fix, verify.",
|
||||
toolNames: ["read_file", "list_files", "grep", "bash", "bash_output", "edit_file", "multi_edit"],
|
||||
systemPromptAddendum:
|
||||
"You are a debugging specialist. Investigate a reported bug by forming a hypothesis, reading the relevant code, and " +
|
||||
"reproducing it with shell commands before touching anything. Apply the minimal fix that addresses the root cause (not " +
|
||||
"the symptom), then verify the fix actually resolves the reproduction. Prefer a small, targeted edit over a rewrite. " +
|
||||
"If you can't reproduce the bug, say so and report what you found instead of guessing at a fix.",
|
||||
},
|
||||
{
|
||||
name: "test-writer",
|
||||
description: "Write focused tests for a feature or bug fix, run them, and iterate until they pass.",
|
||||
toolNames: ["read_file", "list_files", "grep", "write_file", "edit_file", "bash", "bash_output"],
|
||||
systemPromptAddendum:
|
||||
"You are a test-writing specialist. Write focused, meaningful tests (not trivial smoke tests) for the delegated feature " +
|
||||
"or fix, following the project's existing test conventions and runner. Run the tests with shell commands and iterate " +
|
||||
"until they pass — a test that's never run is unfinished. Cover the important edge cases, but don't over-test. If the " +
|
||||
"code under test is wrong, fix it minimally rather than writing a test around the bug.",
|
||||
},
|
||||
];
|
||||
|
||||
export const AGENT_TYPE_NAMES = AGENT_TYPES.map((t) => t.name) as [string, ...string[]];
|
||||
|
||||
/** Looks up a built-in agent type by name. Falls back to general-purpose for an unknown/missing
|
||||
* name so a model that omits the field or typo's it still gets a working sub-agent. */
|
||||
export function getAgentType(name?: string): AgentTypeSpec {
|
||||
if (name) {
|
||||
const found = AGENT_TYPES.find((t) => t.name === name);
|
||||
if (found) return found;
|
||||
}
|
||||
return AGENT_TYPES[0]!;
|
||||
}
|
||||
|
||||
/** Whether a given agent type is read-only (no mutating tools), used to decide worktree isolation:
|
||||
* read-only parallel agents don't need isolation and should see the current working state. */
|
||||
export function isReadOnlyAgentType(name?: string): boolean {
|
||||
const spec = getAgentType(name);
|
||||
return spec.toolNames !== undefined && !spec.toolNames.some((n) => MUTATING_TOOL_NAMES.has(n));
|
||||
}
|
||||
|
||||
/** Mutating tool names, for isReadOnlyAgentType. Kept here (not imported from the tool defs) so this
|
||||
* stays a static decision without instantiating tools. */
|
||||
const MUTATING_TOOL_NAMES = new Set([
|
||||
"write_file",
|
||||
"edit_file",
|
||||
"multi_edit",
|
||||
"notebook_edit",
|
||||
"bash",
|
||||
"bash_kill",
|
||||
"git_commit",
|
||||
"memory_write",
|
||||
]);
|
||||
@@ -0,0 +1,87 @@
|
||||
import { describe, expect, it, vi } from "vitest";
|
||||
import { askQuestionTool } from "./askQuestion.js";
|
||||
import type { AskQuestionAnswer, AskQuestionSpec, ToolContext } from "./types.js";
|
||||
|
||||
const validArgs = {
|
||||
questions: [
|
||||
{
|
||||
question: "Which auth method?",
|
||||
header: "Auth method",
|
||||
options: [
|
||||
{ label: "OAuth", description: "delegate to provider" },
|
||||
{ label: "API key", description: "simple header token" },
|
||||
],
|
||||
},
|
||||
],
|
||||
};
|
||||
|
||||
function ctxWith(askQuestion?: ToolContext["askQuestion"]): ToolContext {
|
||||
return { cwd: "/x", ...(askQuestion ? { askQuestion } : {}) };
|
||||
}
|
||||
|
||||
describe("ask_user_question tool", () => {
|
||||
it("parses a well-formed single question", () => {
|
||||
expect(() => askQuestionTool.schema.parse(validArgs)).not.toThrow();
|
||||
});
|
||||
|
||||
it("rejects fewer than 2 options per question", () => {
|
||||
expect(() =>
|
||||
askQuestionTool.schema.parse({ questions: [{ question: "q?", header: "h", options: [{ label: "only" }] }] }),
|
||||
).toThrow();
|
||||
});
|
||||
|
||||
it("rejects more than 4 options per question", () => {
|
||||
const opts = Array.from({ length: 5 }, (_, i) => ({ label: `o${i}` }));
|
||||
expect(() => askQuestionTool.schema.parse({ questions: [{ question: "q?", header: "h", options: opts }] })).toThrow();
|
||||
});
|
||||
|
||||
it("rejects zero questions", () => {
|
||||
expect(() => askQuestionTool.schema.parse({ questions: [] })).toThrow();
|
||||
});
|
||||
|
||||
it("rejects more than 4 questions", () => {
|
||||
const qs = Array.from({ length: 5 }, (_, i) => ({
|
||||
question: `q${i}?`,
|
||||
header: `h${i}`,
|
||||
options: [{ label: "a" }, { label: "b" }],
|
||||
}));
|
||||
expect(() => askQuestionTool.schema.parse({ questions: qs })).toThrow();
|
||||
});
|
||||
|
||||
it("rejects a header longer than 12 chars", () => {
|
||||
expect(() =>
|
||||
askQuestionTool.schema.parse({
|
||||
questions: [{ question: "q?", header: "this is too long", options: [{ label: "a" }, { label: "b" }] }],
|
||||
}),
|
||||
).toThrow();
|
||||
});
|
||||
|
||||
it("returns the user's answers when askQuestion is wired", async () => {
|
||||
const answers: AskQuestionAnswer[] = [{ question: "Which auth method?", selected: ["OAuth"] }];
|
||||
const askQuestion = vi.fn(async (_questions: AskQuestionSpec[]) => answers);
|
||||
const result = await askQuestionTool.handler(validArgs as any, ctxWith(askQuestion));
|
||||
expect(askQuestion).toHaveBeenCalledTimes(1);
|
||||
// The spec passed to the callback preserves question/header/options and omits undefined fields.
|
||||
const first = askQuestion.mock.calls[0]![0][0]!;
|
||||
expect(first.question).toBe("Which auth method?");
|
||||
expect(first.header).toBe("Auth method");
|
||||
expect(first.options[0]).toEqual({ label: "OAuth", description: "delegate to provider" });
|
||||
expect("multiSelect" in first).toBe(false);
|
||||
expect(result).toEqual({ answers });
|
||||
});
|
||||
|
||||
it("preserves multiSelect when set", async () => {
|
||||
const askQuestion = vi.fn(async (_questions: AskQuestionSpec[]) => [{ question: "q?", selected: ["a", "b"] }]);
|
||||
await askQuestionTool.handler(
|
||||
{ questions: [{ question: "q?", header: "h", options: [{ label: "a" }, { label: "b" }], multiSelect: true }] } as any,
|
||||
ctxWith(askQuestion),
|
||||
);
|
||||
expect(askQuestion.mock.calls[0]![0][0]!.multiSelect).toBe(true);
|
||||
});
|
||||
|
||||
it("returns a clear error (not a hang) when no interactive UI is available", async () => {
|
||||
const result = await askQuestionTool.handler(validArgs as any, ctxWith(undefined));
|
||||
expect("error" in (result as object)).toBe(true);
|
||||
expect((result as { error: string }).error).toMatch(/no interactive UI|Can't ask/i);
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,68 @@
|
||||
import { z } from "zod";
|
||||
import type { AskQuestionAnswer, AskQuestionSpec, ToolDef } from "./types.js";
|
||||
|
||||
const optionSchema = z.object({
|
||||
label: z.string().describe("A concise (1-5 word) label for the option."),
|
||||
description: z
|
||||
.string()
|
||||
.optional()
|
||||
.describe("Explanation of what this option means or the trade-off it implies, shown dimmed under the label."),
|
||||
});
|
||||
|
||||
const questionSchema = z.object({
|
||||
question: z.string().describe("The complete question to ask, ending with a question mark."),
|
||||
header: z
|
||||
.string()
|
||||
.max(12)
|
||||
.describe("A very short label (max ~12 chars) shown as a chip beside the question, e.g. \"Auth method\"."),
|
||||
options: z
|
||||
.array(optionSchema)
|
||||
.min(2)
|
||||
.max(4)
|
||||
.describe("Two to four mutually exclusive options (unless multiSelect). The user can also type a custom \"Other\" answer."),
|
||||
multiSelect: z
|
||||
.boolean()
|
||||
.optional()
|
||||
.describe("Set true to allow several options to be selected instead of just one."),
|
||||
});
|
||||
|
||||
const schema = z.object({
|
||||
questions: z
|
||||
.array(questionSchema)
|
||||
.min(1)
|
||||
.max(4)
|
||||
.describe("One to four questions to ask. The UI asks them one at a time and returns all answers together."),
|
||||
});
|
||||
|
||||
/** Lets the model ask the user a structured multiple-choice question when it is blocked on a decision
|
||||
* that is genuinely the user's to make — one it can't resolve from the code, request, or sensible
|
||||
* defaults. Non-mutating (it changes nothing on the filesystem), so it's allowed in every permission
|
||||
* mode including plan mode. The UI shows each question with its options (plus an implicit "Other"
|
||||
* path for a freeform answer) and returns the selected label(s); in a headless context with no UI
|
||||
* (sub-agents) the callback is absent and the tool fails with a clear "can't ask" error instead of
|
||||
* hanging. Reserve this for real decision points — don't ask questions you could answer yourself by
|
||||
* reading the code or following an obvious default. */
|
||||
export const askQuestionTool: ToolDef<z.infer<typeof schema>> = {
|
||||
name: "ask_user_question",
|
||||
description:
|
||||
"Ask the user a structured multiple-choice question when blocked on a decision only they can make. " +
|
||||
"Pass 1-4 questions, each with 2-4 options and a short header chip. The user can pick an option or type a " +
|
||||
"custom \"Other\" answer. Use this instead of a prose question when a discrete choice would clarify the path. " +
|
||||
"Only ask when the request is genuinely ambiguous after you've explored — don't offload decisions you could " +
|
||||
"make yourself.",
|
||||
schema,
|
||||
mutating: false,
|
||||
handler: async (args, ctx) => {
|
||||
if (!ctx.askQuestion) {
|
||||
return { error: "Can't ask the user a question in this context (no interactive UI). Make a sensible default choice and proceed, or explain the trade-off in prose." };
|
||||
}
|
||||
const specs: AskQuestionSpec[] = args.questions.map((q) => ({
|
||||
question: q.question,
|
||||
header: q.header,
|
||||
options: q.options.map((o) => ({ label: o.label, ...(o.description ? { description: o.description } : {}) })),
|
||||
...(q.multiSelect ? { multiSelect: true } : {}),
|
||||
}));
|
||||
const answers: AskQuestionAnswer[] = await ctx.askQuestion(specs);
|
||||
return { answers };
|
||||
},
|
||||
};
|
||||
@@ -1,5 +1,4 @@
|
||||
import { beforeEach, describe, expect, it, vi } from "vitest";
|
||||
import { execa } from "execa";
|
||||
import type { ResultPromise } from "execa";
|
||||
import { bashTool } from "./bash.js";
|
||||
import { killProcessTree } from "../utils/processTree.js";
|
||||
@@ -81,36 +80,4 @@ describe("bash tool — sub-agent abort", () => {
|
||||
|
||||
expect(killProcessTree).not.toHaveBeenCalled();
|
||||
});
|
||||
});
|
||||
|
||||
describe("bash tool — safety guards", () => {
|
||||
beforeEach(() => {
|
||||
vi.mocked(execa).mockClear();
|
||||
});
|
||||
|
||||
it("refuses a command that wipes the filesystem root before ever spawning it", async () => {
|
||||
const ctx: ToolContext = { cwd: process.cwd() };
|
||||
await expect(bashTool.handler({ command: "rm -rf /" }, ctx)).rejects.toThrow(/Refusing to run/);
|
||||
expect(execa).not.toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it("does not flag an ordinary rm -rf on a project subdirectory", async () => {
|
||||
// preview (not handler) so this doesn't spawn the fake child, which only ever resolves when
|
||||
// killed/backgrounded — nothing here would do either, so awaiting handler() would hang.
|
||||
const ctx: ToolContext = { cwd: process.cwd() };
|
||||
const preview = await bashTool.preview!({ command: "rm -rf node_modules" }, ctx);
|
||||
expect(preview).not.toMatch(/^Blocked:/);
|
||||
});
|
||||
|
||||
it("refuses a cwd override that escapes the working directory", async () => {
|
||||
const ctx: ToolContext = { cwd: process.cwd() };
|
||||
await expect(bashTool.handler({ command: "ls", cwd: "../../" }, ctx)).rejects.toThrow(/Refusing to write outside/);
|
||||
expect(execa).not.toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it("preview surfaces the block reason instead of running the command", async () => {
|
||||
const ctx: ToolContext = { cwd: process.cwd() };
|
||||
const preview = await bashTool.preview!({ command: "mkfs.ext4 /dev/sda1" }, ctx);
|
||||
expect(preview).toMatch(/^Blocked:/);
|
||||
});
|
||||
});
|
||||
+24
-33
@@ -1,11 +1,11 @@
|
||||
import path from "node:path";
|
||||
import { execa } from "execa";
|
||||
import { z } from "zod";
|
||||
import { registerBackgroundJob } from "./backgroundJobs.js";
|
||||
import { riskyBashCommandReason } from "./bashGuard.js";
|
||||
import { resolveWithinCwd } from "./pathGuard.js";
|
||||
import { killProcessTree } from "../utils/processTree.js";
|
||||
import { truncate } from "../utils/truncate.js";
|
||||
import { resolveShell } from "../utils/shell.js";
|
||||
import { assertWithinWorkspace } from "../utils/path.js";
|
||||
import type { ToolDef } from "./types.js";
|
||||
|
||||
const schema = z.object({
|
||||
@@ -22,39 +22,28 @@ function delay(ms: number): Promise<"pending"> {
|
||||
|
||||
export const bashTool: ToolDef<z.infer<typeof schema>> = {
|
||||
name: "bash",
|
||||
description: "Run a shell command and return its stdout, stderr, and exit code. Use for building, running tests, git operations, or inspecting the environment. Output is capped (head+tail preserved); long-running commands can be backgrounded with Ctrl+B and checked with bash_output.",
|
||||
description:
|
||||
"Run a shell command and return its stdout, stderr, and exit code. Be careful with destructive " +
|
||||
"operations (rm, git push, etc.). For long-running commands, increase timeout_ms or use Ctrl+B " +
|
||||
"to background the command while it's running.",
|
||||
schema,
|
||||
mutating: true,
|
||||
preview: async ({ command, cwd }, ctx) => {
|
||||
const riskyReason = riskyBashCommandReason(command);
|
||||
if (riskyReason) return `Blocked: this command ${riskyReason}.`;
|
||||
if (cwd) {
|
||||
try {
|
||||
resolveWithinCwd(ctx.cwd, cwd);
|
||||
} catch (err) {
|
||||
return (err as Error).message;
|
||||
}
|
||||
}
|
||||
return `Run shell command: ${command}${cwd ? ` (cwd: ${cwd})` : ""}`;
|
||||
},
|
||||
preview: async ({ command, cwd }) => `Run shell command: ${command}${cwd ? ` (cwd: ${cwd})` : ""}`,
|
||||
handler: async ({ command, cwd, timeout_ms }, ctx) => {
|
||||
const riskyReason = riskyBashCommandReason(command);
|
||||
if (riskyReason) {
|
||||
throw new Error(`Refusing to run: this command ${riskyReason}.`);
|
||||
}
|
||||
const workDir = cwd ? resolveWithinCwd(ctx.cwd, cwd) : ctx.cwd;
|
||||
const workDir = cwd ? path.resolve(ctx.cwd, cwd) : ctx.cwd;
|
||||
if (cwd) assertWithinWorkspace(workDir, ctx.cwd, cwd);
|
||||
// Timeout is enforced by our own timer rather than execa's built-in `timeout` option, so that
|
||||
// backgrounding via Ctrl+B can cancel it below — execa's own timeout kills the process on a
|
||||
// fixed schedule regardless of what happens to it afterward, which would silently kill a
|
||||
// long-running command right after the user chose to keep it running in the background.
|
||||
const child = execa(command, { shell: resolveShell(), cwd: workDir, reject: false });
|
||||
const stdoutChunks: string[] = [];
|
||||
const stderrChunks: string[] = [];
|
||||
let stdout = "";
|
||||
let stderr = "";
|
||||
const onStdout = (d: Buffer) => {
|
||||
stdoutChunks.push(d.toString());
|
||||
stdout += d.toString();
|
||||
};
|
||||
const onStderr = (d: Buffer) => {
|
||||
stderrChunks.push(d.toString());
|
||||
stderr += d.toString();
|
||||
};
|
||||
child.stdout?.on("data", onStdout);
|
||||
child.stderr?.on("data", onStderr);
|
||||
@@ -88,13 +77,10 @@ export const bashTool: ToolDef<z.infer<typeof schema>> = {
|
||||
for (;;) {
|
||||
if (ctx.backgroundControl?.requested) {
|
||||
clearTimeout(foregroundTimer);
|
||||
// Detach our own capture listeners before handing the streams to the background registry —
|
||||
// otherwise both this closure's listeners and registerBackgroundJob's keep appending to
|
||||
// separate buffers forever, doubling the work and growing memory without bound for a
|
||||
// long-running backgrounded job. The buffers captured so far seed the job.
|
||||
child.stdout?.off("data", onStdout);
|
||||
child.stderr?.off("data", onStderr);
|
||||
const job = registerBackgroundJob(command, workDir, child, stdoutChunks.join(""), stderrChunks.join(""));
|
||||
// Hand the streams off to the background registry; the finally block below will detach
|
||||
// our own capture listeners so both closures don't keep appending to separate buffers
|
||||
// forever, doubling the work and growing memory without bound for a long-running job.
|
||||
const job = registerBackgroundJob(command, workDir, child, stdout, stderr);
|
||||
return {
|
||||
backgrounded: true,
|
||||
jobId: job.id,
|
||||
@@ -106,13 +92,18 @@ export const bashTool: ToolDef<z.infer<typeof schema>> = {
|
||||
clearTimeout(foregroundTimer);
|
||||
return {
|
||||
exitCode: settled.exitCode,
|
||||
stdout: truncate(stdoutChunks.join("")),
|
||||
stderr: truncate(stderrChunks.join("")),
|
||||
stdout: truncate(stdout),
|
||||
stderr: truncate(stderr),
|
||||
timedOut,
|
||||
};
|
||||
}
|
||||
}
|
||||
} finally {
|
||||
// Detach our capture listeners on every exit path so a long-running command doesn't keep
|
||||
// orphaned handlers alive after the tool returns. On backgrounding this also stops the
|
||||
// foreground closure from competing with the background registry for stream data.
|
||||
child.stdout?.off("data", onStdout);
|
||||
child.stderr?.off("data", onStderr);
|
||||
// Remove the abort listener on every exit path. On backgrounding this is what stops a
|
||||
// sub-agent timeout from killing a job the user explicitly chose to keep running; on normal
|
||||
// completion it's just cleanup. (The listener is `{ once: true }`, but it may never fire.)
|
||||
|
||||
@@ -1,45 +0,0 @@
|
||||
import { describe, expect, it } from "vitest";
|
||||
import { riskyBashCommandReason } from "./bashGuard.js";
|
||||
|
||||
describe("riskyBashCommandReason", () => {
|
||||
describe("blocks", () => {
|
||||
const cases: [name: string, command: string][] = [
|
||||
["rm -rf /", "rm -rf /"],
|
||||
["rm -fr / (flag order swapped)", "rm -fr /"],
|
||||
["rm -rf / with a trailing slash-star", "rm -rf /*"],
|
||||
["rm -Rf ~ (home dir)", "rm -Rf ~"],
|
||||
["rm --recursive --force /", "rm --recursive --force /"],
|
||||
["sudo rm -rf /", "sudo rm -rf /"],
|
||||
["classic fork bomb", ":(){ :|:& };:"],
|
||||
["fork bomb with extra whitespace", ": ( ) { : | : & } ; :"],
|
||||
["mkfs.ext4", "mkfs.ext4 /dev/sda1"],
|
||||
["dd to a raw device", "dd if=/dev/zero of=/dev/sda bs=1M"],
|
||||
["redirect onto a raw device", "echo oops > /dev/sda"],
|
||||
["Windows format", "format C:"],
|
||||
["Windows rd /s /q on a drive root", "rd /s /q C:\\"],
|
||||
["PowerShell Remove-Item -Recurse -Force on a drive", "Remove-Item -Recurse -Force C:\\"],
|
||||
];
|
||||
for (const [name, command] of cases) {
|
||||
it(name, () => {
|
||||
expect(riskyBashCommandReason(command)).not.toBeNull();
|
||||
});
|
||||
}
|
||||
});
|
||||
|
||||
describe("does not block", () => {
|
||||
const cases: [name: string, command: string][] = [
|
||||
["rm -rf on a project subdirectory", "rm -rf node_modules"],
|
||||
["rm -rf on a relative build dir", "rm -rf ./dist"],
|
||||
["rm without force/recursive on root-looking arg", "rm /tmp/foo.txt"],
|
||||
["dd between two regular files", "dd if=file.img of=out.img"],
|
||||
["a command that merely contains the word format", "echo 'please format your PR title'"],
|
||||
["a normal git command", "git status"],
|
||||
["listing a directory named format", "ls format"],
|
||||
];
|
||||
for (const [name, command] of cases) {
|
||||
it(name, () => {
|
||||
expect(riskyBashCommandReason(command)).toBeNull();
|
||||
});
|
||||
}
|
||||
});
|
||||
});
|
||||
@@ -1,87 +0,0 @@
|
||||
/** Blocks a small set of unambiguously catastrophic shell commands — wiping the whole filesystem
|
||||
* or a whole drive, formatting a device, a fork bomb — before they ever reach the confirmation
|
||||
* prompt (or, under auto-accept, before they'd run with no prompt at all). This is not a general
|
||||
* command sandbox: it doesn't stop a model from `rm -rf`-ing some *other* directory it shouldn't,
|
||||
* running a slow fork loop that isn't the canonical bomb syntax, or anything merely inadvisable —
|
||||
* only the handful of patterns whose only realistic purpose is destroying the whole machine, where
|
||||
* a false negative is far more likely than a false positive. Deliberately narrow so it doesn't
|
||||
* reject legitimate commands like `rm -rf node_modules` or `dd if=file.img of=out.img`. */
|
||||
|
||||
interface RiskyPattern {
|
||||
test: (command: string) => boolean;
|
||||
reason: string;
|
||||
}
|
||||
|
||||
/** Splits on whitespace for a crude token scan — good enough for a blocklist (not a security
|
||||
* boundary; execa still runs the raw string through a real shell either way) and avoids a brittle
|
||||
* do-everything regex that has to encode flag ordering itself. */
|
||||
function tokenize(command: string): string[] {
|
||||
return command.trim().split(/\s+/);
|
||||
}
|
||||
|
||||
const ROOT_TARGETS = new Set(["/", "/*", "~", "~/", "~/*", "$home", "${home}"]);
|
||||
|
||||
/** `rm -rf /`, `rm -fr ~`, `sudo rm -Rf --no-preserve-root /`, etc. — recursive+forced deletion
|
||||
* whose target is the filesystem root or the whole home directory, in any flag order/spelling. */
|
||||
function isRmWipingRootOrHome(command: string): boolean {
|
||||
const tokens = tokenize(command).map((t) => t.toLowerCase());
|
||||
const rmIdx = tokens.findIndex((t) => t === "rm" || t.endsWith("/rm"));
|
||||
if (rmIdx === -1) return false;
|
||||
const rest = tokens.slice(rmIdx + 1);
|
||||
const isFlag = (t: string) => t.startsWith("-");
|
||||
const hasForce = rest.some((t) => (isFlag(t) && !t.startsWith("--") && t.includes("f")) || t === "--force");
|
||||
const hasRecursive = rest.some((t) => (isFlag(t) && !t.startsWith("--") && (t.includes("r") || t.includes("R"))) || t === "--recursive");
|
||||
const targets = rest.filter((t) => !isFlag(t));
|
||||
return hasForce && hasRecursive && targets.some((t) => ROOT_TARGETS.has(t));
|
||||
}
|
||||
|
||||
const RISKY_PATTERNS: RiskyPattern[] = [
|
||||
{
|
||||
test: isRmWipingRootOrHome,
|
||||
reason: "recursively force-deletes the filesystem root or home directory",
|
||||
},
|
||||
{
|
||||
// Classic bash fork bomb: ":(){ :|:& };:" (whitespace-tolerant).
|
||||
test: (cmd) => /:\s*\(\s*\)\s*\{\s*:\s*\|\s*:\s*&?\s*;?\s*\}\s*;\s*:/.test(cmd),
|
||||
reason: "is a fork bomb (unbounded process spawning)",
|
||||
},
|
||||
{
|
||||
test: (cmd) => /\bmkfs(\.\w+)?\b/i.test(cmd),
|
||||
reason: "formats a filesystem (mkfs)",
|
||||
},
|
||||
{
|
||||
test: (cmd) => /\bdd\b[^\n]*\bof=\/dev\/(sd|hd|nvme|disk|xvd|rdisk)\w*/i.test(cmd),
|
||||
reason: "writes raw data directly to a block device (dd of=/dev/...)",
|
||||
},
|
||||
{
|
||||
test: (cmd) => />\s*\/dev\/(sd|hd|nvme|disk|xvd|rdisk)\w*\b/i.test(cmd),
|
||||
reason: "redirects output directly onto a block device",
|
||||
},
|
||||
{
|
||||
// `format C:`, `format /Y D:` — Windows drive format.
|
||||
test: (cmd) => /\bformat\b[^\n]*\b[a-zA-Z]:/i.test(cmd),
|
||||
reason: "formats a Windows drive (format)",
|
||||
},
|
||||
{
|
||||
// `rd /s /q C:\`, `rmdir /s /q D:\` — recursive quiet delete of a bare drive root.
|
||||
test: (cmd) => /\b(rd|rmdir)\b[^\n]*\/s\b[^\n]*\b[a-zA-Z]:\\?\s*(\/q\b[^\n]*)?$/im.test(cmd),
|
||||
reason: "recursively deletes an entire Windows drive",
|
||||
},
|
||||
{
|
||||
// PowerShell `Remove-Item -Recurse -Force C:\` (or -Path C:\, or $env:SystemDrive), flag order-tolerant.
|
||||
test: (cmd) =>
|
||||
/remove-item\b/i.test(cmd) &&
|
||||
/-recurse\b/i.test(cmd) &&
|
||||
/-force\b/i.test(cmd) &&
|
||||
(/\b[a-zA-Z]:\\?\s*($|['")\s;])/.test(cmd) || /\$env:systemdrive\b/i.test(cmd)),
|
||||
reason: "recursively force-deletes an entire Windows drive (Remove-Item)",
|
||||
},
|
||||
];
|
||||
|
||||
/** Returns a human-readable reason if `command` matches a known catastrophic pattern, else null. */
|
||||
export function riskyBashCommandReason(command: string): string | null {
|
||||
for (const pattern of RISKY_PATTERNS) {
|
||||
if (pattern.test(command)) return pattern.reason;
|
||||
}
|
||||
return null;
|
||||
}
|
||||
@@ -1,68 +0,0 @@
|
||||
import { describe, it, expect, vi, beforeEach } from "vitest";
|
||||
|
||||
// The LSP tools (definition/references/diagnostics) are thin dispatchers over lspManager. We mock
|
||||
// the manager functions so the tests run without spawning a language server, and assert each tool
|
||||
// forwards the right arguments (path resolved against cwd, 1-indexed→handled by the manager,
|
||||
// includeDeclaration default) and returns the manager's result verbatim.
|
||||
|
||||
vi.mock("../codeintel/lspManager.js", () => ({
|
||||
getDefinition: vi.fn(async () => ({ definitions: [{ path: "/abs/a.ts", line: 3, column: 5 }] })),
|
||||
getReferences: vi.fn(async () => ({ references: [{ path: "/abs/a.ts", line: 3, column: 5 }] })),
|
||||
getDiagnostics: vi.fn(async () => ({
|
||||
diagnostics: [{ path: "/abs/a.ts", line: 1, column: 1, severity: "error", message: "oops" }],
|
||||
})),
|
||||
}));
|
||||
|
||||
import { definitionTool, referencesTool, diagnosticsTool } from "./codeIntel.js";
|
||||
import { getDefinition, getReferences, getDiagnostics } from "../codeintel/lspManager.js";
|
||||
|
||||
const ctx = { cwd: "/proj" };
|
||||
|
||||
describe("definition tool", () => {
|
||||
beforeEach(() => vi.clearAllMocks());
|
||||
|
||||
it("forwards path/line/column/cwd to getDefinition and returns its result", async () => {
|
||||
const out = await definitionTool.handler({ path: "src/a.ts", line: 3, column: 5 }, ctx);
|
||||
expect(getDefinition).toHaveBeenCalledWith("src/a.ts", 3, 5, "/proj");
|
||||
expect(out).toEqual({ definitions: [{ path: "/abs/a.ts", line: 3, column: 5 }] });
|
||||
});
|
||||
|
||||
it("is read-only (no permission prompt)", () => {
|
||||
expect(definitionTool.mutating).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe("references tool", () => {
|
||||
beforeEach(() => vi.clearAllMocks());
|
||||
|
||||
it("defaults includeDeclaration to true when omitted", async () => {
|
||||
await referencesTool.handler({ path: "src/a.ts", line: 3, column: 5 }, ctx);
|
||||
expect(getReferences).toHaveBeenCalledWith("src/a.ts", 3, 5, "/proj", true);
|
||||
});
|
||||
|
||||
it("passes an explicit includeDeclaration through", async () => {
|
||||
await referencesTool.handler({ path: "src/a.ts", line: 3, column: 5, include_declaration: false }, ctx);
|
||||
expect(getReferences).toHaveBeenCalledWith("src/a.ts", 3, 5, "/proj", false);
|
||||
});
|
||||
|
||||
it("returns the manager's references result", async () => {
|
||||
const out = await referencesTool.handler({ path: "src/a.ts", line: 3, column: 5 }, ctx);
|
||||
expect(out).toEqual({ references: [{ path: "/abs/a.ts", line: 3, column: 5 }] });
|
||||
});
|
||||
});
|
||||
|
||||
describe("diagnostics tool", () => {
|
||||
beforeEach(() => vi.clearAllMocks());
|
||||
|
||||
it("forwards path/cwd to getDiagnostics and returns its result", async () => {
|
||||
const out = await diagnosticsTool.handler({ path: "src/a.ts" }, ctx);
|
||||
expect(getDiagnostics).toHaveBeenCalledWith("src/a.ts", "/proj");
|
||||
expect(out).toEqual({
|
||||
diagnostics: [{ path: "/abs/a.ts", line: 1, column: 1, severity: "error", message: "oops" }],
|
||||
});
|
||||
});
|
||||
|
||||
it("is read-only", () => {
|
||||
expect(diagnosticsTool.mutating).toBe(false);
|
||||
});
|
||||
});
|
||||
@@ -1,61 +0,0 @@
|
||||
import { z } from "zod";
|
||||
import type { ToolDef } from "./types.js";
|
||||
import { getDefinition, getReferences, getDiagnostics } from "../codeintel/lspManager.js";
|
||||
|
||||
// All three tools are read-only LSP queries. They share a common shape: point them at a file
|
||||
// (relative to cwd) and a 1-indexed line/column, and they ask the language server for the answer.
|
||||
// The server is lazily started on first use per language (tsserver, pyright, gopls, clangd,
|
||||
// rust-analyzer) and reused across the whole session — see lspManager.ts for the lifecycle.
|
||||
//
|
||||
// The error messages from lspManager are written to be actionable (e.g. "install
|
||||
// typescript-language-server"), so we let them surface verbatim rather than wrapping them — a
|
||||
// generic "LSP unavailable" would hide the one piece of info the model needs to recover.
|
||||
|
||||
const positionSchema = z.object({
|
||||
path: z
|
||||
.string()
|
||||
.describe("File path, relative to the working directory. Must match an extension with a configured LSP server (.ts/.tsx/.js/.jsx/.py/.go/.rs/.c/.cpp/…)."),
|
||||
line: z.number().int().min(1).describe("1-indexed line number of the symbol to query."),
|
||||
column: z.number().int().min(1).describe("1-indexed column number of the symbol to query."),
|
||||
});
|
||||
|
||||
export const definitionTool: ToolDef<z.infer<typeof positionSchema>> = {
|
||||
name: "definition",
|
||||
description:
|
||||
"Resolve where a symbol is DEFINED using the language server (LSP). Use when grep finds a call site but you need the actual declaration — e.g. a function/variable/type name at a line:column. Returns one or more file:line:column locations (empty list if the server couldn't resolve it, which is a legitimate 'not found', not an error). Requires the relevant language server on PATH (typescript-language-server, pyright-langserver, gopls, clangd, or rust-analyzer).",
|
||||
schema: positionSchema,
|
||||
mutating: false,
|
||||
handler: async (args, ctx) => getDefinition(args.path, args.line, args.column, ctx.cwd),
|
||||
};
|
||||
|
||||
const referencesSchema = positionSchema.extend({
|
||||
include_declaration: z
|
||||
.boolean()
|
||||
.optional()
|
||||
.describe("Whether to include the symbol's own declaration among the references. Defaults to true (matches most IDE 'find all references' behavior)."),
|
||||
});
|
||||
|
||||
export const referencesTool: ToolDef<z.infer<typeof referencesSchema>> = {
|
||||
name: "references",
|
||||
description:
|
||||
"Find every reference to a symbol using the language server (LSP) — the same as an IDE's 'find all references'. Use to enumerate all call/usage sites of a function/variable/type at a line:column before a rename or to gauge impact. Returns a list of file:line:column locations. Requires the relevant language server on PATH.",
|
||||
schema: referencesSchema,
|
||||
mutating: false,
|
||||
handler: async (args, ctx) =>
|
||||
getReferences(args.path, args.line, args.column, ctx.cwd, args.include_declaration ?? true),
|
||||
};
|
||||
|
||||
const diagnosticsSchema = z.object({
|
||||
path: z
|
||||
.string()
|
||||
.describe("File path, relative to the working directory, to check for type/syntax errors."),
|
||||
});
|
||||
|
||||
export const diagnosticsTool: ToolDef<z.infer<typeof diagnosticsSchema>> = {
|
||||
name: "diagnostics",
|
||||
description:
|
||||
"Get the latest type/syntax diagnostics (errors and warnings) the language server has published for a file — equivalent to an editor's Problems panel. Use right after an edit_file/write_file to verify the change didn't introduce a type error, or when `tsc --noEmit`/`pyright` would be the alternative. Forces a document sync first so the snapshot is current. Returns severity (error/warning/information/hint), line, column, and message for each diagnostic. Requires the relevant language server on PATH.",
|
||||
schema: diagnosticsSchema,
|
||||
mutating: false,
|
||||
handler: async (args, ctx) => getDiagnostics(args.path, ctx.cwd),
|
||||
};
|
||||
@@ -0,0 +1,95 @@
|
||||
import { z } from "zod";
|
||||
import type { ToolDef } from "./types.js";
|
||||
|
||||
export const cronCreateTool: ToolDef<z.infer<typeof cronCreateSchema>> = {
|
||||
name: "cron_create",
|
||||
description:
|
||||
"Schedule a prompt to run on a recurring cron schedule (5-field cron in the user's LOCAL timezone: minute hour " +
|
||||
"day-of-month month day-of-week, e.g. '0 9 * * 1-5' = weekdays at 9am). Use for recurring checks, reminders, or " +
|
||||
"self-paced loops. The prompt fires only while the REPL is idle. Recurring jobs auto-expire after 7 days. Set " +
|
||||
"recurring: false for a one-shot that fires once then deletes itself. Set durable: true to persist across restarts. " +
|
||||
"Returns the new job id.",
|
||||
schema: z.object({
|
||||
cron: z
|
||||
.string()
|
||||
.min(1)
|
||||
.describe("5-field cron expression (minute hour day-of-month month day-of-week) in local time."),
|
||||
prompt: z.string().min(1).describe("The prompt to enqueue when the job fires."),
|
||||
recurring: z.boolean().optional().describe("True (default) to fire on every match; false to fire once then delete."),
|
||||
durable: z
|
||||
.boolean()
|
||||
.optional()
|
||||
.describe("True to persist the job to disk so it survives a restart (default false = session-only)."),
|
||||
}),
|
||||
// Scheduling is reversible (cron_delete) and not a destructive filesystem op — no confirmation prompt.
|
||||
mutating: false,
|
||||
handler: async (args, ctx) => {
|
||||
if (!ctx.cronStore) return { error: "Scheduling is not available in this context." };
|
||||
try {
|
||||
const job = ctx.cronStore.create(args);
|
||||
return { id: job.id, job };
|
||||
} catch (err) {
|
||||
return { error: (err as Error).message };
|
||||
}
|
||||
},
|
||||
};
|
||||
|
||||
const cronCreateSchema = z.object({
|
||||
cron: z.string().min(1),
|
||||
prompt: z.string().min(1),
|
||||
recurring: z.boolean().optional(),
|
||||
durable: z.boolean().optional(),
|
||||
});
|
||||
|
||||
export const cronListTool: ToolDef<z.infer<typeof cronListSchema>> = {
|
||||
name: "cron_list",
|
||||
description: "List all scheduled cron jobs with their id, schedule, prompt, and whether they're recurring/durable.",
|
||||
schema: z.object({}),
|
||||
mutating: false,
|
||||
handler: async (_args, ctx) => {
|
||||
return { jobs: ctx.cronStore?.list() ?? [] };
|
||||
},
|
||||
};
|
||||
|
||||
const cronListSchema = z.object({});
|
||||
|
||||
export const cronDeleteTool: ToolDef<z.infer<typeof cronDeleteSchema>> = {
|
||||
name: "cron_delete",
|
||||
description: "Cancel a scheduled cron job by id (from cron_list or cron_create's return). Returns { deleted: id } on success.",
|
||||
schema: z.object({ id: z.string().min(1) }),
|
||||
mutating: false,
|
||||
handler: async (args, ctx) => {
|
||||
const store = ctx.cronStore;
|
||||
if (!store) return { error: "Scheduling is not available in this context." };
|
||||
return store.delete(args.id) ? { deleted: args.id } : { error: `Job ${args.id} not found.` };
|
||||
},
|
||||
};
|
||||
|
||||
const cronDeleteSchema = z.object({ id: z.string().min(1) });
|
||||
|
||||
export const scheduleWakeupTool: ToolDef<z.infer<typeof scheduleWakeupSchema>> = {
|
||||
name: "schedule_wakeup",
|
||||
description:
|
||||
"Schedule a one-shot prompt to fire after delaySeconds (60-3600), for self-paced loops that check back on external " +
|
||||
"state. Pass stop: true to cancel ALL pending wakeups and end the loop. The prompt fires once then is removed. " +
|
||||
"Only fires while the REPL is idle.",
|
||||
schema: z.object({
|
||||
delaySeconds: z.number().int().min(1),
|
||||
prompt: z.string().min(1),
|
||||
stop: z.boolean().optional(),
|
||||
reason: z.string().optional(),
|
||||
}),
|
||||
mutating: false,
|
||||
handler: async (args, ctx) => {
|
||||
if (!ctx.cronStore) return { error: "Scheduling is not available in this context." };
|
||||
const result = ctx.cronStore.scheduleWakeup(args);
|
||||
return result;
|
||||
},
|
||||
};
|
||||
|
||||
const scheduleWakeupSchema = z.object({
|
||||
delaySeconds: z.number().int().min(1),
|
||||
prompt: z.string().min(1),
|
||||
stop: z.boolean().optional(),
|
||||
reason: z.string().optional(),
|
||||
});
|
||||
@@ -1,121 +0,0 @@
|
||||
import { mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs";
|
||||
import os from "node:os";
|
||||
import path from "node:path";
|
||||
import { afterEach, beforeEach, describe, expect, it } from "vitest";
|
||||
import { detectEol, editFileTool, fromLF, toLF } from "./editFile.js";
|
||||
import type { ToolContext } from "./types.js";
|
||||
|
||||
describe("editFile tool — path containment", () => {
|
||||
let cwd: string;
|
||||
let ctx: ToolContext;
|
||||
|
||||
beforeEach(() => {
|
||||
cwd = mkdtempSync(path.join(os.tmpdir(), "locode-editfile-"));
|
||||
ctx = { cwd };
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
rmSync(cwd, { recursive: true, force: true });
|
||||
});
|
||||
|
||||
it("edits a file inside the working directory", async () => {
|
||||
writeFileSync(path.join(cwd, "note.txt"), "hello world");
|
||||
const result = (await editFileTool.handler({ path: "note.txt", old_string: "world", new_string: "there" }, ctx)) as { replacements: number };
|
||||
expect(result.replacements).toBe(1);
|
||||
});
|
||||
|
||||
it("refuses to edit a file outside the working directory via ../ traversal", async () => {
|
||||
await expect(
|
||||
editFileTool.handler({ path: "../escape.txt", old_string: "a", new_string: "b" }, ctx),
|
||||
).rejects.toThrow(/outside the working directory/);
|
||||
});
|
||||
|
||||
it("preview reports the block instead of reading the target file", async () => {
|
||||
const preview = await editFileTool.preview!({ path: "../escape.txt", old_string: "a", new_string: "b" }, ctx);
|
||||
expect(preview).toMatch(/outside the working directory/);
|
||||
});
|
||||
|
||||
it("handler suggests the closest match when old_string is not found", async () => {
|
||||
writeFileSync(
|
||||
path.join(cwd, "code.ts"),
|
||||
"function greet(name: string): string {\n return `Hello, ${name}!`;\n}\n",
|
||||
);
|
||||
// Close but wrong: single quotes instead of backticks, "Hi" instead of "Hello".
|
||||
await expect(
|
||||
editFileTool.handler(
|
||||
{ path: "code.ts", old_string: "return 'Hi, ${name}!';", new_string: "return `Hi, ${name}!`;" },
|
||||
ctx,
|
||||
),
|
||||
).rejects.toThrow(/closest match/);
|
||||
});
|
||||
|
||||
it("preview warns and shows the closest match when old_string is not found", async () => {
|
||||
writeFileSync(path.join(cwd, "note.txt"), "the quick brown fox jumps over the lazy dog");
|
||||
const preview = await editFileTool.preview!(
|
||||
{ path: "note.txt", old_string: "the quick red fox jumps over the lazy cat", new_string: "x" },
|
||||
ctx,
|
||||
);
|
||||
expect(preview).toMatch(/not found/);
|
||||
expect(preview).toMatch(/closest match/);
|
||||
expect(preview).toContain("quick brown fox");
|
||||
});
|
||||
|
||||
it("does not suggest a match when nothing is remotely similar", async () => {
|
||||
writeFileSync(path.join(cwd, "note.txt"), "aaaaaaaaaaaaaaaaaaaaaaaa");
|
||||
await expect(
|
||||
editFileTool.handler(
|
||||
{ path: "note.txt", old_string: "completely different text xyz", new_string: "b" },
|
||||
ctx,
|
||||
),
|
||||
).rejects.toThrow(/not found/);
|
||||
// No "closest match" suffix when similarity is below the threshold.
|
||||
await expect(
|
||||
editFileTool.handler(
|
||||
{ path: "note.txt", old_string: "completely different text xyz", new_string: "b" },
|
||||
ctx,
|
||||
),
|
||||
).rejects.not.toThrow(/closest match/);
|
||||
});
|
||||
|
||||
// read_file shows the model LF-normalized content regardless of the file's real line endings
|
||||
// (see readFile.ts), so old_string/new_string from a model are always LF — matching must happen
|
||||
// in that same space or every CRLF file in a project like this one fails with "not found".
|
||||
it("matches an LF old_string against a CRLF file (mirrors what read_file shows the model)", async () => {
|
||||
writeFileSync(path.join(cwd, "code.ts"), "function greet() {\r\n return 1;\r\n}\r\n");
|
||||
const result = (await editFileTool.handler(
|
||||
{ path: "code.ts", old_string: " return 1;", new_string: " return 2;" },
|
||||
ctx,
|
||||
)) as { replacements: number };
|
||||
expect(result.replacements).toBe(1);
|
||||
const onDisk = readFileSync(path.join(cwd, "code.ts"), "utf-8");
|
||||
expect(onDisk).toBe("function greet() {\r\n return 2;\r\n}\r\n");
|
||||
});
|
||||
|
||||
it("preserves CRLF line endings on disk after an edit spanning multiple lines", async () => {
|
||||
writeFileSync(path.join(cwd, "code.ts"), "a\r\nb\r\nc\r\n");
|
||||
await editFileTool.handler({ path: "code.ts", old_string: "a\nb", new_string: "a\nx\nb" }, ctx);
|
||||
const onDisk = readFileSync(path.join(cwd, "code.ts"), "utf-8");
|
||||
expect(onDisk).toBe("a\r\nx\r\nb\r\nc\r\n");
|
||||
});
|
||||
|
||||
it("leaves a pure-LF file untouched by EOL conversion", async () => {
|
||||
writeFileSync(path.join(cwd, "code.ts"), "a\nb\nc\n");
|
||||
await editFileTool.handler({ path: "code.ts", old_string: "b", new_string: "x" }, ctx);
|
||||
const onDisk = readFileSync(path.join(cwd, "code.ts"), "utf-8");
|
||||
expect(onDisk).toBe("a\nx\nc\n");
|
||||
});
|
||||
});
|
||||
|
||||
describe("EOL helpers", () => {
|
||||
it("detectEol finds CRLF, defaults to LF otherwise", () => {
|
||||
expect(detectEol("a\r\nb")).toBe("\r\n");
|
||||
expect(detectEol("a\nb")).toBe("\n");
|
||||
expect(detectEol("a")).toBe("\n");
|
||||
});
|
||||
|
||||
it("toLF/fromLF round-trip", () => {
|
||||
expect(toLF("a\r\nb\r\nc")).toBe("a\nb\nc");
|
||||
expect(fromLF("a\nb\nc", "\r\n")).toBe("a\r\nb\r\nc");
|
||||
expect(fromLF("a\nb\nc", "\n")).toBe("a\nb\nc");
|
||||
});
|
||||
});
|
||||
+23
-159
@@ -1,8 +1,9 @@
|
||||
import { createPatch } from "diff";
|
||||
import { randomBytes } from "node:crypto";
|
||||
import { readFile as fsReadFile, rename as fsRename, unlink as fsUnlink, writeFile as fsWriteFile } from "node:fs/promises";
|
||||
import path from "node:path";
|
||||
import { z } from "zod";
|
||||
import { resolveWithinCwd } from "./pathGuard.js";
|
||||
import { assertWithinWorkspace } from "../utils/path.js";
|
||||
import type { ToolDef } from "./types.js";
|
||||
|
||||
const schema = z.object({
|
||||
@@ -12,26 +13,6 @@ const schema = z.object({
|
||||
replace_all: z.boolean().optional().describe("Replace every occurrence instead of requiring a unique match."),
|
||||
});
|
||||
|
||||
/** `read_file` shows the model LF-normalized content (`content.split(/\r?\n/).join(...)` — see
|
||||
* readFile.ts), regardless of the file's actual line endings on disk. A model's `old_string`/
|
||||
* `new_string` are built from what it read, so they're always LF. Matching that against this
|
||||
* tool's raw (real `\r\n`-preserving) file read would fail on every CRLF file in the project —
|
||||
* which is most of them (see the repo's CRLF/LF notes). Detect the file's line ending once, do
|
||||
* all matching/editing in LF space (so `old_string` from the model lines up), then convert the
|
||||
* result back before writing so the file's on-disk convention is preserved rather than silently
|
||||
* flipped to LF. */
|
||||
export function detectEol(raw: string): "\r\n" | "\n" {
|
||||
return raw.includes("\r\n") ? "\r\n" : "\n";
|
||||
}
|
||||
|
||||
export function toLF(s: string): string {
|
||||
return s.replace(/\r\n/g, "\n");
|
||||
}
|
||||
|
||||
export function fromLF(s: string, eol: "\r\n" | "\n"): string {
|
||||
return eol === "\n" ? s : s.replace(/\n/g, eol);
|
||||
}
|
||||
|
||||
export function countOccurrences(haystack: string, needle: string): number {
|
||||
return needle === "" ? 0 : haystack.split(needle).length - 1;
|
||||
}
|
||||
@@ -44,161 +25,41 @@ export function applyEdit(original: string, oldString: string, newString: string
|
||||
return replaceAll ? original.split(oldString).join(newString) : original.replace(oldString, () => newString);
|
||||
}
|
||||
|
||||
/**
|
||||
* Normalise a candidate snippet for fuzzy comparison: collapse runs of whitespace to single spaces
|
||||
* and trim. This makes the similarity score tolerant to indentation/line-ending differences, which
|
||||
* are the most common reasons a local model's old_string almost-matches but not quite.
|
||||
*/
|
||||
function normaliseForCompare(s: string): string {
|
||||
return s.replace(/\s+/g, " ").trim();
|
||||
}
|
||||
|
||||
/**
|
||||
* Compute a Levenshtein distance limited to `maxDist` — early-exits once the distance exceeds it,
|
||||
* making it O(n*m) worst case but far cheaper in practice when we only care about "close enough".
|
||||
*/
|
||||
function boundedLevenshtein(a: string, b: string, maxDist: number): number {
|
||||
const al = a.length;
|
||||
const bl = b.length;
|
||||
if (Math.abs(al - bl) > maxDist) return maxDist + 1;
|
||||
if (al === 0) return bl;
|
||||
if (bl === 0) return al;
|
||||
let prev: number[] = new Array<number>(bl + 1);
|
||||
let curr: number[] = new Array<number>(bl + 1);
|
||||
for (let j = 0; j <= bl; j++) prev[j] = j;
|
||||
for (let i = 1; i <= al; i++) {
|
||||
curr[0] = i;
|
||||
let rowMin = i;
|
||||
const ai = a.charCodeAt(i - 1);
|
||||
for (let j = 1; j <= bl; j++) {
|
||||
const cost = ai === b.charCodeAt(j - 1) ? 0 : 1;
|
||||
const del = (prev[j] ?? 0) + 1;
|
||||
const ins = (curr[j - 1] ?? 0) + 1;
|
||||
const sub = (prev[j - 1] ?? 0) + cost;
|
||||
curr[j] = Math.min(del, ins, sub);
|
||||
const cell = curr[j] ?? 0;
|
||||
if (cell < rowMin) rowMin = cell;
|
||||
}
|
||||
// If every cell in this row already exceeds maxDist, the final answer can only be worse.
|
||||
if (rowMin > maxDist) return maxDist + 1;
|
||||
[prev, curr] = [curr, prev];
|
||||
}
|
||||
return prev[bl] ?? maxDist + 1;
|
||||
}
|
||||
|
||||
interface SimilarMatch {
|
||||
/** 0..1 similarity ratio (1 = identical, 0 = unrelated). */
|
||||
score: number;
|
||||
/** The exact text from the file at the best matching window. */
|
||||
snippet: string;
|
||||
/** 1-indexed line number where the snippet starts. */
|
||||
line: number;
|
||||
}
|
||||
|
||||
/**
|
||||
* Find the region of `content` most similar to `needle`. Slides a window of the needle's length
|
||||
* (±50%) across the file in word steps, scoring normalised text with bounded Levenshtein. Returns
|
||||
* the best candidate when its similarity is at least 0.5 — clearly worth suggesting to the model.
|
||||
* Returns null when nothing is close enough, in which case the caller falls back to the plain
|
||||
* "not found" message.
|
||||
*/
|
||||
function findSimilarMatch(content: string, needle: string): SimilarMatch | null {
|
||||
const needleNorm = normaliseForCompare(needle);
|
||||
if (needleNorm.length < 3) return null;
|
||||
|
||||
const words = needleNorm.split(" ");
|
||||
const minLen = Math.floor(needleNorm.length * 0.5);
|
||||
const maxLen = Math.ceil(needleNorm.length * 1.5);
|
||||
|
||||
let best: SimilarMatch | null = null;
|
||||
let bestDist = Infinity;
|
||||
|
||||
// Walk the file by character, treating every position as a potential window start is O(n*len)
|
||||
// and too slow for big files. Instead, step at every Nth character (≈ word boundaries) to keep
|
||||
// it cheap while still landing near real matches.
|
||||
const step = Math.max(1, Math.floor(needleNorm.length / 8));
|
||||
const contentLen = content.length;
|
||||
|
||||
for (let start = 0; start < contentLen; start += step) {
|
||||
for (let len = minLen; len <= maxLen; len += step) {
|
||||
const end = Math.min(start + len, contentLen);
|
||||
const candidate = content.slice(start, end);
|
||||
const candNorm = normaliseForCompare(candidate);
|
||||
if (candNorm.length < minLen) continue;
|
||||
|
||||
// Only spend Levenshtein effort if the lengths are plausibly close.
|
||||
const maxDist = Math.floor(needleNorm.length * 0.5);
|
||||
const dist = boundedLevenshtein(needleNorm, candNorm, maxDist);
|
||||
if (dist >= bestDist) continue;
|
||||
|
||||
bestDist = dist;
|
||||
const score = 1 - dist / Math.max(needleNorm.length, candNorm.length);
|
||||
// 1-indexed line: count newlines before `start`.
|
||||
let line = 1;
|
||||
for (let k = 0; k < start; k++) if (content.charCodeAt(k) === 10) line++;
|
||||
best = { score, snippet: candidate.trim(), line };
|
||||
}
|
||||
}
|
||||
|
||||
if (best && best.score >= 0.5) return best;
|
||||
return null;
|
||||
}
|
||||
|
||||
/** Build a "did you mean" suffix for error/preview messages. Returns "" if nothing useful. */
|
||||
function similarHint(original: string, oldString: string): string {
|
||||
const m = findSimilarMatch(original, oldString);
|
||||
if (!m) return "";
|
||||
// Truncate long snippets so the message stays readable.
|
||||
const snippet =
|
||||
m.snippet.length > 300 ? `${m.snippet.slice(0, 300)}…` : m.snippet;
|
||||
return `\n\nThe closest match in the file (line ${m.line}, ~${Math.round(m.score * 100)}% similar):\n"""\n${snippet}\n"""\nUse this exact text (or a unique subset of it) as old_string.`;
|
||||
}
|
||||
|
||||
export const editFileTool: ToolDef<z.infer<typeof schema>> = {
|
||||
name: "edit_file",
|
||||
description:
|
||||
"Replace exact text in a file. old_string must match exactly. Unless replace_all is set, it must be unique — include enough context. " +
|
||||
"Use for small, targeted changes to an existing file. On a mismatch, the closest similar text is suggested to help retry.",
|
||||
"Replace exact text in a file. old_string must match exactly. Read the file first, then use enough " +
|
||||
"surrounding context in old_string to make it unique. Unless replace_all is set, duplicate matches are " +
|
||||
"rejected. Use this for small, targeted changes; prefer write_file for new files or full rewrites.",
|
||||
schema,
|
||||
mutating: true,
|
||||
preview: async ({ path: filePath, old_string, new_string, replace_all }, ctx) => {
|
||||
let resolved: string;
|
||||
const resolved = path.resolve(ctx.cwd, filePath);
|
||||
assertWithinWorkspace(resolved, ctx.cwd, filePath);
|
||||
let original: string;
|
||||
try {
|
||||
resolved = resolveWithinCwd(ctx.cwd, filePath);
|
||||
} catch (err) {
|
||||
return (err as Error).message;
|
||||
}
|
||||
let raw: string;
|
||||
try {
|
||||
raw = await fsReadFile(resolved, "utf-8");
|
||||
original = await fsReadFile(resolved, "utf-8");
|
||||
} catch {
|
||||
return `File ${resolved} does not exist.`;
|
||||
}
|
||||
const eol = detectEol(raw);
|
||||
const original = toLF(raw);
|
||||
const oldLF = toLF(old_string);
|
||||
const newLF = toLF(new_string);
|
||||
const occurrences = countOccurrences(original, oldLF);
|
||||
const occurrences = countOccurrences(original, old_string);
|
||||
if (occurrences === 0) {
|
||||
return `Warning: old_string not found in ${resolved} — this edit will fail.${similarHint(original, oldLF)}`;
|
||||
return `Warning: old_string not found in ${resolved} — this edit will fail.`;
|
||||
}
|
||||
if (occurrences > 1 && !replace_all) {
|
||||
return `Warning: old_string appears ${occurrences} times in ${resolved} — this edit will fail unless replace_all is set.`;
|
||||
}
|
||||
const updated = fromLF(applyEdit(original, oldLF, newLF, replace_all), eol);
|
||||
return createPatch(resolved, raw, updated, "", "");
|
||||
const updated = applyEdit(original, old_string, new_string, replace_all);
|
||||
return createPatch(resolved, original, updated, "", "");
|
||||
},
|
||||
handler: async ({ path: filePath, old_string, new_string, replace_all }, ctx) => {
|
||||
const resolved = resolveWithinCwd(ctx.cwd, filePath);
|
||||
const raw = await fsReadFile(resolved, "utf-8");
|
||||
const eol = detectEol(raw);
|
||||
const original = toLF(raw);
|
||||
const oldLF = toLF(old_string);
|
||||
const newLF = toLF(new_string);
|
||||
const occurrences = countOccurrences(original, oldLF);
|
||||
const resolved = path.resolve(ctx.cwd, filePath);
|
||||
assertWithinWorkspace(resolved, ctx.cwd, filePath);
|
||||
const original = await fsReadFile(resolved, "utf-8");
|
||||
const occurrences = countOccurrences(original, old_string);
|
||||
if (occurrences === 0) {
|
||||
throw new Error(
|
||||
`old_string not found in ${filePath}. Make sure it matches the file exactly, including whitespace.${similarHint(original, oldLF)}`,
|
||||
`old_string not found in ${filePath}. Make sure it matches the file exactly, including whitespace.`,
|
||||
);
|
||||
}
|
||||
if (occurrences > 1 && !replace_all) {
|
||||
@@ -206,7 +67,7 @@ export const editFileTool: ToolDef<z.infer<typeof schema>> = {
|
||||
`old_string appears ${occurrences} times in ${filePath}. Provide more surrounding context to make it unique, or set replace_all: true.`,
|
||||
);
|
||||
}
|
||||
const updated = fromLF(applyEdit(original, oldLF, newLF, replace_all), eol);
|
||||
const updated = applyEdit(original, old_string, new_string, replace_all);
|
||||
// Write to a temp file in the same directory, then rename — rename is atomic within a single
|
||||
// directory, so a crash mid-write can't leave the user's source file half-overwritten (the live
|
||||
// file stays intact until the rename swaps in the full new content). Clean up the temp file if
|
||||
@@ -219,6 +80,9 @@ export const editFileTool: ToolDef<z.infer<typeof schema>> = {
|
||||
await fsUnlink(tmp).catch(() => {});
|
||||
throw err;
|
||||
}
|
||||
if (ctx.setLastEdit) {
|
||||
ctx.setLastEdit({ path: filePath, previousContent: original });
|
||||
}
|
||||
return { path: resolved, replacements: replace_all ? occurrences : 1 };
|
||||
},
|
||||
};
|
||||
};
|
||||
|
||||
@@ -0,0 +1,44 @@
|
||||
import { z } from "zod";
|
||||
import type { ToolDef } from "./types.js";
|
||||
|
||||
const schema = z.object({
|
||||
plan: z
|
||||
.string()
|
||||
.describe(
|
||||
"The full implementation plan in prose: the files you would change, the approach for each, and the key edits. " +
|
||||
"Be concrete and actionable so the user can review it at a glance.",
|
||||
),
|
||||
});
|
||||
|
||||
/** The structured way to exit plan mode: the model calls this once it has finished researching and
|
||||
* has a concrete plan, instead of presenting the plan as a final prose message. The handler shows
|
||||
* the plan to the user via the same Approve/Reject prompt the prose path uses (see maybePresentPlan
|
||||
* in agent/loop.ts); on approval plan mode ends and the model proceeds to implement in the same
|
||||
* turn, on rejection it stays in plan mode and can refine. The prose path remains as a fallback for
|
||||
* models that present a plan without calling this tool. Non-mutating: it changes permission mode, not
|
||||
* the filesystem, so it passes the plan-mode mutating-tool gate. Only callable in plan mode — the
|
||||
* ctx callback is withheld otherwise, so a call at the wrong time returns a clear error. */
|
||||
export const exitPlanModeTool: ToolDef<z.infer<typeof schema>> = {
|
||||
name: "exit_plan_mode",
|
||||
description:
|
||||
"Exit plan mode by presenting your implementation plan for the user's approval. Call this once you've " +
|
||||
"finished researching and have a concrete plan (which files you'd change and how). On approval, plan mode " +
|
||||
"ends and you implement the plan in this same turn. On rejection, stay in plan mode, refine the plan (explore " +
|
||||
"more if needed), and call exit_plan_mode again. Only available in plan mode — don't call it otherwise.",
|
||||
schema,
|
||||
mutating: false,
|
||||
handler: async (args, ctx) => {
|
||||
if (!ctx.exitPlanMode) {
|
||||
return { error: "exit_plan_mode is only available while plan mode is active." };
|
||||
}
|
||||
const { approved } = await ctx.exitPlanMode(args.plan);
|
||||
if (approved) {
|
||||
return { result: "Plan approved. Plan mode is now off — proceed to implement the plan now." };
|
||||
}
|
||||
return {
|
||||
result:
|
||||
"The user rejected the plan. Stay in plan mode, refine it (explore more if needed), and call " +
|
||||
"exit_plan_mode again when ready. Do not call any mutating tool yet.",
|
||||
};
|
||||
},
|
||||
};
|
||||
+4
-2
@@ -29,7 +29,8 @@ const statusSchema = z.object({
|
||||
export const gitStatusTool: ToolDef<z.infer<typeof statusSchema>> = {
|
||||
name: "git_status",
|
||||
description:
|
||||
"Inspect the git repo: `status`, `diff`, `log`, `show`, or `branches`. Read-only — no confirmation needed.",
|
||||
"Inspect the git repo: `status`, `diff`, `log`, `show`, or `branches`. Read-only — no confirmation needed. " +
|
||||
"Always check status/diff before mutating git_commit operations.",
|
||||
schema: statusSchema,
|
||||
mutating: false,
|
||||
handler: async ({ operation, paths, staged, ref, maxCount }, ctx) => {
|
||||
@@ -109,7 +110,8 @@ export const gitCommitTool: ToolDef<z.infer<typeof commitSchema>> = {
|
||||
name: "git_commit",
|
||||
description:
|
||||
"Git operations: `add`, `commit`, `create_branch`, `checkout`, `push`, `reset`, `stash`, `merge`, `rebase`, " +
|
||||
"`delete_branch`. Mutating operations require user confirmation with a preview.",
|
||||
"`delete_branch`. Mutating operations require user confirmation with a preview. For commit, run git_status " +
|
||||
"or git_commit add first, then provide a clear, concise message.",
|
||||
schema: commitSchema,
|
||||
mutating: true,
|
||||
preview: async (args, ctx) => buildPreview(args, ctx.cwd),
|
||||
|
||||
+4
-1
@@ -14,7 +14,10 @@ const schema = z.object({
|
||||
|
||||
export const grepTool: ToolDef<z.infer<typeof schema>> = {
|
||||
name: "grep",
|
||||
description: "Search file contents for a regular expression pattern using ripgrep. Use to find where a symbol/function/word is used across the codebase, or to locate files containing specific text. Faster than reading files one by one.",
|
||||
description:
|
||||
"Search file contents for a regular expression pattern using ripgrep. This is the best first step " +
|
||||
"when exploring a codebase: use it to find where a symbol, function, or pattern is used, then " +
|
||||
"read only the relevant files. If results are too broad, refine with `path` or `glob`.",
|
||||
schema,
|
||||
mutating: false,
|
||||
handler: async ({ pattern, path: searchPath, glob, case_insensitive, max_results }, ctx) => {
|
||||
|
||||
+23
-8
@@ -1,19 +1,25 @@
|
||||
import { agentTool } from "./agentTool.js";
|
||||
import { definitionTool, referencesTool, diagnosticsTool } from "./codeIntel.js";
|
||||
import { askQuestionTool } from "./askQuestion.js";
|
||||
import { bashTool } from "./bash.js";
|
||||
import { bashKillTool } from "./bashKill.js";
|
||||
import { bashOutputTool } from "./bashOutput.js";
|
||||
import { cronCreateTool, cronDeleteTool, cronListTool, scheduleWakeupTool } from "./cron.js";
|
||||
import { editFileTool } from "./editFile.js";
|
||||
import { multiEditTool } from "./multiEdit.js";
|
||||
import { exitPlanModeTool } from "./exitPlanMode.js";
|
||||
import { gitCommitTool, gitStatusTool } from "./git.js";
|
||||
import { grepTool } from "./grep.js";
|
||||
import { listFilesTool } from "./listFiles.js";
|
||||
import { memoryTool, memoryWriteTool } from "./memory.js";
|
||||
import { multiEditTool } from "./multiEdit.js";
|
||||
import { notebookEditTool } from "./notebookEdit.js";
|
||||
import { readFileTool } from "./readFile.js";
|
||||
import { todoWriteTool } from "./todoWrite.js";
|
||||
import { taskCreateTool, taskListTool, taskGetTool, taskUpdateTool } from "./task.js";
|
||||
import { sendMessageTool } from "./sendMessage.js";
|
||||
import { listTeammatesTool } from "./teammates.js";
|
||||
import { taskCreateTool, taskGetTool, taskListTool, taskUpdateTool } from "./task.js";
|
||||
import { webFetchTool } from "./webFetch.js";
|
||||
import { workflowTool } from "./workflow.js";
|
||||
import { webSearchTool } from "./webSearch.js";
|
||||
import { enterWorktreeTool, exitWorktreeTool } from "./worktreeSession.js";
|
||||
import { writeFileTool } from "./writeFile.js";
|
||||
import type { ToolDef } from "./types.js";
|
||||
|
||||
@@ -21,9 +27,6 @@ export const TOOLS: ToolDef[] = [
|
||||
readFileTool,
|
||||
listFilesTool,
|
||||
grepTool,
|
||||
definitionTool,
|
||||
referencesTool,
|
||||
diagnosticsTool,
|
||||
webSearchTool,
|
||||
webFetchTool,
|
||||
gitStatusTool,
|
||||
@@ -35,12 +38,24 @@ export const TOOLS: ToolDef[] = [
|
||||
bashOutputTool,
|
||||
bashKillTool,
|
||||
gitCommitTool,
|
||||
todoWriteTool,
|
||||
taskCreateTool,
|
||||
taskListTool,
|
||||
taskGetTool,
|
||||
taskUpdateTool,
|
||||
memoryTool,
|
||||
memoryWriteTool,
|
||||
agentTool,
|
||||
sendMessageTool,
|
||||
listTeammatesTool,
|
||||
workflowTool,
|
||||
exitPlanModeTool,
|
||||
askQuestionTool,
|
||||
cronCreateTool,
|
||||
cronListTool,
|
||||
cronDeleteTool,
|
||||
scheduleWakeupTool,
|
||||
enterWorktreeTool,
|
||||
exitWorktreeTool,
|
||||
];
|
||||
|
||||
export const TOOL_REGISTRY: Map<string, ToolDef> = new Map(TOOLS.map((t) => [t.name, t]));
|
||||
|
||||
@@ -17,7 +17,10 @@ const MAX_MATCHES = 500;
|
||||
|
||||
export const listFilesTool: ToolDef<z.infer<typeof schema>> = {
|
||||
name: "list_files",
|
||||
description: "List files matching a glob pattern (e.g. `src/**/*.ts`). Use to explore the project structure or find files by name/extension before reading them.",
|
||||
description:
|
||||
"List files matching a glob pattern. Use this to understand directory structure or find files " +
|
||||
"by name. For searching file contents, use grep instead. Large result sets are truncated; narrow " +
|
||||
"the pattern if you get too many matches.",
|
||||
schema,
|
||||
mutating: false,
|
||||
handler: async ({ pattern, cwd }, ctx) => {
|
||||
|
||||
@@ -0,0 +1,123 @@
|
||||
import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
|
||||
import { existsSync, mkdtempSync, readFileSync, rmSync } from "node:fs";
|
||||
import path from "node:path";
|
||||
import os from "node:os";
|
||||
import { _setConfigFilePathForTest } from "../config/store.js";
|
||||
import { userMemoryDir, userMemoryIndexPath } from "../utils/userMemory.js";
|
||||
import { memoryTool, memoryWriteTool } from "./memory.js";
|
||||
import type { ToolContext } from "./types.js";
|
||||
|
||||
let tempDir: string;
|
||||
const noCtx = {} as ToolContext;
|
||||
|
||||
beforeEach(() => {
|
||||
tempDir = mkdtempSync(path.join(os.tmpdir(), "locode-memory-tool-"));
|
||||
_setConfigFilePathForTest(path.join(tempDir, "config.json"));
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
_setConfigFilePathForTest(undefined);
|
||||
if (existsSync(tempDir)) rmSync(tempDir, { recursive: true, force: true });
|
||||
});
|
||||
|
||||
describe("memory tool (read)", () => {
|
||||
it("reports empty memory when nothing is saved", async () => {
|
||||
const result = await memoryTool.handler({}, noCtx);
|
||||
expect(result).toEqual({ content: "(memory is empty)", count: 0 });
|
||||
});
|
||||
|
||||
it("reads a specific fact by name", async () => {
|
||||
await memoryWriteTool.handler(
|
||||
{ action: "write", name: "prefers-concise", description: "short answers", type: "user", content: "Be brief." },
|
||||
noCtx,
|
||||
);
|
||||
const result = (await memoryTool.handler({ name: "prefers-concise" }, noCtx)) as { content: string };
|
||||
expect(result.content).toContain("Be brief.");
|
||||
expect(result.content).toContain("name: prefers-concise");
|
||||
});
|
||||
|
||||
it("returns a not-found message for an unknown name", async () => {
|
||||
const result = (await memoryTool.handler({ name: "nope" }, noCtx)) as { content: string };
|
||||
expect(result.content).toContain("No memory fact named 'nope'");
|
||||
});
|
||||
|
||||
it("lists all facts with no argument", async () => {
|
||||
await memoryWriteTool.handler({ action: "write", name: "a-fact", description: "a", type: "user", content: "aa" }, noCtx);
|
||||
await memoryWriteTool.handler({ action: "write", name: "b-fact", description: "b", type: "project", content: "bb" }, noCtx);
|
||||
const result = (await memoryTool.handler({}, noCtx)) as { content: string; count: number };
|
||||
expect(result.count).toBe(2);
|
||||
expect(result.content).toContain("aa");
|
||||
expect(result.content).toContain("bb");
|
||||
});
|
||||
});
|
||||
|
||||
describe("memory_write tool (write/delete)", () => {
|
||||
it("creates a typed fact file + index on write", async () => {
|
||||
const result = await memoryWriteTool.handler(
|
||||
{ action: "write", name: "react-stack", description: "uses react", type: "project", content: "Stack is React + vitest." },
|
||||
noCtx,
|
||||
);
|
||||
expect(result).toEqual({ written: "react-stack", type: "project", indexUpdated: true });
|
||||
expect(existsSync(path.join(userMemoryDir(), "react-stack.md"))).toBe(true);
|
||||
expect(readFileSync(userMemoryIndexPath(), "utf-8")).toContain("react-stack.md");
|
||||
});
|
||||
|
||||
it("overwrites an existing fact on write", async () => {
|
||||
await memoryWriteTool.handler({ action: "write", name: "flip", description: "old", type: "user", content: "old body" }, noCtx);
|
||||
await memoryWriteTool.handler({ action: "write", name: "flip", description: "new", type: "reference", content: "new body" }, noCtx);
|
||||
const file = readFileSync(path.join(userMemoryDir(), "flip.md"), "utf-8");
|
||||
expect(file).toContain("new body");
|
||||
expect(file).toContain("description: new");
|
||||
// Index has a single line for flip.
|
||||
expect(readFileSync(userMemoryIndexPath(), "utf-8").match(/flip\.md/g)).toHaveLength(1);
|
||||
});
|
||||
|
||||
it("rejects a write missing required fields", async () => {
|
||||
const result = await memoryWriteTool.handler({ action: "write", name: "x", content: "y" }, noCtx);
|
||||
expect(result).toEqual({ error: "write requires non-empty 'description', 'type', and 'content'." });
|
||||
});
|
||||
|
||||
it("deletes an existing fact and its index line", async () => {
|
||||
await memoryWriteTool.handler({ action: "write", name: "gone", description: "x", type: "user", content: "yy" }, noCtx);
|
||||
const result = await memoryWriteTool.handler({ action: "delete", name: "gone" }, noCtx);
|
||||
expect(result).toEqual({ deleted: "gone" });
|
||||
expect(existsSync(path.join(userMemoryDir(), "gone.md"))).toBe(false);
|
||||
expect(readFileSync(userMemoryIndexPath(), "utf-8")).not.toContain("gone");
|
||||
});
|
||||
|
||||
it("reports an error deleting a missing fact", async () => {
|
||||
const result = await memoryWriteTool.handler({ action: "delete", name: "nope" }, noCtx);
|
||||
expect(result).toEqual({ error: "No memory fact named 'nope' to delete." });
|
||||
});
|
||||
|
||||
it("produces a diff preview for a write", async () => {
|
||||
const preview = await memoryWriteTool.preview!(
|
||||
{ action: "write", name: "fresh", description: "d", type: "user", content: "body text" },
|
||||
noCtx,
|
||||
);
|
||||
expect(preview).toContain("Create memory fact fresh.md");
|
||||
expect(preview).toContain("body text");
|
||||
});
|
||||
|
||||
it("produces a diff preview when overwriting an existing fact", async () => {
|
||||
await memoryWriteTool.handler({ action: "write", name: "p", description: "d", type: "user", content: "old" }, noCtx);
|
||||
const preview = await memoryWriteTool.preview!(
|
||||
{ action: "write", name: "p", description: "d", type: "user", content: "new" },
|
||||
noCtx,
|
||||
);
|
||||
expect(preview).toContain("@@");
|
||||
expect(preview).toContain("+new");
|
||||
expect(preview).toContain("-old");
|
||||
});
|
||||
|
||||
it("refreshes the session user-memory cache via setUserMemory after a write", async () => {
|
||||
const setUserMemory = vi.fn();
|
||||
const ctx = { setUserMemory } as unknown as ToolContext;
|
||||
await memoryWriteTool.handler({ action: "write", name: "cached", description: "d", type: "user", content: "remember this" }, ctx);
|
||||
expect(setUserMemory).toHaveBeenCalledOnce();
|
||||
// The refreshed cache is the bounded index form loadUserMemory produces, carrying the new line.
|
||||
const arg = setUserMemory.mock.calls[0]![0] as string | null;
|
||||
expect(arg).toContain("Personal memory index");
|
||||
expect(arg).toContain("cached");
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,131 @@
|
||||
import { createPatch } from "diff";
|
||||
import { z } from "zod";
|
||||
import {
|
||||
deleteMemoryEntry,
|
||||
listMemoryEntries,
|
||||
loadUserMemory,
|
||||
MEMORY_TYPES,
|
||||
readMemoryEntry,
|
||||
writeMemoryEntry,
|
||||
type MemoryType,
|
||||
} from "../utils/userMemory.js";
|
||||
import type { ToolDef } from "./types.js";
|
||||
|
||||
// The memory files live in the user's locode config dir (see utils/userMemory.ts), NOT under the
|
||||
// project workspace. So unlike write_file/edit_file these tools do NOT call assertWithinWorkspace —
|
||||
// they always operate on files under the fixed `memory/` dir resolved via userMemoryDir() (test-
|
||||
// override-aware), and the only user-supplied identifier is a kebab-case `name` validated against
|
||||
// [a-z0-9-]+, so there's no traversal surface.
|
||||
//
|
||||
// Split into two tools because ToolDef.mutating is a static per-tool flag (the confirmation/plan-mode
|
||||
// gate keys off it): `memory` is a read-only no-prompt tool usable even in plan mode, while
|
||||
// `memory_write` mutates the user-level files and goes through the normal confirm gate.
|
||||
|
||||
async function readIndexOrLegacy(): Promise<string> {
|
||||
const loaded = await loadUserMemory();
|
||||
return loaded ?? "(memory is empty)";
|
||||
}
|
||||
|
||||
export const memoryTool: ToolDef<{ name?: string }> = {
|
||||
name: "memory",
|
||||
description:
|
||||
"Read the user's personal memory (the typed per-fact files in the locode config dir's memory/ directory). " +
|
||||
"With no arguments, lists every saved fact (name, type, description, and full body). With a 'name', reads just " +
|
||||
"that one fact. The lightweight index is already in your system prompt each turn — use this tool to pull a " +
|
||||
"fact's full body when the index hook tells you it's relevant. Read-only — use memory_write to save or delete.",
|
||||
schema: z.object({
|
||||
name: z
|
||||
.string()
|
||||
.optional()
|
||||
.describe("The slug name of a specific fact to read. Omit to list all saved facts."),
|
||||
}),
|
||||
mutating: false,
|
||||
handler: async (args) => {
|
||||
if (args.name) {
|
||||
const entry = await readMemoryEntry(args.name);
|
||||
if (!entry) return { content: `No memory fact named '${args.name}'.` };
|
||||
return { content: `---\nname: ${entry.name}\ndescription: ${entry.description}\ntype: ${entry.type}\n---\n\n${entry.body}` };
|
||||
}
|
||||
const entries = await listMemoryEntries();
|
||||
if (entries.length === 0) {
|
||||
// No typed facts — surface the legacy freeform file if one exists so the model isn't blind to it.
|
||||
const legacy = await readIndexOrLegacy();
|
||||
return { content: legacy, count: 0 };
|
||||
}
|
||||
const rendered = entries.map((e) => `${e.name} (${e.type}): ${e.description}\n${e.body}`).join("\n\n---\n\n");
|
||||
return { content: rendered, count: entries.length };
|
||||
},
|
||||
};
|
||||
|
||||
const writeSchema = z.object({
|
||||
action: z
|
||||
.enum(["write", "delete"])
|
||||
.describe("write: create or overwrite a typed memory fact; delete: remove one."),
|
||||
name: z
|
||||
.string()
|
||||
.describe("The fact's slug (kebab-case, [a-z0-9-]+). Used as the filename and the frontmatter 'name'."),
|
||||
description: z
|
||||
.string()
|
||||
.optional()
|
||||
.describe("One-line summary used as the recall hook in the always-in-prompt index. Required for action='write'."),
|
||||
type: z
|
||||
.enum(MEMORY_TYPES as [MemoryType, ...MemoryType[]])
|
||||
.optional()
|
||||
.describe("Fact category: user (who the user is), feedback (working-style guidance), project (ongoing work), reference (external pointers). Required for action='write'."),
|
||||
content: z
|
||||
.string()
|
||||
.optional()
|
||||
.describe("The fact body. Required for action='write'. Keep it a concise, self-contained fact; for feedback/project include a 'Why:' and 'How to apply:' line."),
|
||||
});
|
||||
|
||||
export const memoryWriteTool: ToolDef<z.infer<typeof writeSchema>> = {
|
||||
name: "memory_write",
|
||||
description:
|
||||
"Write or delete a typed personal memory fact (file under the locode config dir's memory/ directory). Each fact " +
|
||||
"is one file with frontmatter (name/description/type) + a body; a one-line index is folded into every future " +
|
||||
"session's system prompt so you can recall it. Use 'write' to save a durable fact worth remembering across " +
|
||||
"sessions (a stated preference, a correction, a project convention) and 'delete' to remove one. Do not use this " +
|
||||
"for transient per-task notes.",
|
||||
schema: writeSchema,
|
||||
mutating: true,
|
||||
preview: async (args) => {
|
||||
if (args.action === "delete") {
|
||||
const existing = await readMemoryEntry(args.name);
|
||||
if (!existing) return `No memory fact named '${args.name}' — nothing to delete.`;
|
||||
return createPatch(`${args.name}.md`, serializeForPreview(existing), "", "", "");
|
||||
}
|
||||
if (!args.description || !args.type || !args.content) {
|
||||
return "write requires 'name', 'description', 'type', and 'content'.";
|
||||
}
|
||||
const existing = await readMemoryEntry(args.name);
|
||||
const next = serializeForPreview({ name: args.name, description: args.description, type: args.type, body: args.content });
|
||||
if (!existing) return `Create memory fact ${args.name}.md:\n${next}`;
|
||||
return createPatch(`${args.name}.md`, serializeForPreview(existing), next, "", "");
|
||||
},
|
||||
handler: async (args, ctx) => {
|
||||
if (args.action === "delete") {
|
||||
const removed = await deleteMemoryEntry(args.name);
|
||||
if (!removed) return { error: `No memory fact named '${args.name}' to delete.` };
|
||||
if (ctx.setUserMemory) ctx.setUserMemory(await loadUserMemory());
|
||||
return { deleted: args.name };
|
||||
}
|
||||
// write
|
||||
if (!args.description || !args.type || !args.content) {
|
||||
return { error: "write requires non-empty 'description', 'type', and 'content'." };
|
||||
}
|
||||
const content = args.content.trim();
|
||||
if (!content) return { error: "'content' must not be empty." };
|
||||
const entry = await writeMemoryEntry({
|
||||
name: args.name,
|
||||
description: args.description.trim(),
|
||||
type: args.type,
|
||||
body: content,
|
||||
});
|
||||
if (ctx.setUserMemory) ctx.setUserMemory(await loadUserMemory());
|
||||
return { written: entry.name, type: entry.type, indexUpdated: true };
|
||||
},
|
||||
};
|
||||
|
||||
function serializeForPreview(entry: { name: string; description: string; type: MemoryType; body: string }): string {
|
||||
return `---\nname: ${entry.name}\ndescription: ${entry.description}\ntype: ${entry.type}\n---\n\n${entry.body.trim()}\n`;
|
||||
}
|
||||
+117
-129
@@ -1,153 +1,141 @@
|
||||
import { mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs";
|
||||
import os from "node:os";
|
||||
import { mkdtemp, readFile, rm, writeFile } from "node:fs/promises";
|
||||
import { tmpdir } from "node:os";
|
||||
import path from "node:path";
|
||||
import { afterEach, beforeEach, describe, expect, it } from "vitest";
|
||||
import { describe, expect, it } from "vitest";
|
||||
import { multiEditTool } from "./multiEdit.js";
|
||||
import type { ToolContext } from "./types.js";
|
||||
|
||||
describe("multiEdit tool", () => {
|
||||
let cwd: string;
|
||||
let ctx: ToolContext;
|
||||
async function makeCwd(): Promise<string> {
|
||||
return mkdtemp(path.join(tmpdir(), "locode-multiedit-"));
|
||||
}
|
||||
|
||||
beforeEach(() => {
|
||||
cwd = mkdtempSync(path.join(os.tmpdir(), "locode-multiedit-"));
|
||||
ctx = { cwd };
|
||||
describe("multiEditTool", () => {
|
||||
it("rejects paths that escape the working directory", async () => {
|
||||
const cwd = await makeCwd();
|
||||
try {
|
||||
await expect(
|
||||
multiEditTool.handler({ path: "../outside.txt", edits: [{ old_string: "a", new_string: "b" }] }, { cwd }),
|
||||
).rejects.toThrow("Path resolves outside the working directory");
|
||||
} finally {
|
||||
await rm(cwd, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
rmSync(cwd, { recursive: true, force: true });
|
||||
});
|
||||
|
||||
it("applies an ordered batch to one file in a single atomic write", async () => {
|
||||
writeFileSync(path.join(cwd, "code.ts"), "const A = 1;\nconst B = 2;\nconst C = 3;\n");
|
||||
const result = (await multiEditTool.handler(
|
||||
{
|
||||
path: "code.ts",
|
||||
edits: [
|
||||
{ old_string: "const A = 1;", new_string: "const A = 10;" },
|
||||
{ old_string: "const C = 3;", new_string: "const C = 30;" },
|
||||
],
|
||||
},
|
||||
ctx,
|
||||
)) as { applied: number };
|
||||
expect(result.applied).toBe(2);
|
||||
expect(readFileSync(path.join(cwd, "code.ts"), "utf-8")).toBe(
|
||||
"const A = 10;\nconst B = 2;\nconst C = 30;\n",
|
||||
);
|
||||
});
|
||||
|
||||
it("an earlier edit can change the text a later edit matches", async () => {
|
||||
writeFileSync(path.join(cwd, "f.txt"), "alpha\n");
|
||||
await multiEditTool.handler(
|
||||
{
|
||||
path: "f.txt",
|
||||
edits: [
|
||||
{ old_string: "alpha", new_string: "beta" },
|
||||
{ old_string: "beta", new_string: "gamma" },
|
||||
],
|
||||
},
|
||||
ctx,
|
||||
);
|
||||
expect(readFileSync(path.join(cwd, "f.txt"), "utf-8")).toBe("gamma\n");
|
||||
});
|
||||
|
||||
it("errors on the first edit that doesn't match, naming the edit index", async () => {
|
||||
writeFileSync(path.join(cwd, "f.txt"), "alpha\n");
|
||||
await expect(
|
||||
multiEditTool.handler(
|
||||
it("applies several edits to one file in order, as a single atomic write", async () => {
|
||||
const cwd = await makeCwd();
|
||||
try {
|
||||
const file = path.join(cwd, "src.txt");
|
||||
await writeFile(file, "alpha\nbeta\ngamma\n");
|
||||
const result = (await multiEditTool.handler(
|
||||
{
|
||||
path: "f.txt",
|
||||
path: "src.txt",
|
||||
edits: [
|
||||
{ old_string: "alpha", new_string: "beta" },
|
||||
{ old_string: "missing", new_string: "x" },
|
||||
{ old_string: "alpha", new_string: "ALPHA" },
|
||||
{ old_string: "beta", new_string: "BETA" },
|
||||
{ old_string: "gamma", new_string: "GAMMA" },
|
||||
],
|
||||
},
|
||||
ctx,
|
||||
),
|
||||
).rejects.toThrow(/Edit 2: old_string not found/);
|
||||
{ cwd },
|
||||
)) as { path: string; applied: number };
|
||||
expect(result.applied).toBe(3);
|
||||
expect(await readFile(file, "utf-8")).toBe("ALPHA\nBETA\nGAMMA\n");
|
||||
} finally {
|
||||
await rm(cwd, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
it("errors when an old_string is ambiguous and replace_all is not set", async () => {
|
||||
writeFileSync(path.join(cwd, "f.txt"), "dup\ndup\n");
|
||||
await expect(
|
||||
multiEditTool.handler(
|
||||
it("applies later edits to the result of earlier ones", async () => {
|
||||
const cwd = await makeCwd();
|
||||
try {
|
||||
const file = path.join(cwd, "src.txt");
|
||||
await writeFile(file, "foo\n");
|
||||
// First edit renames the line; second edit matches the renamed text.
|
||||
await multiEditTool.handler(
|
||||
{
|
||||
path: "f.txt",
|
||||
edits: [{ old_string: "dup", new_string: "x" }],
|
||||
path: "src.txt",
|
||||
edits: [
|
||||
{ old_string: "foo", new_string: "bar" },
|
||||
{ old_string: "bar", new_string: "baz" },
|
||||
],
|
||||
},
|
||||
ctx,
|
||||
),
|
||||
).rejects.toThrow(/Edit 1: old_string appears 2 times/);
|
||||
{ cwd },
|
||||
);
|
||||
expect(await readFile(file, "utf-8")).toBe("baz\n");
|
||||
} finally {
|
||||
await rm(cwd, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
it("replace_all applies to all occurrences within the batch step", async () => {
|
||||
writeFileSync(path.join(cwd, "f.txt"), "dup\ndup\n");
|
||||
await multiEditTool.handler(
|
||||
{
|
||||
path: "f.txt",
|
||||
edits: [{ old_string: "dup", new_string: "x", replace_all: true }],
|
||||
},
|
||||
ctx,
|
||||
);
|
||||
expect(readFileSync(path.join(cwd, "f.txt"), "utf-8")).toBe("x\nx\n");
|
||||
it("fails on the first non-unique match and writes nothing", async () => {
|
||||
const cwd = await makeCwd();
|
||||
try {
|
||||
const file = path.join(cwd, "src.txt");
|
||||
const original = "dup\ndup\nunique\n";
|
||||
await writeFile(file, original);
|
||||
await expect(
|
||||
multiEditTool.handler(
|
||||
{
|
||||
path: "src.txt",
|
||||
edits: [
|
||||
{ old_string: "dup", new_string: "x" }, // ambiguous, no replace_all
|
||||
{ old_string: "unique", new_string: "UNIQUE" },
|
||||
],
|
||||
},
|
||||
{ cwd },
|
||||
),
|
||||
).rejects.toThrow(/appears 2 times/);
|
||||
// The failed batch must not have written anything — the file is unchanged.
|
||||
expect(await readFile(file, "utf-8")).toBe(original);
|
||||
} finally {
|
||||
await rm(cwd, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
it("refuses to edit outside the working directory", async () => {
|
||||
await expect(
|
||||
multiEditTool.handler(
|
||||
{ path: "../escape.txt", edits: [{ old_string: "a", new_string: "b" }] },
|
||||
ctx,
|
||||
),
|
||||
).rejects.toThrow(/outside the working directory/);
|
||||
it("respects replace_all within a batch edit", async () => {
|
||||
const cwd = await makeCwd();
|
||||
try {
|
||||
const file = path.join(cwd, "src.txt");
|
||||
await writeFile(file, "dup\ndup\n");
|
||||
await multiEditTool.handler(
|
||||
{ path: "src.txt", edits: [{ old_string: "dup", new_string: "x", replace_all: true }] },
|
||||
{ cwd },
|
||||
);
|
||||
expect(await readFile(file, "utf-8")).toBe("x\nx\n");
|
||||
} finally {
|
||||
await rm(cwd, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
it("preview produces a unified diff of the full batch", async () => {
|
||||
writeFileSync(path.join(cwd, "f.txt"), "one\ntwo\n");
|
||||
const preview = await multiEditTool.preview!(
|
||||
{
|
||||
path: "f.txt",
|
||||
edits: [
|
||||
{ old_string: "one", new_string: "ONE" },
|
||||
{ old_string: "two", new_string: "TWO" },
|
||||
],
|
||||
},
|
||||
ctx,
|
||||
);
|
||||
expect(preview).toMatch(/-one/);
|
||||
expect(preview).toMatch(/\+ONE/);
|
||||
expect(preview).toMatch(/-two/);
|
||||
expect(preview).toMatch(/\+TWO/);
|
||||
it("produces a diff preview spanning all edits", async () => {
|
||||
const cwd = await makeCwd();
|
||||
try {
|
||||
const file = path.join(cwd, "src.txt");
|
||||
await writeFile(file, "a\nb\n");
|
||||
const preview = await multiEditTool.preview!(
|
||||
{ path: "src.txt", edits: [{ old_string: "a", new_string: "A" }, { old_string: "b", new_string: "B" }] },
|
||||
{ cwd },
|
||||
);
|
||||
expect(preview).toContain("@@");
|
||||
expect(preview).toContain("+A");
|
||||
expect(preview).toContain("+B");
|
||||
} finally {
|
||||
await rm(cwd, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
it("preview warns when a batch edit will fail", async () => {
|
||||
writeFileSync(path.join(cwd, "f.txt"), "one\n");
|
||||
const preview = await multiEditTool.preview!(
|
||||
{
|
||||
path: "f.txt",
|
||||
edits: [{ old_string: "missing", new_string: "x" }],
|
||||
},
|
||||
ctx,
|
||||
);
|
||||
expect(preview).toMatch(/Edit 1: old_string not found/);
|
||||
});
|
||||
|
||||
// Same LF-vs-CRLF mismatch as editFile.test.ts: read_file always shows the model LF content, so
|
||||
// a batch's old_string/new_string must match against a CRLF file's LF-normalized text, and the
|
||||
// result written back must preserve the file's original CRLF convention.
|
||||
it("matches LF old_strings against a CRLF file and preserves CRLF on write", async () => {
|
||||
writeFileSync(path.join(cwd, "code.ts"), "const A = 1;\r\nconst B = 2;\r\nconst C = 3;\r\n");
|
||||
await multiEditTool.handler(
|
||||
{
|
||||
path: "code.ts",
|
||||
edits: [
|
||||
{ old_string: "const A = 1;", new_string: "const A = 10;" },
|
||||
{ old_string: "const C = 3;", new_string: "const C = 30;" },
|
||||
],
|
||||
},
|
||||
ctx,
|
||||
);
|
||||
expect(readFileSync(path.join(cwd, "code.ts"), "utf-8")).toBe(
|
||||
"const A = 10;\r\nconst B = 2;\r\nconst C = 30;\r\n",
|
||||
);
|
||||
it("records the pre-edit content for /undo via setLastEdit", async () => {
|
||||
const cwd = await makeCwd();
|
||||
try {
|
||||
const file = path.join(cwd, "src.txt");
|
||||
const original = "a\nb\n";
|
||||
await writeFile(file, original);
|
||||
let captured: { path: string; previousContent: string } | undefined;
|
||||
await multiEditTool.handler(
|
||||
{ path: "src.txt", edits: [{ old_string: "a", new_string: "A" }] },
|
||||
{ cwd, setLastEdit: (e: { path: string; previousContent: string }) => (captured = e) } as any,
|
||||
);
|
||||
expect(captured).toEqual({ path: "src.txt", previousContent: original });
|
||||
} finally {
|
||||
await rm(cwd, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
});
|
||||
+25
-37
@@ -1,41 +1,32 @@
|
||||
import { createPatch } from "diff";
|
||||
import { randomBytes } from "node:crypto";
|
||||
import { readFile as fsReadFile, rename as fsRename, unlink as fsUnlink, writeFile as fsWriteFile } from "node:fs/promises";
|
||||
import path from "node:path";
|
||||
import { z } from "zod";
|
||||
import { resolveWithinCwd } from "./pathGuard.js";
|
||||
import { applyEdit, countOccurrences, detectEol, fromLF, toLF } from "./editFile.js";
|
||||
import { assertWithinWorkspace } from "../utils/path.js";
|
||||
import { applyEdit, countOccurrences } from "./editFile.js";
|
||||
import type { ToolDef } from "./types.js";
|
||||
|
||||
// A single edit within a multi_edit batch. Mirrors edit_file's args minus `path` (which is shared
|
||||
// across the whole batch). Each edit is applied in array order to the result of the previous one, so
|
||||
// an earlier edit can shift the text a later edit matches — that's why each old_string is checked
|
||||
// against the running result, not the original file.
|
||||
// across the whole batch). Each edit is applied in array order to the result of the previous one.
|
||||
const editSchema = z.object({
|
||||
old_string: z.string().describe("Exact text to replace. Must match the current file content exactly at this point in the batch — earlier edits may have shifted it."),
|
||||
old_string: z.string().describe("Exact text to replace. Must match the current file content exactly at this point in the batch."),
|
||||
new_string: z.string().describe("Replacement text."),
|
||||
replace_all: z.boolean().optional().describe("Replace every occurrence instead of requiring a unique match."),
|
||||
});
|
||||
|
||||
const schema = z.object({
|
||||
path: z.string().describe("File path to edit, relative to the working directory or absolute."),
|
||||
edits: z.array(editSchema).min(1).describe("Ordered list of edits to apply to the same file, one after another. Each edit sees the result of the previous one."),
|
||||
edits: z.array(editSchema).min(1).describe("Ordered list of edits to apply to the same file, one after another."),
|
||||
});
|
||||
|
||||
/** Applies a batch of edits to an in-memory string, validating each. Throws on the first edit that
|
||||
* doesn't match uniquely (unless its replace_all is set) or doesn't match at all. Edits apply to the
|
||||
* running result, so an earlier edit can change the text a later edit matches. `original` and every
|
||||
* edit's old_string/new_string must already be LF-normalized (see editFile.ts's detectEol/toLF —
|
||||
* read_file shows the model LF-only content regardless of the file's real line endings, so matching
|
||||
* must happen in that same space). */
|
||||
function applyBatch(
|
||||
original: string,
|
||||
edits: { old_string: string; new_string: string; replace_all?: boolean }[],
|
||||
filePath: string,
|
||||
): string {
|
||||
* running result, so an earlier edit can change the text a later edit matches. */
|
||||
function applyBatch(original: string, edits: { old_string: string; new_string: string; replace_all?: boolean }[], filePath: string): string {
|
||||
let current = original;
|
||||
edits.forEach((edit, i) => {
|
||||
const oldLF = toLF(edit.old_string);
|
||||
const occurrences = countOccurrences(current, oldLF);
|
||||
const occurrences = countOccurrences(current, edit.old_string);
|
||||
if (occurrences === 0) {
|
||||
throw new Error(
|
||||
`Edit ${i + 1}: old_string not found in ${filePath}. Earlier edits may have shifted the text — re-read the file and adjust. Make sure it matches exactly, including whitespace.`,
|
||||
@@ -46,7 +37,7 @@ function applyBatch(
|
||||
`Edit ${i + 1}: old_string appears ${occurrences} times in ${filePath}. Provide more surrounding context to make it unique, or set replace_all: true.`,
|
||||
);
|
||||
}
|
||||
current = applyEdit(current, oldLF, toLF(edit.new_string), edit.replace_all);
|
||||
current = applyEdit(current, edit.old_string, edit.new_string, edit.replace_all);
|
||||
});
|
||||
return current;
|
||||
}
|
||||
@@ -56,39 +47,33 @@ export const multiEditTool: ToolDef<z.infer<typeof schema>> = {
|
||||
description:
|
||||
"Apply several edits to the same file in one call, in order. Each edit is {old_string, new_string, replace_all?}. " +
|
||||
"Use this instead of repeated edit_file calls when you have multiple distinct changes to one file — it's one confirmation " +
|
||||
"and one atomic write instead of N round-trips. Each old_string must match uniquely at its point in the batch " +
|
||||
"(unless replace_all is set). Read the file first.",
|
||||
"and one atomic write. Each old_string must match uniquely at its point in the batch (unless replace_all is set). Read the file first.",
|
||||
schema,
|
||||
mutating: true,
|
||||
preview: async ({ path: filePath, edits }, ctx) => {
|
||||
let resolved: string;
|
||||
const resolved = path.resolve(ctx.cwd, filePath);
|
||||
assertWithinWorkspace(resolved, ctx.cwd, filePath);
|
||||
let original: string;
|
||||
try {
|
||||
resolved = resolveWithinCwd(ctx.cwd, filePath);
|
||||
} catch (err) {
|
||||
return (err as Error).message;
|
||||
}
|
||||
let raw: string;
|
||||
try {
|
||||
raw = await fsReadFile(resolved, "utf-8");
|
||||
original = await fsReadFile(resolved, "utf-8");
|
||||
} catch {
|
||||
return `File ${resolved} does not exist.`;
|
||||
}
|
||||
try {
|
||||
const eol = detectEol(raw);
|
||||
const updated = fromLF(applyBatch(toLF(raw), edits, filePath), eol);
|
||||
return createPatch(resolved, raw, updated, "", "");
|
||||
const updated = applyBatch(original, edits, filePath);
|
||||
return createPatch(resolved, original, updated, "", "");
|
||||
} catch (err) {
|
||||
return `Warning: ${(err as Error).message} — this edit will fail.`;
|
||||
}
|
||||
},
|
||||
handler: async ({ path: filePath, edits }, ctx) => {
|
||||
const resolved = resolveWithinCwd(ctx.cwd, filePath);
|
||||
const raw = await fsReadFile(resolved, "utf-8");
|
||||
const eol = detectEol(raw);
|
||||
const updated = fromLF(applyBatch(toLF(raw), edits, filePath), eol);
|
||||
const resolved = path.resolve(ctx.cwd, filePath);
|
||||
assertWithinWorkspace(resolved, ctx.cwd, filePath);
|
||||
const original = await fsReadFile(resolved, "utf-8");
|
||||
const updated = applyBatch(original, edits, filePath);
|
||||
// Atomic write via temp+rename (same rationale as edit_file): a crash mid-write can't leave the
|
||||
// user's source file half-overwritten — the live file stays intact until the rename swaps in the
|
||||
// full new content. Clean up the temp file if anything fails so a stray `.tmp` doesn't accumulate.
|
||||
// full new content. Clean up the temp file if anything fails.
|
||||
const tmp = `${resolved}.locode-${randomBytes(4).toString("hex")}.tmp`;
|
||||
try {
|
||||
await fsWriteFile(tmp, updated, "utf-8");
|
||||
@@ -97,6 +82,9 @@ export const multiEditTool: ToolDef<z.infer<typeof schema>> = {
|
||||
await fsUnlink(tmp).catch(() => {});
|
||||
throw err;
|
||||
}
|
||||
if (ctx.setLastEdit) {
|
||||
ctx.setLastEdit({ path: filePath, previousContent: original });
|
||||
}
|
||||
return { path: resolved, applied: edits.length };
|
||||
},
|
||||
};
|
||||
@@ -6,21 +6,19 @@ import { notebookEditTool } from "./notebookEdit.js";
|
||||
|
||||
/** A minimal valid nbformat 4 notebook with two code cells. */
|
||||
function minimalNotebook(): string {
|
||||
return (
|
||||
JSON.stringify(
|
||||
{
|
||||
nbformat: 4,
|
||||
nbformat_minor: 5,
|
||||
metadata: {},
|
||||
cells: [
|
||||
{ cell_type: "code", id: "c1", source: ["print('a')\n"], metadata: {}, outputs: [], execution_count: null },
|
||||
{ cell_type: "code", id: "c2", source: ["print('b')\n"], metadata: {}, outputs: [], execution_count: null },
|
||||
],
|
||||
},
|
||||
null,
|
||||
2,
|
||||
) + "\n"
|
||||
);
|
||||
return JSON.stringify(
|
||||
{
|
||||
nbformat: 4,
|
||||
nbformat_minor: 5,
|
||||
metadata: {},
|
||||
cells: [
|
||||
{ cell_type: "code", id: "c1", source: ["print('a')\n"], metadata: {}, outputs: [], execution_count: null },
|
||||
{ cell_type: "code", id: "c2", source: ["print('b')\n"], metadata: {}, outputs: [], execution_count: null },
|
||||
],
|
||||
},
|
||||
null,
|
||||
2,
|
||||
) + "\n";
|
||||
}
|
||||
|
||||
async function makeCwd(): Promise<string> {
|
||||
@@ -36,8 +34,11 @@ describe("notebookEditTool", () => {
|
||||
const cwd = await makeCwd();
|
||||
try {
|
||||
await expect(
|
||||
notebookEditTool.handler({ notebook_path: "../outside.ipynb", edit_mode: "delete" }, { cwd }),
|
||||
).rejects.toThrow(/outside the working directory/);
|
||||
notebookEditTool.handler(
|
||||
{ notebook_path: "../outside.ipynb", edit_mode: "delete" },
|
||||
{ cwd },
|
||||
),
|
||||
).rejects.toThrow("Path resolves outside the working directory");
|
||||
} finally {
|
||||
await rm(cwd, { recursive: true, force: true });
|
||||
}
|
||||
@@ -114,20 +115,22 @@ describe("notebookEditTool", () => {
|
||||
}
|
||||
});
|
||||
|
||||
it("deletes a cell by cell_id, and a missing id fails", async () => {
|
||||
it("deletes a cell by cell_id and writes nothing if not found", async () => {
|
||||
const cwd = await makeCwd();
|
||||
try {
|
||||
const file = path.join(cwd, "nb.ipynb");
|
||||
await writeFile(file, minimalNotebook());
|
||||
const original = minimalNotebook();
|
||||
await writeFile(file, original);
|
||||
await notebookEditTool.handler({ notebook_path: "nb.ipynb", cell_id: "c1", edit_mode: "delete" }, { cwd });
|
||||
const cells = parseCells(await readFile(file, "utf-8"));
|
||||
expect(cells).toHaveLength(1);
|
||||
expect(cells[0]!.id).toBe("c2");
|
||||
|
||||
// A missing id fails.
|
||||
// A missing id fails and leaves the file unchanged.
|
||||
await expect(
|
||||
notebookEditTool.handler({ notebook_path: "nb.ipynb", cell_id: "nope", edit_mode: "delete" }, { cwd }),
|
||||
).rejects.toThrow(/not found/);
|
||||
expect(await readFile(file, "utf-8")).not.toBe(original);
|
||||
} finally {
|
||||
await rm(cwd, { recursive: true, force: true });
|
||||
}
|
||||
@@ -159,6 +162,8 @@ describe("notebookEditTool", () => {
|
||||
{ cwd },
|
||||
);
|
||||
expect(preview).toContain("@@");
|
||||
// The diff is over the notebook JSON, so the changed source appears JSON-quoted/indented
|
||||
// on +/- lines rather than as bare text — assert on the content, not the leading marker.
|
||||
expect(preview).toContain("print('A')");
|
||||
expect(preview).toContain("print('a')");
|
||||
} finally {
|
||||
@@ -166,19 +171,18 @@ describe("notebookEditTool", () => {
|
||||
}
|
||||
});
|
||||
|
||||
it("switching a code cell to markdown drops code-only fields", async () => {
|
||||
it("records pre-edit content for /undo via setLastEdit", async () => {
|
||||
const cwd = await makeCwd();
|
||||
try {
|
||||
const file = path.join(cwd, "nb.ipynb");
|
||||
await writeFile(file, minimalNotebook());
|
||||
const original = minimalNotebook();
|
||||
await writeFile(file, original);
|
||||
let captured: { path: string; previousContent: string } | undefined;
|
||||
await notebookEditTool.handler(
|
||||
{ notebook_path: "nb.ipynb", cell_index: 0, edit_mode: "replace", cell_type: "markdown", new_source: "prose" },
|
||||
{ cwd },
|
||||
{ notebook_path: "nb.ipynb", cell_index: 1, edit_mode: "delete" },
|
||||
{ cwd, setLastEdit: (e: { path: string; previousContent: string }) => (captured = e) } as any,
|
||||
);
|
||||
const cell = parseCells(await readFile(file, "utf-8"))[0]!;
|
||||
expect(cell.cell_type).toBe("markdown");
|
||||
expect(cell).not.toHaveProperty("execution_count");
|
||||
expect(cell).not.toHaveProperty("outputs");
|
||||
expect(captured).toEqual({ path: "nb.ipynb", previousContent: original });
|
||||
} finally {
|
||||
await rm(cwd, { recursive: true, force: true });
|
||||
}
|
||||
|
||||
+11
-12
@@ -1,15 +1,15 @@
|
||||
import { createPatch } from "diff";
|
||||
import { randomBytes } from "node:crypto";
|
||||
import { readFile as fsReadFile, rename as fsRename, unlink as fsUnlink, writeFile as fsWriteFile } from "node:fs/promises";
|
||||
import path from "node:path";
|
||||
import { z } from "zod";
|
||||
import { resolveWithinCwd } from "./pathGuard.js";
|
||||
import { assertWithinWorkspace } from "../utils/path.js";
|
||||
import type { ToolDef } from "./types.js";
|
||||
|
||||
// .ipynb is a JSON document (nbformat 4): { nbformat, nbformat_minor, metadata, cells: Cell[] }.
|
||||
// Each cell is { cell_type: "code"|"markdown"|"raw", id?, source, metadata, outputs?, execution_count? }.
|
||||
// `source` is a list of strings where every line except the last carries a trailing "\n" (nbformat
|
||||
// convention). We convert the model's single-string new_source to/from that array form, so the
|
||||
// model never has to hand-write nbformat's line-array quirk — it just gives the full cell text.
|
||||
// convention). We convert the model's single-string new_source to/from that array form.
|
||||
|
||||
type Notebook = { nbformat: number; nbformat_minor: number; metadata: Record<string, unknown>; cells: Cell[] };
|
||||
type Cell = { cell_type: string; id?: string; source: string[]; metadata: Record<string, unknown>; outputs?: unknown[]; execution_count?: unknown };
|
||||
@@ -108,12 +108,8 @@ export const notebookEditTool: ToolDef<z.infer<typeof schema>> = {
|
||||
schema,
|
||||
mutating: true,
|
||||
preview: async (args, ctx) => {
|
||||
let resolved: string;
|
||||
try {
|
||||
resolved = resolveWithinCwd(ctx.cwd, args.notebook_path);
|
||||
} catch (err) {
|
||||
return (err as Error).message;
|
||||
}
|
||||
const resolved = path.resolve(ctx.cwd, args.notebook_path);
|
||||
assertWithinWorkspace(resolved, ctx.cwd, args.notebook_path);
|
||||
let original: string;
|
||||
try {
|
||||
original = await fsReadFile(resolved, "utf-8");
|
||||
@@ -134,7 +130,8 @@ export const notebookEditTool: ToolDef<z.infer<typeof schema>> = {
|
||||
}
|
||||
},
|
||||
handler: async (args, ctx) => {
|
||||
const resolved = resolveWithinCwd(ctx.cwd, args.notebook_path);
|
||||
const resolved = path.resolve(ctx.cwd, args.notebook_path);
|
||||
assertWithinWorkspace(resolved, ctx.cwd, args.notebook_path);
|
||||
const original = await fsReadFile(resolved, "utf-8");
|
||||
let notebook: Notebook;
|
||||
try {
|
||||
@@ -145,8 +142,7 @@ export const notebookEditTool: ToolDef<z.infer<typeof schema>> = {
|
||||
if (!Array.isArray(notebook.cells)) throw new Error(`${resolved} has no cells array — not a valid .ipynb.`);
|
||||
applyNotebookEdit(notebook, args, args.notebook_path);
|
||||
const updated = JSON.stringify(notebook, null, 2) + "\n";
|
||||
// Atomic write via temp+rename (same rationale as edit_file/multi_edit): a crash mid-write can't
|
||||
// leave the notebook half-overwritten. Clean up the temp file if anything fails.
|
||||
// Atomic write via temp+rename (same rationale as edit_file/multi_edit).
|
||||
const tmp = `${resolved}.locode-${randomBytes(4).toString("hex")}.tmp`;
|
||||
try {
|
||||
await fsWriteFile(tmp, updated, "utf-8");
|
||||
@@ -155,6 +151,9 @@ export const notebookEditTool: ToolDef<z.infer<typeof schema>> = {
|
||||
await fsUnlink(tmp).catch(() => {});
|
||||
throw err;
|
||||
}
|
||||
if (ctx.setLastEdit) {
|
||||
ctx.setLastEdit({ path: args.notebook_path, previousContent: original });
|
||||
}
|
||||
return { path: resolved, edit_mode: args.edit_mode ?? "replace", cell_count: notebook.cells.length };
|
||||
},
|
||||
};
|
||||
@@ -1,36 +0,0 @@
|
||||
import path from "node:path";
|
||||
import { describe, expect, it } from "vitest";
|
||||
import { PathOutsideCwdError, resolveWithinCwd } from "./pathGuard.js";
|
||||
|
||||
describe("resolveWithinCwd", () => {
|
||||
const cwd = path.resolve("/project");
|
||||
|
||||
it("resolves a plain relative path inside cwd", () => {
|
||||
expect(resolveWithinCwd(cwd, "src/index.ts")).toBe(path.join(cwd, "src", "index.ts"));
|
||||
});
|
||||
|
||||
it("resolves an absolute path that happens to already be inside cwd", () => {
|
||||
const inside = path.join(cwd, "foo.txt");
|
||||
expect(resolveWithinCwd(cwd, inside)).toBe(inside);
|
||||
});
|
||||
|
||||
it("resolves cwd itself", () => {
|
||||
expect(resolveWithinCwd(cwd, ".")).toBe(cwd);
|
||||
});
|
||||
|
||||
it("rejects a ../ escape", () => {
|
||||
expect(() => resolveWithinCwd(cwd, "../outside.txt")).toThrow(PathOutsideCwdError);
|
||||
});
|
||||
|
||||
it("rejects a deeper ../../ escape", () => {
|
||||
expect(() => resolveWithinCwd(cwd, "sub/../../outside.txt")).toThrow(PathOutsideCwdError);
|
||||
});
|
||||
|
||||
it("rejects an absolute path outside cwd", () => {
|
||||
expect(() => resolveWithinCwd(cwd, path.resolve("/etc/passwd"))).toThrow(PathOutsideCwdError);
|
||||
});
|
||||
|
||||
it("rejects the filesystem root", () => {
|
||||
expect(() => resolveWithinCwd(cwd, path.parse(cwd).root)).toThrow(PathOutsideCwdError);
|
||||
});
|
||||
});
|
||||
@@ -1,27 +0,0 @@
|
||||
import path from "node:path";
|
||||
|
||||
/** Thrown by resolveWithinCwd — kept as its own class only so callers can recognize it (via
|
||||
* instanceof) if they ever need to react differently than a plain thrown Error. */
|
||||
export class PathOutsideCwdError extends Error {}
|
||||
|
||||
/** Resolves `targetPath` against `cwd` and hard-blocks the result if it would land outside the
|
||||
* project root (the working directory locode was launched in) — an absolute path elsewhere on
|
||||
* disk, a `../` escape, or (on Windows) a path on a different drive all reject. This applies
|
||||
* unconditionally, regardless of permission mode: even `auto-accept` skips a tool's `preview`
|
||||
* entirely (see gateAndRun in agent/loop.ts), so this check has to live in each tool's `handler`
|
||||
* — which always runs — to actually hold as a floor rather than just a confirmation-dialog hint.
|
||||
* It's deliberately not configurable; a model tricked (or simply mistaken) into targeting a path
|
||||
* outside the project shouldn't be one auto-approved call away from touching it. */
|
||||
export function resolveWithinCwd(cwd: string, targetPath: string): string {
|
||||
const resolved = path.resolve(cwd, targetPath);
|
||||
const rel = path.relative(cwd, resolved);
|
||||
// rel === "" is targetPath resolving to cwd itself — fine. Anything starting with ".." walked
|
||||
// upward out of cwd; an absolute rel (Windows: a different drive, e.g. "D:\foo") never went
|
||||
// through cwd's tree in the first place. Either way, it's outside.
|
||||
if (rel !== "" && (rel.startsWith(`..${path.sep}`) || rel === ".." || path.isAbsolute(rel))) {
|
||||
throw new PathOutsideCwdError(
|
||||
`Refusing to write outside the working directory: "${targetPath}" resolves to ${resolved}, which is not inside ${cwd}.`,
|
||||
);
|
||||
}
|
||||
return resolved;
|
||||
}
|
||||
+6
-47
@@ -2,25 +2,9 @@ import { readFile as fsReadFile, stat as fsStat } from "node:fs/promises";
|
||||
import path from "node:path";
|
||||
import { z } from "zod";
|
||||
import { imageMimeType, MAX_IMAGE_BYTES } from "../utils/image.js";
|
||||
import { assertWithinWorkspace } from "../utils/path.js";
|
||||
import type { ToolDef } from "./types.js";
|
||||
|
||||
function withinWorkspace(resolved: string, workspace: string): boolean {
|
||||
const rel = path.relative(workspace, resolved);
|
||||
return !rel.startsWith("..") && !path.isAbsolute(rel);
|
||||
}
|
||||
|
||||
// Reject files larger than this so we never accidentally OOM on a huge binary or log file.
|
||||
const MAX_FILE_SIZE = 10 * 1024 * 1024; // 10 MB
|
||||
|
||||
// Common binary extensions — if the file extension matches, reject without reading.
|
||||
const BINARY_EXTENSIONS = new Set([
|
||||
".exe", ".dll", ".so", ".dylib", ".bin", ".dat", ".o", ".obj", ".pyc", ".pyo",
|
||||
".class", ".jar", ".war", ".zip", ".tar", ".gz", ".bz2", ".7z", ".rar",
|
||||
".iso", ".dmg", ".pdb", ".lib", ".a", ".woff", ".woff2", ".eot", ".ttf", ".otf",
|
||||
".pdf", ".doc", ".docx", ".xls", ".xlsx", ".ppt", ".pptx",
|
||||
".sqlite", ".db", ".ico", ".cur",
|
||||
]);
|
||||
|
||||
// Cap on how much text a single read_file call returns, so a huge file can't blow up the context
|
||||
// in one call. Cut on a line boundary (never mid-line) and report the exact next offset, so the
|
||||
// model can page through the rest with `offset` instead of re-reading the same truncated prefix in
|
||||
@@ -39,25 +23,17 @@ export const readFileTool: ToolDef<z.infer<typeof schema>> = {
|
||||
description:
|
||||
"Read a local file. Text files return 1-indexed lines; large files are paginated (use nextOffset for next page). " +
|
||||
"Image files (png, jpg, jpeg, gif, webp, bmp) are returned as image content (requires vision-capable model). " +
|
||||
"Use to inspect file contents before editing, or to understand existing code. Prefer this over bash cat for files.",
|
||||
"When you need to inspect many files, use grep first to find the relevant ones and only read_file the files or page ranges you actually need — " +
|
||||
"the session has a per-turn tool-call budget, and unnecessary full-file reads burn through it quickly.",
|
||||
schema,
|
||||
mutating: false,
|
||||
handler: async ({ path: filePath, offset, limit }, ctx) => {
|
||||
const resolved = path.resolve(ctx.cwd, filePath);
|
||||
if (!withinWorkspace(resolved, ctx.cwd)) {
|
||||
throw new Error(`File ${filePath} resolves outside the workspace.`);
|
||||
}
|
||||
|
||||
// --- Size guard: reject files over MAX_FILE_SIZE before reading ---
|
||||
const stats = await fsStat(resolved);
|
||||
if (stats.size > MAX_FILE_SIZE) {
|
||||
throw new Error(
|
||||
`${filePath} is ${(stats.size / 1_048_576).toFixed(1)}MB, over the ${MAX_FILE_SIZE / 1_048_576}MB read limit.`,
|
||||
);
|
||||
}
|
||||
assertWithinWorkspace(resolved, ctx.cwd, filePath);
|
||||
|
||||
const mimeType = imageMimeType(resolved);
|
||||
if (mimeType) {
|
||||
const stats = await fsStat(resolved);
|
||||
if (stats.size > MAX_IMAGE_BYTES) {
|
||||
throw new Error(
|
||||
`${filePath} is ${(stats.size / 1_048_576).toFixed(1)}MB, over the ${MAX_IMAGE_BYTES / 1_048_576}MB limit for image reads.`,
|
||||
@@ -67,25 +43,8 @@ export const readFileTool: ToolDef<z.infer<typeof schema>> = {
|
||||
return { path: resolved, image: true, mimeType, bytes: buffer.byteLength, base64: buffer.toString("base64") };
|
||||
}
|
||||
|
||||
// --- Binary guard: reject by extension ---
|
||||
const ext = path.extname(resolved).toLowerCase();
|
||||
if (BINARY_EXTENSIONS.has(ext)) {
|
||||
throw new Error(
|
||||
`${filePath} looks like a binary file (${ext}). Use bash for binary inspection.`,
|
||||
);
|
||||
}
|
||||
|
||||
const content = await fsReadFile(resolved, "utf-8");
|
||||
|
||||
// --- Binary guard: null-byte heuristic (catches extensionless binaries) ---
|
||||
const nullIndex = content.indexOf("\0");
|
||||
if (nullIndex !== -1) {
|
||||
throw new Error(
|
||||
`${filePath} appears to be a binary file (null byte at position ${nullIndex}). Use bash for binary inspection.`,
|
||||
);
|
||||
}
|
||||
|
||||
const lines = content.split(/\r?\n/);
|
||||
const lines = content.split("\n");
|
||||
const start = offset ? offset - 1 : 0;
|
||||
const requestedEnd = limit ? Math.min(start + limit, lines.length) : lines.length;
|
||||
|
||||
|
||||
@@ -0,0 +1,72 @@
|
||||
import { z } from "zod";
|
||||
import type { ToolDef } from "./types.js";
|
||||
|
||||
const schema = z
|
||||
.object({
|
||||
agentId: z
|
||||
.string()
|
||||
.optional()
|
||||
.describe(
|
||||
"The agentId returned by a prior 'agent'/'agent__*' call. Only resumable agents (single, " +
|
||||
"shared-cwd delegations) return an agentId — parallel (worktree-isolated) agents do not. " +
|
||||
"Provide this OR `name`.",
|
||||
),
|
||||
name: z
|
||||
.string()
|
||||
.optional()
|
||||
.describe(
|
||||
"The name of a teammate you created by passing `name` to a prior `agent` call. Addressing a " +
|
||||
"teammate by name avoids tracking its agentId. Provide this OR `agentId`. Use list_teammates " +
|
||||
"to see your named teammates.",
|
||||
),
|
||||
message: z
|
||||
.string()
|
||||
.describe("The follow-up instruction for the sub-agent. It retains the context of the original delegation."),
|
||||
})
|
||||
.refine((d) => d.agentId || d.name, {
|
||||
message: "Provide either an `agentId` (returned by a prior agent call) or a `name` (of a named teammate).",
|
||||
});
|
||||
|
||||
/** Continues a previously-spawned resumable sub-agent with a follow-up message, preserving its
|
||||
* context — the cheaper alternative to re-delegating from scratch when a sub-agent's first answer
|
||||
* was close but needs a correction, or when it ran out of budget mid-task. Only sub-agents that ran
|
||||
* in the shared cwd (single/sequential delegations) are resumable and return an agentId; parallel
|
||||
* worktree-isolated agents are fire-and-forget. The target may be identified by its agentId OR, if
|
||||
* it was spawned with a `name`, by that name. Read-only: it has no filesystem side effects beyond
|
||||
* what the continued sub-agent itself does (and those still go through the normal confirm gate). */
|
||||
export const sendMessageTool: ToolDef<z.infer<typeof schema>> = {
|
||||
name: "send_message",
|
||||
description:
|
||||
"Continue a previously-spawned resumable sub-agent (one that returned an agentId, or one you named via " +
|
||||
"the `agent` tool's `name` arg) with a follow-up message, preserving its context. Cheaper than " +
|
||||
"re-delegating from scratch. Use it to refine a sub-agent's answer, ask a follow-up, or continue one that " +
|
||||
"ran out of its step budget. Identify it by `agentId` OR by `name` (a teammate name). Only shared-cwd agents " +
|
||||
"are resumable; parallel worktree-isolated agents don't expose an agentId.",
|
||||
schema,
|
||||
mutating: false,
|
||||
handler: async (args, ctx) => {
|
||||
if (!ctx.resumeSubAgent) {
|
||||
throw new Error("Sub-agent continuation is not available in this context.");
|
||||
}
|
||||
// Resolve the target: a name takes precedence (it's the human-friendly handle), but fall back to
|
||||
// an explicit agentId. If a name is given but not on the roster, return a clean error instead of
|
||||
// calling resumeSubAgent with an undefined agentId.
|
||||
let agentId = args.agentId;
|
||||
if (args.name) {
|
||||
const resolved = ctx.resolveTeammate?.(args.name);
|
||||
if (!resolved) {
|
||||
return {
|
||||
error: `No teammate named "${args.name}" was found. Use list_teammates to see named teammates, or pass the agentId returned by the original agent call.`,
|
||||
};
|
||||
}
|
||||
agentId = resolved;
|
||||
}
|
||||
if (!agentId) {
|
||||
return {
|
||||
error: "Provide either a `name` (of a named teammate) or an `agentId` (returned by a prior agent call) to identify the sub-agent to continue.",
|
||||
};
|
||||
}
|
||||
const result = await ctx.resumeSubAgent(agentId, args.message);
|
||||
return { agentId, result, ...(args.name ? { name: args.name } : {}) };
|
||||
},
|
||||
};
|
||||
+196
-92
@@ -1,116 +1,220 @@
|
||||
import { describe, it, expect } from "vitest";
|
||||
import { taskCreateTool, taskListTool, taskGetTool, taskUpdateTool, TaskStore } from "./task.js";
|
||||
import { describe, expect, it, vi } from "vitest";
|
||||
import { TaskStore, taskCreateTool, taskGetTool, taskListTool, taskUpdateTool } from "./task.js";
|
||||
import type { Task, TaskSummary } from "./task.js";
|
||||
import type { ToolContext } from "./types.js";
|
||||
|
||||
function ctxWithStore(): { ctx: ToolContext; store: TaskStore } {
|
||||
const store = new TaskStore();
|
||||
return { ctx: { taskStore: store } as ToolContext, store };
|
||||
function ctxWith(store?: TaskStore): ToolContext {
|
||||
return { cwd: "/x", ...(store ? { taskStore: store } : {}) };
|
||||
}
|
||||
|
||||
describe("task tools", () => {
|
||||
it("task_create creates a pending task and returns it with an id", async () => {
|
||||
const { ctx, store } = ctxWithStore();
|
||||
const out = (await taskCreateTool.handler({ subject: "Fix bug", description: "details" }, ctx)) as {
|
||||
id: string;
|
||||
task: { status: string; blocks: string[]; blockedBy: string[] };
|
||||
describe("TaskStore", () => {
|
||||
it("creates tasks with sequential ids starting at t1 and pending status", () => {
|
||||
const s = new TaskStore();
|
||||
const a = s.create({ subject: "A", description: "do A" });
|
||||
const b = s.create({ subject: "B", description: "do B", activeForm: "doing B" });
|
||||
expect(a.id).toBe("t1");
|
||||
expect(b.id).toBe("t2");
|
||||
expect(a.status).toBe("pending");
|
||||
expect(b.activeForm).toBe("doing B");
|
||||
expect(a.blocks).toEqual([]);
|
||||
expect(a.blockedBy).toEqual([]);
|
||||
});
|
||||
|
||||
it("emits a snapshot via the change emitter after each mutation", () => {
|
||||
const s = new TaskStore();
|
||||
const snaps: ReturnType<TaskStore["list"]>[] = [];
|
||||
s.setEmitter((tasks) => snaps.push(tasks));
|
||||
s.create({ subject: "A", description: "x" });
|
||||
s.create({ subject: "B", description: "y" });
|
||||
expect(snaps).toHaveLength(2);
|
||||
expect(snaps[1]!.map((t) => t.id)).toEqual(["t1", "t2"]);
|
||||
});
|
||||
|
||||
it("updates status, subject, owner, and activeForm", () => {
|
||||
const s = new TaskStore();
|
||||
const t = s.create({ subject: "A", description: "x" });
|
||||
const u = s.update(t.id, { status: "in_progress", owner: "agent-1", activeForm: "working" });
|
||||
expect(u?.status).toBe("in_progress");
|
||||
expect(u?.owner).toBe("agent-1");
|
||||
expect(u?.activeForm).toBe("working");
|
||||
});
|
||||
|
||||
it("links dependencies via addBlocks/addBlockedBy, ignoring self-refs, unknown ids, and duplicates", () => {
|
||||
const s = new TaskStore();
|
||||
const a = s.create({ subject: "A", description: "x" });
|
||||
const b = s.create({ subject: "B", description: "y" });
|
||||
// A blocks B: add B's blockedBy=[A] and A's blocks=[B].
|
||||
s.update(b.id, { addBlockedBy: [a.id] });
|
||||
s.update(a.id, { addBlocks: [b.id] });
|
||||
expect(s.get(a.id)!.blocks).toEqual([b.id]);
|
||||
expect(s.get(b.id)!.blockedBy).toEqual([a.id]);
|
||||
// Self-ref, unknown id, and duplicate are all ignored.
|
||||
s.update(a.id, { addBlocks: [a.id, "t99", b.id] });
|
||||
expect(s.get(a.id)!.blocks).toEqual([b.id]);
|
||||
});
|
||||
|
||||
it("prevents a direct 2-cycle when adding a blockedBy dependency", () => {
|
||||
const s = new TaskStore();
|
||||
const a = s.create({ subject: "A", description: "x" });
|
||||
const b = s.create({ subject: "B", description: "y" });
|
||||
s.update(a.id, { addBlockedBy: [b.id] }); // A waits on B
|
||||
// Now B waiting on A would create a 2-cycle — skipped silently.
|
||||
s.update(b.id, { addBlockedBy: [a.id] });
|
||||
expect(s.get(b.id)!.blockedBy).toEqual([]);
|
||||
});
|
||||
|
||||
it("merge-patches metadata, deleting keys set to null", () => {
|
||||
const s = new TaskStore();
|
||||
const t = s.create({ subject: "A", description: "x", metadata: { keep: 1, drop: 2 } });
|
||||
s.update(t.id, { metadata: { added: 3, drop: null } });
|
||||
expect(s.get(t.id)!.metadata).toEqual({ keep: 1, added: 3 });
|
||||
});
|
||||
|
||||
it("deletes a task (status: 'deleted') and prunes dangling block/blockedBy refs", () => {
|
||||
const s = new TaskStore();
|
||||
const a = s.create({ subject: "A", description: "x" });
|
||||
const b = s.create({ subject: "B", description: "y" });
|
||||
s.update(b.id, { addBlockedBy: [a.id] });
|
||||
s.update(a.id, { addBlocks: [b.id] });
|
||||
expect(s.update(a.id, { status: "deleted" })).toBeUndefined();
|
||||
expect(s.get(a.id)).toBeUndefined();
|
||||
// B's blockedBy no longer references the deleted A.
|
||||
expect(s.get(b.id)!.blockedBy).toEqual([]);
|
||||
expect(s.get(b.id)!.blocks).toEqual([]);
|
||||
});
|
||||
|
||||
it("update returns undefined for an unknown id", () => {
|
||||
const s = new TaskStore();
|
||||
expect(s.update("t99", { status: "in_progress" })).toBeUndefined();
|
||||
});
|
||||
});
|
||||
|
||||
describe("task_create tool", () => {
|
||||
it("is non-mutating (no confirmation prompt)", () => {
|
||||
expect(taskCreateTool.mutating).toBe(false);
|
||||
});
|
||||
|
||||
it("creates a task and returns its id + snapshot", async () => {
|
||||
const store = new TaskStore();
|
||||
const result = (await taskCreateTool.handler(
|
||||
{ subject: "Fix bug", description: "root-cause then patch", activeForm: "Fixing bug" },
|
||||
ctxWith(store),
|
||||
)) as { id: string; task: Task };
|
||||
expect(result.id).toBe("t1");
|
||||
expect(result.task.subject).toBe("Fix bug");
|
||||
expect(result.task.activeForm).toBe("Fixing bug");
|
||||
// Mutating the returned snapshot must not affect the store.
|
||||
result.task.blocks.push("t99");
|
||||
expect(store.get("t1")!.blocks).toEqual([]);
|
||||
});
|
||||
|
||||
it("returns a clear error when no task store is available", async () => {
|
||||
const result = (await taskCreateTool.handler({ subject: "x", description: "y" }, ctxWith(undefined))) as {
|
||||
error: string;
|
||||
};
|
||||
expect(out.id).toMatch(/^t\d+$/);
|
||||
expect(out.task.status).toBe("pending");
|
||||
expect(out.task.blocks).toEqual([]);
|
||||
expect(out.task.blockedBy).toEqual([]);
|
||||
expect(store.list()).toHaveLength(1);
|
||||
expect(result.error).toMatch(/not available/i);
|
||||
});
|
||||
});
|
||||
|
||||
describe("task_list tool", () => {
|
||||
it("returns summaries of all tasks", async () => {
|
||||
const store = new TaskStore();
|
||||
store.create({ subject: "A", description: "x" });
|
||||
store.create({ subject: "B", description: "y" });
|
||||
const result = (await taskListTool.handler({}, ctxWith(store))) as { tasks: TaskSummary[] };
|
||||
expect(result.tasks).toHaveLength(2);
|
||||
expect(result.tasks.map((t) => t.subject)).toEqual(["A", "B"]);
|
||||
});
|
||||
|
||||
it("task_list returns all task summaries", async () => {
|
||||
const { ctx } = ctxWithStore();
|
||||
await taskCreateTool.handler({ subject: "A", description: "x" }, ctx);
|
||||
await taskCreateTool.handler({ subject: "B", description: "y" }, ctx);
|
||||
const out = (await taskListTool.handler({}, ctx)) as { tasks: { subject: string }[] };
|
||||
expect(out.tasks.map((t) => t.subject)).toEqual(["A", "B"]);
|
||||
it("returns an empty list (not an error) when no store is available", async () => {
|
||||
const result = (await taskListTool.handler({}, ctxWith(undefined))) as { tasks: TaskSummary[] };
|
||||
expect(result.tasks).toEqual([]);
|
||||
});
|
||||
});
|
||||
|
||||
describe("task_get tool", () => {
|
||||
it("returns the full task including dependencies and metadata", async () => {
|
||||
const store = new TaskStore();
|
||||
const a = store.create({ subject: "A", description: "x" });
|
||||
const b = store.create({ subject: "B", description: "y" });
|
||||
store.update(b.id, { addBlockedBy: [a.id] });
|
||||
const result = (await taskGetTool.handler({ taskId: b.id }, ctxWith(store))) as { task?: Task };
|
||||
expect(result.task!.blockedBy).toEqual([a.id]);
|
||||
});
|
||||
|
||||
it("task_get returns full details and errors on unknown id", async () => {
|
||||
const { ctx } = ctxWithStore();
|
||||
const created = (await taskCreateTool.handler({ subject: "A", description: "long desc" }, ctx)) as {
|
||||
id: string;
|
||||
task: { description: string };
|
||||
it("returns an error for an unknown id", async () => {
|
||||
const store = new TaskStore();
|
||||
const result = (await taskGetTool.handler({ taskId: "t99" }, ctxWith(store))) as { error: string };
|
||||
expect(result.error).toMatch(/not found/i);
|
||||
});
|
||||
});
|
||||
|
||||
describe("task_update tool", () => {
|
||||
it("updates status and returns the updated task", async () => {
|
||||
const store = new TaskStore();
|
||||
const t = store.create({ subject: "A", description: "x" });
|
||||
const result = (await taskUpdateTool.handler({ taskId: t.id, status: "completed" }, ctxWith(store))) as {
|
||||
task?: Task;
|
||||
};
|
||||
const got = (await taskGetTool.handler({ taskId: created.id }, ctx)) as { task: { description: string } };
|
||||
expect(got.task.description).toBe("long desc");
|
||||
const miss = (await taskGetTool.handler({ taskId: "nope" }, ctx)) as { error: string };
|
||||
expect(miss.error).toMatch(/not found/);
|
||||
expect(result.task!.status).toBe("completed");
|
||||
});
|
||||
|
||||
it("task_update sets status and marks in_progress/completed", async () => {
|
||||
const { ctx } = ctxWithStore();
|
||||
const created = (await taskCreateTool.handler({ subject: "A", description: "x" }, ctx)) as { id: string };
|
||||
const upd = (await taskUpdateTool.handler({ taskId: created.id, status: "in_progress" }, ctx)) as {
|
||||
task: { status: string };
|
||||
it("deletes when status is 'deleted' and returns { deleted }", async () => {
|
||||
const store = new TaskStore();
|
||||
const t = store.create({ subject: "A", description: "x" });
|
||||
const result = (await taskUpdateTool.handler({ taskId: t.id, status: "deleted" }, ctxWith(store))) as {
|
||||
deleted: string;
|
||||
};
|
||||
expect(upd.task.status).toBe("in_progress");
|
||||
const done = (await taskUpdateTool.handler({ taskId: created.id, status: "completed" }, ctx)) as {
|
||||
task: { status: string };
|
||||
expect(result.deleted).toBe(t.id);
|
||||
expect(store.get(t.id)).toBeUndefined();
|
||||
});
|
||||
|
||||
it("adds dependencies via addBlockedBy", async () => {
|
||||
const store = new TaskStore();
|
||||
const a = store.create({ subject: "A", description: "x" });
|
||||
const b = store.create({ subject: "B", description: "y" });
|
||||
await taskUpdateTool.handler({ taskId: b.id, addBlockedBy: [a.id] }, ctxWith(store));
|
||||
expect(store.get(b.id)!.blockedBy).toEqual([a.id]);
|
||||
});
|
||||
|
||||
it("returns an error for an unknown id", async () => {
|
||||
const store = new TaskStore();
|
||||
const result = (await taskUpdateTool.handler({ taskId: "t99", status: "in_progress" }, ctxWith(store))) as {
|
||||
error: string;
|
||||
};
|
||||
expect(done.task.status).toBe("completed");
|
||||
expect(result.error).toMatch(/not found/i);
|
||||
});
|
||||
|
||||
it("addBlocks/addBlockedBy link two tasks both ways", async () => {
|
||||
const { ctx } = ctxWithStore();
|
||||
const a = (await taskCreateTool.handler({ subject: "A", description: "x" }, ctx)) as { id: string };
|
||||
const b = (await taskCreateTool.handler({ subject: "B", description: "y" }, ctx)) as { id: string };
|
||||
// B is blocked by A (one-directional: sets B.blockedBy, not A.blocks)
|
||||
await taskUpdateTool.handler({ taskId: b.id, addBlockedBy: [a.id] }, ctx);
|
||||
const bAfter = (await taskGetTool.handler({ taskId: b.id }, ctx)) as { task: { blockedBy: string[] } };
|
||||
expect(bAfter.task.blockedBy).toContain(a.id);
|
||||
// Add the back-ref explicitly: A blocks B.
|
||||
await taskUpdateTool.handler({ taskId: a.id, addBlocks: [b.id] }, ctx);
|
||||
const aAfter = (await taskGetTool.handler({ taskId: a.id }, ctx)) as { task: { blocks: string[] } };
|
||||
expect(aAfter.task.blocks).toContain(b.id);
|
||||
it("returns a clear error when no task store is available", async () => {
|
||||
const result = (await taskUpdateTool.handler({ taskId: "t1", status: "in_progress" }, ctxWith(undefined))) as {
|
||||
error: string;
|
||||
};
|
||||
expect(result.error).toMatch(/not available/i);
|
||||
});
|
||||
|
||||
it("ignores self-refs, unknown ids, and direct 2-cycles", async () => {
|
||||
const { ctx } = ctxWithStore();
|
||||
const a = (await taskCreateTool.handler({ subject: "A", description: "x" }, ctx)) as { id: string };
|
||||
const b = (await taskCreateTool.handler({ subject: "B", description: "y" }, ctx)) as { id: string };
|
||||
// self-ref ignored
|
||||
await taskUpdateTool.handler({ taskId: a.id, addBlockedBy: [a.id] }, ctx);
|
||||
expect(((await taskGetTool.handler({ taskId: a.id }, ctx)) as { task: { blockedBy: string[] } }).task.blockedBy).toEqual([]);
|
||||
// unknown id ignored
|
||||
await taskUpdateTool.handler({ taskId: a.id, addBlockedBy: ["zzz"] }, ctx);
|
||||
expect(((await taskGetTool.handler({ taskId: a.id }, ctx)) as { task: { blockedBy: string[] } }).task.blockedBy).toEqual([]);
|
||||
// B waits on A; now make A wait on B — should be skipped (2-cycle)
|
||||
await taskUpdateTool.handler({ taskId: b.id, addBlockedBy: [a.id] }, ctx);
|
||||
await taskUpdateTool.handler({ taskId: a.id, addBlockedBy: [b.id] }, ctx);
|
||||
expect(((await taskGetTool.handler({ taskId: a.id }, ctx)) as { task: { blockedBy: string[] } }).task.blockedBy).toEqual([]);
|
||||
it("the store emitter fires after a tool-driven update", async () => {
|
||||
const store = new TaskStore();
|
||||
const fired = vi.fn();
|
||||
store.setEmitter(fired);
|
||||
const t = store.create({ subject: "A", description: "x" });
|
||||
fired.mockClear(); // create already fired once; isolate the update
|
||||
await taskUpdateTool.handler({ taskId: t.id, status: "in_progress" }, ctxWith(store));
|
||||
expect(fired).toHaveBeenCalledTimes(1);
|
||||
});
|
||||
});
|
||||
|
||||
it("status deleted removes the task and prunes dangling refs", async () => {
|
||||
const { ctx } = ctxWithStore();
|
||||
const a = (await taskCreateTool.handler({ subject: "A", description: "x" }, ctx)) as { id: string };
|
||||
const b = (await taskCreateTool.handler({ subject: "B", description: "y" }, ctx)) as { id: string };
|
||||
await taskUpdateTool.handler({ taskId: b.id, addBlockedBy: [a.id] }, ctx);
|
||||
const del = (await taskUpdateTool.handler({ taskId: a.id, status: "deleted" }, ctx)) as { deleted: string };
|
||||
expect(del.deleted).toBe(a.id);
|
||||
// B no longer blocked by the removed A
|
||||
const bAfter = (await taskGetTool.handler({ taskId: b.id }, ctx)) as { task: { blockedBy: string[] } };
|
||||
expect(bAfter.task.blockedBy).toEqual([]);
|
||||
expect(((await taskListTool.handler({}, ctx)) as { tasks: unknown[] }).tasks).toHaveLength(1);
|
||||
describe("task tool schema validation", () => {
|
||||
it("rejects an empty subject on task_create", () => {
|
||||
expect(() => taskCreateTool.schema.parse({ subject: "", description: "x" })).toThrow();
|
||||
});
|
||||
|
||||
it("metadata merge-patch: set keys, null deletes", async () => {
|
||||
const { ctx } = ctxWithStore();
|
||||
const a = (await taskCreateTool.handler({ subject: "A", description: "x", metadata: { k: 1 } }, ctx)) as { id: string };
|
||||
await taskUpdateTool.handler({ taskId: a.id, metadata: { k2: "v" } }, ctx);
|
||||
let t = (await taskGetTool.handler({ taskId: a.id }, ctx)) as { task: { metadata: Record<string, unknown> } };
|
||||
expect(t.task.metadata).toEqual({ k: 1, k2: "v" });
|
||||
await taskUpdateTool.handler({ taskId: a.id, metadata: { k: null } }, ctx);
|
||||
t = (await taskGetTool.handler({ taskId: a.id }, ctx)) as { task: { metadata: Record<string, unknown> } };
|
||||
expect(t.task.metadata).toEqual({ k2: "v" });
|
||||
it("rejects an empty taskId on task_get/task_update", () => {
|
||||
expect(() => taskGetTool.schema.parse({ taskId: "" })).toThrow();
|
||||
expect(() => taskUpdateTool.schema.parse({ taskId: "" })).toThrow();
|
||||
});
|
||||
|
||||
it("returns an error when taskStore is absent", async () => {
|
||||
const ctx = {} as ToolContext;
|
||||
const out = (await taskCreateTool.handler({ subject: "A", description: "x" }, ctx)) as { error: string };
|
||||
expect(out.error).toMatch(/not available/);
|
||||
it("rejects an unknown status value on task_update", () => {
|
||||
expect(() => taskUpdateTool.schema.parse({ taskId: "t1", status: "done" })).toThrow();
|
||||
});
|
||||
it("accepts 'deleted' as a status on task_update", () => {
|
||||
expect(() => taskUpdateTool.schema.parse({ taskId: "t1", status: "deleted" })).not.toThrow();
|
||||
});
|
||||
});
|
||||
+63
-72
@@ -7,8 +7,8 @@ export type TaskStatus = "pending" | "in_progress" | "completed";
|
||||
|
||||
/** A structured, trackable unit of work. Tasks form a dependency graph via `blocks`/`blockedBy`
|
||||
* (each lists the other's task ids), can be owned/claimed by a named agent, and carry free-form
|
||||
* metadata. Unlike the flat todo list, tasks are created and updated incrementally (not replaced
|
||||
* wholesale) so dependencies and ownership can be expressed. */
|
||||
* metadata. Unlike the old flat todo list, tasks are created and updated incrementally (not
|
||||
* replaced wholesale) so dependencies and ownership can be expressed. */
|
||||
export interface Task {
|
||||
id: string;
|
||||
subject: string;
|
||||
@@ -36,21 +36,15 @@ export interface TaskSummary {
|
||||
blockedBy: string[];
|
||||
}
|
||||
|
||||
/** JSON-serializable shape for persisting a TaskStore. */
|
||||
export interface TaskStoreSnapshot {
|
||||
seq: number;
|
||||
tasks: Task[];
|
||||
}
|
||||
|
||||
/** In-memory task store. The task tools operate on it via `ctx.taskStore`. Mutations emit a
|
||||
* snapshot through an optional `onChange` callback the loop wires up, so each create/update can
|
||||
* refresh a UI checklist. Purely in-memory (per-session); not persisted. */
|
||||
/** In-memory task store held by the Session. The tools below operate on it via `ctx.taskStore`.
|
||||
* Mutations emit a snapshot to the UI through the `onChange` emitter the loop wires up, so each
|
||||
* create/update renders a fresh checklist item in the scrollback (matching the old todo behavior). */
|
||||
export class TaskStore {
|
||||
private tasks = new Map<string, Task>();
|
||||
private seq = 0;
|
||||
private emitter: ((tasks: TaskSummary[]) => void) | undefined;
|
||||
|
||||
/** Wired by the host so the store can broadcast a snapshot after each mutation. */
|
||||
/** Wired by the agent loop so the store can broadcast a snapshot after each mutation. */
|
||||
setEmitter(emit: (tasks: TaskSummary[]) => void): void {
|
||||
this.emitter = emit;
|
||||
}
|
||||
@@ -160,24 +154,6 @@ export class TaskStore {
|
||||
}
|
||||
this.emit();
|
||||
}
|
||||
|
||||
/** Serialize the entire store for persistence. */
|
||||
toJSON(): TaskStoreSnapshot {
|
||||
return {
|
||||
seq: this.seq,
|
||||
tasks: [...this.tasks.values()].map(serializeTask),
|
||||
};
|
||||
}
|
||||
|
||||
/** Restore a store from a previously-serialized snapshot. */
|
||||
static fromJSON(snapshot: TaskStoreSnapshot): TaskStore {
|
||||
const store = new TaskStore();
|
||||
store.seq = snapshot.seq;
|
||||
for (const task of snapshot.tasks) {
|
||||
store.tasks.set(task.id, { ...task, blocks: [...task.blocks], blockedBy: [...task.blockedBy], metadata: task.metadata ? { ...task.metadata } : undefined });
|
||||
}
|
||||
return store;
|
||||
}
|
||||
}
|
||||
|
||||
/** Merge-patches metadata: a null value deletes the key, any other value sets it. */
|
||||
@@ -206,23 +182,21 @@ function serializeTask(t: Task): Task {
|
||||
|
||||
const metadataSchema = z.record(z.string(), z.any()).optional();
|
||||
|
||||
const taskCreateSchema = z.object({
|
||||
subject: z.string().min(1).describe("A brief, actionable title in imperative form (e.g. 'Fix authentication bug')."),
|
||||
description: z.string().describe("What needs to be done, in enough detail to act on."),
|
||||
activeForm: z
|
||||
.string()
|
||||
.optional()
|
||||
.describe("Present-continuous label shown in the spinner while in_progress (e.g. 'Running tests'). Optional."),
|
||||
metadata: metadataSchema,
|
||||
});
|
||||
|
||||
export const taskCreateTool: ToolDef<z.infer<typeof taskCreateSchema>> = {
|
||||
name: "task_create",
|
||||
description:
|
||||
"Create a structured task to track a unit of multi-step work. Use for non-trivial work (3+ steps) so progress is " +
|
||||
"visible and dependencies can be expressed. Returns the new task with its id. Call task_list to see all tasks, " +
|
||||
"task_get for full details, and task_update to set status, add dependencies (addBlocks/addBlockedBy), or claim ownership.",
|
||||
schema: taskCreateSchema,
|
||||
"visible and dependencies can be expressed. Returns the new task with its id. Call task_list to see all tasks, " +
|
||||
"task_get for full details, and task_update to set status, add dependencies (addBlocks/addBlockedBy), or claim ownership.",
|
||||
schema: z.object({
|
||||
subject: z.string().min(1).describe("A brief, actionable title in imperative form (e.g. 'Fix authentication bug')."),
|
||||
description: z.string().describe("What needs to be done, in enough detail to act on."),
|
||||
activeForm: z
|
||||
.string()
|
||||
.optional()
|
||||
.describe("Present-continuous label shown in the spinner while in_progress (e.g. 'Running tests'). Optional."),
|
||||
metadata: metadataSchema,
|
||||
}),
|
||||
// Purely informational (tracks state in-memory, never touches the filesystem) — no confirmation prompt.
|
||||
mutating: false,
|
||||
handler: async (args, ctx) => {
|
||||
@@ -232,28 +206,33 @@ export const taskCreateTool: ToolDef<z.infer<typeof taskCreateSchema>> = {
|
||||
},
|
||||
};
|
||||
|
||||
const taskListSchema = z.object({});
|
||||
const taskCreateSchema = z.object({
|
||||
subject: z.string().min(1),
|
||||
description: z.string(),
|
||||
activeForm: z.string().optional(),
|
||||
metadata: metadataSchema,
|
||||
});
|
||||
|
||||
export const taskListTool: ToolDef<z.infer<typeof taskListSchema>> = {
|
||||
name: "task_list",
|
||||
description:
|
||||
"List all tasks with their id, subject, status, owner, and what blocks them. Use this to see overall progress and " +
|
||||
"find the next available task to claim.",
|
||||
schema: taskListSchema,
|
||||
"find the next available task to claim.",
|
||||
schema: z.object({}),
|
||||
mutating: false,
|
||||
handler: async (_args, ctx) => {
|
||||
return { tasks: ctx.taskStore?.list() ?? [] };
|
||||
},
|
||||
};
|
||||
|
||||
const taskGetSchema = z.object({ taskId: z.string().min(1) });
|
||||
const taskListSchema = z.object({});
|
||||
|
||||
export const taskGetTool: ToolDef<z.infer<typeof taskGetSchema>> = {
|
||||
name: "task_get",
|
||||
description:
|
||||
"Get a task's full details (description, activeForm, blocks, blockedBy, metadata). Use before starting a task to " +
|
||||
"verify its blockedBy list is empty — if it isn't, the blocking tasks must complete first.",
|
||||
schema: taskGetSchema,
|
||||
"verify its blockedBy list is empty — if it isn't, the blocking tasks must complete first.",
|
||||
schema: z.object({ taskId: z.string().min(1) }),
|
||||
mutating: false,
|
||||
handler: async (args, ctx) => {
|
||||
const task = ctx.taskStore?.get(args.taskId);
|
||||
@@ -261,6 +240,40 @@ export const taskGetTool: ToolDef<z.infer<typeof taskGetSchema>> = {
|
||||
},
|
||||
};
|
||||
|
||||
const taskGetSchema = z.object({ taskId: z.string().min(1) });
|
||||
|
||||
export const taskUpdateTool: ToolDef<z.infer<typeof taskUpdateSchema>> = {
|
||||
name: "task_update",
|
||||
description:
|
||||
"Update a task: set status (pending|in_progress|completed — or 'deleted' to remove it), rename subject/description, " +
|
||||
"set owner to claim it, add dependencies via addBlocks/addBlockedBy (task ids), or merge-patch metadata (set a key " +
|
||||
"to null to delete it). Mark a task in_progress when starting it and completed when done. Verify blockedBy is empty " +
|
||||
"before starting. Returns the updated task, or { deleted: id } when status is 'deleted'.",
|
||||
schema: z.object({
|
||||
taskId: z.string().min(1),
|
||||
status: z.enum(["pending", "in_progress", "completed", "deleted"]).optional(),
|
||||
subject: z.string().optional(),
|
||||
description: z.string().optional(),
|
||||
activeForm: z.string().optional(),
|
||||
owner: z.string().optional(),
|
||||
addBlocks: z.array(z.string()).optional(),
|
||||
addBlockedBy: z.array(z.string()).optional(),
|
||||
metadata: metadataSchema,
|
||||
}),
|
||||
mutating: false,
|
||||
handler: async (args, ctx) => {
|
||||
const store = ctx.taskStore;
|
||||
if (!store) return { error: "Task tracking is not available in this context." };
|
||||
if (!store.get(args.taskId)) return { error: `Task ${args.taskId} not found.` };
|
||||
if (args.status === "deleted") {
|
||||
store.update(args.taskId, args);
|
||||
return { deleted: args.taskId };
|
||||
}
|
||||
const task = store.update(args.taskId, args);
|
||||
return task ? { task: serializeTask(task) } : { deleted: args.taskId };
|
||||
},
|
||||
};
|
||||
|
||||
const taskUpdateSchema = z.object({
|
||||
taskId: z.string().min(1),
|
||||
status: z.enum(["pending", "in_progress", "completed", "deleted"]).optional(),
|
||||
@@ -271,26 +284,4 @@ const taskUpdateSchema = z.object({
|
||||
addBlocks: z.array(z.string()).optional(),
|
||||
addBlockedBy: z.array(z.string()).optional(),
|
||||
metadata: metadataSchema,
|
||||
});
|
||||
|
||||
export const taskUpdateTool: ToolDef<z.infer<typeof taskUpdateSchema>> = {
|
||||
name: "task_update",
|
||||
description:
|
||||
"Update a task: set status (pending|in_progress|completed — or 'deleted' to remove it), rename subject/description, " +
|
||||
"set owner to claim it, add dependencies via addBlocks/addBlockedBy (task ids), or merge-patch metadata (set a key " +
|
||||
"to null to delete it). Mark a task in_progress when starting it and completed when done. Verify blockedBy is empty " +
|
||||
"before starting. Returns the updated task, or { deleted: id } when status is 'deleted'.",
|
||||
schema: taskUpdateSchema,
|
||||
mutating: false,
|
||||
handler: async (args, ctx) => {
|
||||
const store = ctx.taskStore;
|
||||
if (!store) return { error: "Task tracking is not available in this context." };
|
||||
if (!store.get(args.taskId)) return { error: `Task ${args.taskId} not found.` };
|
||||
if (args.status === "deleted") {
|
||||
store.update(args.taskId, args);
|
||||
return { deleted: args.taskId };
|
||||
}
|
||||
const task = store.update(args.taskId, args);
|
||||
return task ? { task: serializeTask(task) } : { deleted: args.taskId };
|
||||
},
|
||||
};
|
||||
});
|
||||
@@ -0,0 +1,165 @@
|
||||
import { describe, expect, it, vi } from "vitest";
|
||||
import { agentTool } from "./agentTool.js";
|
||||
import { sendMessageTool } from "./sendMessage.js";
|
||||
import { listTeammatesTool } from "./teammates.js";
|
||||
import type { ToolContext } from "./types.js";
|
||||
|
||||
// A mock ctx whose teammate roster is a real Map, so registration/resolution/list behave like the
|
||||
// session-backed wiring in loop.ts (which sets registerTeammate/resolveTeammate/listTeammates over
|
||||
// session.namedAgents). runSubAgent/resumeSubAgent are vi mocks the individual tests program.
|
||||
function ctxWithRoster(opts?: {
|
||||
runSubAgent?: ToolContext["runSubAgent"];
|
||||
resumeSubAgent?: ToolContext["resumeSubAgent"];
|
||||
}): { ctx: ToolContext; roster: Map<string, string> } {
|
||||
const roster = new Map<string, string>();
|
||||
const ctx: ToolContext = {
|
||||
cwd: "/x",
|
||||
registerTeammate: (name, agentId) => {
|
||||
roster.set(name, agentId);
|
||||
},
|
||||
resolveTeammate: (name) => roster.get(name),
|
||||
listTeammates: () => [...roster.entries()].map(([name, agentId]) => ({ name, agentId })),
|
||||
...(opts?.runSubAgent ? { runSubAgent: opts.runSubAgent } : {}),
|
||||
...(opts?.resumeSubAgent ? { resumeSubAgent: opts.resumeSubAgent } : {}),
|
||||
};
|
||||
return { ctx, roster };
|
||||
}
|
||||
|
||||
describe("agent tool — named teammates", () => {
|
||||
it("registers a name on the roster when the delegation is resumable, and echoes name + agentId", async () => {
|
||||
const runSubAgent = vi.fn().mockResolvedValue({ agentId: "agent-1", result: "did it", resumable: true });
|
||||
const { ctx, roster } = ctxWithRoster({ runSubAgent });
|
||||
const result = (await agentTool.handler(
|
||||
{ description: "research X", prompt: "find X", name: "researcher" },
|
||||
ctx,
|
||||
)) as { name: string; agentId: string; result: string };
|
||||
expect(result.name).toBe("researcher");
|
||||
expect(result.agentId).toBe("agent-1");
|
||||
expect(roster.get("researcher")).toBe("agent-1");
|
||||
expect(runSubAgent).toHaveBeenCalledOnce();
|
||||
});
|
||||
|
||||
it("does NOT register a name (and omits name/agentId) when the delegation is NOT resumable (parallel/isolated)", async () => {
|
||||
const runSubAgent = vi.fn().mockResolvedValue({ agentId: "agent-2", result: "analyzed", resumable: false });
|
||||
const { ctx, roster } = ctxWithRoster({ runSubAgent });
|
||||
const result = (await agentTool.handler(
|
||||
{ description: "review Y", prompt: "review Y", name: "reviewer" },
|
||||
ctx,
|
||||
)) as { result: string; name?: string; agentId?: string };
|
||||
expect(result.result).toBe("analyzed");
|
||||
expect(result.name).toBeUndefined();
|
||||
expect(result.agentId).toBeUndefined();
|
||||
expect(roster.has("reviewer")).toBe(false);
|
||||
});
|
||||
|
||||
it("returns an error WITHOUT running when the name is already a teammate (no clobber)", async () => {
|
||||
const runSubAgent = vi.fn();
|
||||
const { ctx, roster } = ctxWithRoster({ runSubAgent });
|
||||
roster.set("researcher", "agent-1");
|
||||
const result = (await agentTool.handler(
|
||||
{ description: "research Z", prompt: "find Z", name: "researcher" },
|
||||
ctx,
|
||||
)) as { error: string };
|
||||
expect(result.error).toMatch(/already exists/i);
|
||||
expect(runSubAgent).not.toHaveBeenCalled();
|
||||
// The existing teammate is untouched.
|
||||
expect(roster.get("researcher")).toBe("agent-1");
|
||||
});
|
||||
|
||||
it("still runs (and silently ignores the name) when no roster is wired (non-session context)", async () => {
|
||||
// No resolveTeammate/registerTeammate on ctx — simulates a hand-built, non-session context.
|
||||
const runSubAgent = vi.fn().mockResolvedValue({ agentId: "agent-9", result: "ok", resumable: true });
|
||||
const ctx: ToolContext = { cwd: "/x", runSubAgent };
|
||||
const result = (await agentTool.handler(
|
||||
{ description: "solo", prompt: "do solo", name: "lonely" },
|
||||
ctx,
|
||||
)) as { agentId: string; name?: string };
|
||||
expect(runSubAgent).toHaveBeenCalledOnce();
|
||||
expect(result.agentId).toBe("agent-9");
|
||||
// No roster to register on → name simply not echoed as a teammate handle.
|
||||
expect(result.name).toBeUndefined();
|
||||
});
|
||||
});
|
||||
|
||||
describe("send_message — name vs agentId", () => {
|
||||
it("resolves a teammate by name, then calls resumeSubAgent with its agentId", async () => {
|
||||
const resumeSubAgent = vi.fn().mockResolvedValue("followed up");
|
||||
const { ctx, roster } = ctxWithRoster({ resumeSubAgent });
|
||||
roster.set("researcher", "agent-1");
|
||||
const result = (await sendMessageTool.handler({ name: "researcher", message: "go deeper" }, ctx)) as {
|
||||
agentId: string;
|
||||
name: string;
|
||||
result: string;
|
||||
};
|
||||
expect(resumeSubAgent).toHaveBeenCalledWith("agent-1", "go deeper");
|
||||
expect(result.agentId).toBe("agent-1");
|
||||
expect(result.name).toBe("researcher");
|
||||
expect(result.result).toBe("followed up");
|
||||
});
|
||||
|
||||
it("returns a clean error (without resuming) for an unknown name", async () => {
|
||||
const resumeSubAgent = vi.fn();
|
||||
const { ctx } = ctxWithRoster({ resumeSubAgent });
|
||||
const result = (await sendMessageTool.handler({ name: "ghost", message: "boo" }, ctx)) as { error: string };
|
||||
expect(result.error).toMatch(/No teammate named "ghost"/);
|
||||
expect(resumeSubAgent).not.toHaveBeenCalled();
|
||||
});
|
||||
|
||||
it("falls back to an explicit agentId when no name is given", async () => {
|
||||
const resumeSubAgent = vi.fn().mockResolvedValue("resumed by id");
|
||||
const { ctx } = ctxWithRoster({ resumeSubAgent });
|
||||
const result = (await sendMessageTool.handler({ agentId: "agent-7", message: "continue" }, ctx)) as {
|
||||
agentId: string;
|
||||
result: string;
|
||||
name?: string;
|
||||
};
|
||||
expect(resumeSubAgent).toHaveBeenCalledWith("agent-7", "continue");
|
||||
expect(result.agentId).toBe("agent-7");
|
||||
expect(result.name).toBeUndefined();
|
||||
});
|
||||
|
||||
it("returns a clean error when neither name nor agentId is provided", async () => {
|
||||
const resumeSubAgent = vi.fn();
|
||||
const { ctx } = ctxWithRoster({ resumeSubAgent });
|
||||
const result = (await sendMessageTool.handler({ message: "to whom?" }, ctx)) as { error: string };
|
||||
expect(result.error).toMatch(/either/i);
|
||||
expect(resumeSubAgent).not.toHaveBeenCalled();
|
||||
});
|
||||
});
|
||||
|
||||
describe("list_teammates tool", () => {
|
||||
it("returns the roster from ctx.listTeammates", async () => {
|
||||
const { ctx, roster } = ctxWithRoster();
|
||||
roster.set("researcher", "agent-1");
|
||||
roster.set("implementer", "agent-2");
|
||||
const result = (await listTeammatesTool.handler({}, ctx)) as { teammates: { name: string; agentId: string }[] };
|
||||
expect(result.teammates).toEqual([
|
||||
{ name: "researcher", agentId: "agent-1" },
|
||||
{ name: "implementer", agentId: "agent-2" },
|
||||
]);
|
||||
});
|
||||
|
||||
it("returns an empty list when no roster is wired (non-session context)", async () => {
|
||||
const result = (await listTeammatesTool.handler({}, { cwd: "/x" })) as { teammates: unknown[] };
|
||||
expect(result.teammates).toEqual([]);
|
||||
});
|
||||
|
||||
it("is non-mutating (no confirmation prompt)", () => {
|
||||
expect(listTeammatesTool.mutating).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe("teammate tool schema validation", () => {
|
||||
it("send_message requires at least one of name/agentId", () => {
|
||||
expect(() => sendMessageTool.schema.parse({ message: "x" })).toThrow();
|
||||
});
|
||||
it("send_message accepts a name only", () => {
|
||||
expect(() => sendMessageTool.schema.parse({ name: "researcher", message: "x" })).not.toThrow();
|
||||
});
|
||||
it("send_message accepts an agentId only", () => {
|
||||
expect(() => sendMessageTool.schema.parse({ agentId: "agent-1", message: "x" })).not.toThrow();
|
||||
});
|
||||
it("list_teammates takes no arguments", () => {
|
||||
expect(() => listTeammatesTool.schema.parse({})).not.toThrow();
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,19 @@
|
||||
import { z } from "zod";
|
||||
import type { ToolDef } from "./types.js";
|
||||
|
||||
const schema = z.object({}).describe("Takes no arguments.");
|
||||
|
||||
/** Lists the session's named teammates — sub-agents you spawned via the `agent` tool with a `name`
|
||||
* that are still resumable. Each entry is `{ name, agentId }`; address one with `send_message` by
|
||||
* its `name` (or its `agentId`). Returns an empty list when there are none, or in a non-session
|
||||
* context where no roster is maintained. Non-mutating: it only reads the name → agentId index. */
|
||||
export const listTeammatesTool: ToolDef<z.infer<typeof schema>> = {
|
||||
name: "list_teammates",
|
||||
description:
|
||||
"List your named teammates — sub-agents you spawned with a `name` via the `agent` tool that are still " +
|
||||
"resumable. Each has a name and agentId; continue one with send_message by its name. Returns an empty list " +
|
||||
"when none exist or in a non-session context.",
|
||||
schema,
|
||||
mutating: false,
|
||||
handler: async (_args, ctx) => ({ teammates: ctx.listTeammates?.() ?? [] }),
|
||||
};
|
||||
@@ -1,25 +0,0 @@
|
||||
import { describe, expect, it, vi } from "vitest";
|
||||
import { todoWriteTool } from "./todoWrite.js";
|
||||
|
||||
describe("todo_write", () => {
|
||||
it("is non-mutating (no confirmation prompt)", () => {
|
||||
expect(todoWriteTool.mutating).toBe(false);
|
||||
});
|
||||
|
||||
it("forwards the full list to ctx.setTodos and echoes it back", async () => {
|
||||
const setTodos = vi.fn();
|
||||
const todos = [
|
||||
{ content: "read the config", status: "completed" as const },
|
||||
{ content: "write the fix", status: "in_progress" as const },
|
||||
{ content: "run tests", status: "pending" as const },
|
||||
];
|
||||
const result = await todoWriteTool.handler({ todos }, { cwd: "/tmp", setTodos });
|
||||
expect(setTodos).toHaveBeenCalledWith(todos);
|
||||
expect(result).toEqual({ todos });
|
||||
});
|
||||
|
||||
it("doesn't throw when setTodos is absent from the context", async () => {
|
||||
const todos = [{ content: "a task", status: "pending" as const }];
|
||||
await expect(todoWriteTool.handler({ todos }, { cwd: "/tmp" })).resolves.toEqual({ todos });
|
||||
});
|
||||
});
|
||||
@@ -1,28 +0,0 @@
|
||||
import { z } from "zod";
|
||||
import type { ToolDef } from "./types.js";
|
||||
|
||||
const todoItemSchema = z.object({
|
||||
content: z.string().describe("Short description of the task."),
|
||||
status: z.enum(["pending", "in_progress", "completed"]),
|
||||
});
|
||||
|
||||
const schema = z.object({
|
||||
todos: z
|
||||
.array(todoItemSchema)
|
||||
.describe("Full checklist (replaces previous list, not a diff)."),
|
||||
});
|
||||
|
||||
export const todoWriteTool: ToolDef<z.infer<typeof schema>> = {
|
||||
name: "todo_write",
|
||||
description:
|
||||
"Show a task checklist for multi-step work (3+ steps). One item 'in_progress' at a time, " +
|
||||
"mark 'completed' when done. Pass the full list each time. Skip for trivial requests.",
|
||||
schema,
|
||||
// Purely informational (like Claude Code's TodoWrite) — never touches the filesystem or asks
|
||||
// the user anything, so it shouldn't interrupt the flow with a confirmation prompt.
|
||||
mutating: false,
|
||||
handler: async (args, ctx) => {
|
||||
ctx.setTodos?.(args.todos);
|
||||
return { todos: args.todos };
|
||||
},
|
||||
};
|
||||
+105
-26
@@ -1,5 +1,6 @@
|
||||
import type { z } from "zod";
|
||||
import type { TaskStore } from "./task.js";
|
||||
import type { CronStore } from "../scheduler/cron.js";
|
||||
|
||||
export interface SubAgentTask {
|
||||
/** Short (3-6 word) label shown in the UI while the sub-agent runs. */
|
||||
@@ -8,46 +9,58 @@ export interface SubAgentTask {
|
||||
prompt: string;
|
||||
}
|
||||
|
||||
/** Result of one sub-agent in a parallel batch: either its final text, or the error that
|
||||
* terminated it (timeout, MaxIterationsError, a thrown tool error, etc.). Kept separate from a
|
||||
* plain string so the parent model can see at a glance which sub-tasks succeeded and which it
|
||||
* needs to retry or work around — one failed sub-task shouldn't discard the (potentially
|
||||
* expensive) results of its siblings. */
|
||||
export interface SubAgentResult {
|
||||
description: string;
|
||||
/** The sub-agent's final text answer, or undefined if it failed before producing one. */
|
||||
result?: string;
|
||||
/** Present when the sub-agent failed. A timeout, MaxIterationsError, or any other thrown error
|
||||
* surfaces here rather than rejecting the whole batch. */
|
||||
error?: string;
|
||||
}
|
||||
|
||||
export interface SubAgentOverrides {
|
||||
/** Replaces locode's generic sub-agent system prompt entirely — used by plugin-defined agents
|
||||
* (agents/*.md) that ship their own identity/instructions instead of the generic "delegate a
|
||||
* task" framing. */
|
||||
systemPrompt?: string;
|
||||
/** Prepended (as an addendum) to the generic sub-agent system prompt instead of replacing it, so
|
||||
* the standard tool-use discipline survives. Used by built-in specialist agent types
|
||||
* (agentTypes.ts) like explore/code-reviewer. Ignored when `systemPrompt` is also set. */
|
||||
systemPromptAddendum?: string;
|
||||
/** Restricts the sub-agent's toolset to tools with these names (unknown names are silently
|
||||
* ignored); omit to inherit the parent's full toolset minus `agent`/plugin-agent tools. */
|
||||
toolNames?: string[];
|
||||
/** Working directory the sub-agent runs against. Set by the parallel-fan-out dispatch path to a
|
||||
* throwaway git worktree (see utils/worktree.ts) so concurrent sub-agents in one batch can't
|
||||
* collide on files. Omit (the default for a single/sequential delegation) to run in the parent's
|
||||
* cwd and persist edits. */
|
||||
cwd?: string;
|
||||
}
|
||||
|
||||
export interface TodoItem {
|
||||
content: string;
|
||||
status: "pending" | "in_progress" | "completed";
|
||||
/** Result of a sub-agent run: the agent's id (for later continuation via `resumeSubAgent`/the
|
||||
* `send_message` tool), its final answer text, and whether it's resumable. A sub-agent is resumable
|
||||
* only when it ran in the shared cwd (sequential single call, or a read-only parallel agent) — a
|
||||
* worktree-isolated parallel agent's cwd is cleaned up after the batch, so it can't be continued. */
|
||||
export interface SubAgentResult {
|
||||
agentId: string;
|
||||
result: string;
|
||||
resumable: boolean;
|
||||
}
|
||||
|
||||
export interface ToolContext {
|
||||
cwd: string;
|
||||
/** Only present when running inside a session capable of spawning sub-agents (used by the `agent`
|
||||
* tool). Runs a SINGLE sub-agent and returns its final text; a parallel batch is the `agent`
|
||||
* tool's own responsibility (it calls this once per task). */
|
||||
runSubAgent?: (task: SubAgentTask, overrides?: SubAgentOverrides) => Promise<string>;
|
||||
/** Replaces the session's task checklist (used by the `todo_write` tool). Absent only if a
|
||||
* future tool context is built without one — every session-backed context provides it. */
|
||||
setTodos?: (todos: TodoItem[]) => void;
|
||||
/** The session's structured task store (used by the task_create/list/get/update tools). Absent
|
||||
* only if a tool context is built without one — every session-backed context provides it. */
|
||||
/** Only present when running inside a session capable of spawning sub-agents (used by the `agent` tool). */
|
||||
runSubAgent?: (task: SubAgentTask, overrides?: SubAgentOverrides) => Promise<SubAgentResult>;
|
||||
/** Continues a previously-spawned resumable sub-agent (one that returned an agentId) with a
|
||||
* follow-up message, preserving its context. Used by the `send_message` tool. Rejects with a
|
||||
* clear error if the agentId is unknown or wasn't resumable. */
|
||||
resumeSubAgent?: (agentId: string, message: string) => Promise<string>;
|
||||
/** Registers a named teammate: maps `name` → `agentId` on the session's roster so `send_message`
|
||||
* can address it by name and `list_teammates` can show it. Used by the `agent` tool when the
|
||||
* model supplies a `name` and the delegation is resumable (shared-cwd). Absent in non-session
|
||||
* contexts, in which case the name is silently ignored (the agent still runs, just unnamed). */
|
||||
registerTeammate?: (name: string, agentId: string) => void;
|
||||
/** Resolves a teammate name to its agentId (or undefined if no such named teammate exists). Used by
|
||||
* `send_message` to address a teammate by name instead of agentId, and by the `agent` tool to
|
||||
* pre-flight a name collision before spawning a new teammate. Absent in non-session contexts. */
|
||||
resolveTeammate?: (name: string) => string | undefined;
|
||||
/** Lists all named teammates ({ name, agentId }) for the `list_teammates` tool. Absent in
|
||||
* non-session contexts — the tool returns an empty list then. */
|
||||
listTeammates?: () => { name: string; agentId: string }[];
|
||||
/** The session's structured task store, used by the `task_create`/`task_list`/`task_get`/
|
||||
* `task_update` tools to track multi-step work with dependencies and ownership. Absent only in a
|
||||
* non-session context (e.g. a hand-built test context) — the task tools return a clear error then. */
|
||||
taskStore?: TaskStore;
|
||||
/** Set only while this specific call is a backgroundable tool (currently just `bash`) — the tool
|
||||
* polls `requested` and, once true, detaches into the background job registry instead of
|
||||
@@ -58,6 +71,72 @@ export interface ToolContext {
|
||||
* tree so the command can't keep running (and keep the event loop alive on exit) after the
|
||||
* parent has already abandoned the turn. Undefined for top-level turns. */
|
||||
signal?: AbortSignal;
|
||||
/** Records the previous content of a file edited or written by this tool, so the user can later
|
||||
* roll back the most recent mutation via the /undo slash command. Only present in session-backed
|
||||
* contexts that provide a Session object. */
|
||||
setLastEdit?: (edit: { path: string; previousContent: string }) => void;
|
||||
/** Refreshes the session's cached user memory after the `memory_write` tool changes memory.md,
|
||||
* so a later system-prompt rebuild (compaction, /mode or /perm switch) re-folds the new content
|
||||
* instead of the pre-write snapshot. Only present in session-backed contexts. */
|
||||
setUserMemory?: (memory: string | null) => void;
|
||||
/** Surfaces a notice line to the parent UI's scrollback. Used by long-running, headless tools
|
||||
* (notably `workflow`) to report progress — `log()`/`phase()` inside a workflow script reach the
|
||||
* user through this. Absent in non-session contexts. */
|
||||
emitNotice?: (text: string, isError?: boolean) => void;
|
||||
/** Presents a plan for user approval and, on approval, exits plan mode so the caller can implement.
|
||||
* Used by the `exit_plan_mode` tool. Returns whether the user approved; on rejection plan mode
|
||||
* stays active so the model can refine and re-present. Absent outside plan mode (so the tool fails
|
||||
* cleanly with "only available in plan mode" if the model calls it at the wrong time). */
|
||||
exitPlanMode?: (plan: string) => Promise<{ approved: boolean }>;
|
||||
/** Asks the user a structured multiple-choice question (or a short sequence of them) when the model
|
||||
* is blocked on a decision only the user can make. Used by the `ask_user_question` tool. Resolves
|
||||
* with the selected option label(s) per question; an empty selection means the user skipped
|
||||
* (treated as "Other"/custom input in the UI, surfaced back as the typed text). Absent in
|
||||
* non-session contexts (headless sub-agents can't prompt). */
|
||||
askQuestion?: (questions: AskQuestionSpec[]) => Promise<AskQuestionAnswer[]>;
|
||||
/** The session's cron/wakeup scheduler, used by the `cron_create`/`cron_list`/`cron_delete`/
|
||||
* `schedule_wakeup` tools. Absent in a non-session context (e.g. a hand-built test context) — the
|
||||
* scheduling tools return a clear error then. */
|
||||
cronStore?: CronStore;
|
||||
/** Switches the session's working directory to `newCwd` (used by `enter_worktree`/`exit_worktree`)
|
||||
* and notifies the UI so its live cwd display, /undo resolution, git-info, and @mention handling
|
||||
* follow the switch. Absent in non-session contexts — the worktree tools fail cleanly then. */
|
||||
setCwd?: (newCwd: string) => void;
|
||||
/** Reads the active interactive worktree tracking ({ dir, branch, originalCwd }) or undefined when
|
||||
* not in a worktree session. Used by `enter_worktree` (refuse re-entry) and `exit_worktree`
|
||||
* (restore cwd + remove). */
|
||||
getWorktree?: () => { dir: string; branch: string; originalCwd: string } | undefined;
|
||||
/** Sets or clears the active worktree tracking on the session (and notifies the App so a /model
|
||||
* switch can re-attach it). `enter_worktree` sets it; `exit_worktree` passes undefined to clear. */
|
||||
setWorktree?: (worktree: { dir: string; branch: string; originalCwd: string } | undefined) => void;
|
||||
}
|
||||
|
||||
/** One selectable option in an {@link AskQuestionSpec}. `description` is shown dimmed under the label
|
||||
* to explain a trade-off or implication, so the user can compare options at a glance. */
|
||||
export interface AskQuestionOption {
|
||||
label: string;
|
||||
description?: string;
|
||||
}
|
||||
|
||||
/** A single question the model asks the user via the `ask_user_question` tool. `header` is a short
|
||||
* (≤ ~12 char) chip shown beside the question for scanability; `options` is 2-4 choices; when
|
||||
* `multiSelect` is true the user may pick several (otherwise exactly one). The UI also offers an
|
||||
* implicit "Other" path so the user can type a custom answer not in the list. */
|
||||
export interface AskQuestionSpec {
|
||||
question: string;
|
||||
header: string;
|
||||
options: AskQuestionOption[];
|
||||
multiSelect?: boolean;
|
||||
}
|
||||
|
||||
/** The user's answer to one {@link AskQuestionSpec}: the labels of the selected option(s), in the
|
||||
* order the model listed them. An empty array with `custom` set means the user typed a freeform
|
||||
* answer via "Other" instead of picking a listed option. */
|
||||
export interface AskQuestionAnswer {
|
||||
question: string;
|
||||
selected: string[];
|
||||
/** A freeform answer the user typed via "Other" instead of picking a listed option. */
|
||||
custom?: string;
|
||||
}
|
||||
|
||||
export interface ToolDef<T = any> {
|
||||
|
||||
@@ -19,8 +19,7 @@ function isBinaryContentType(contentType: string): boolean {
|
||||
export const webFetchTool: ToolDef<z.infer<typeof schema>> = {
|
||||
name: "web_fetch",
|
||||
description:
|
||||
"Fetch a URL and return readable text (HTML tags/scripts/styles stripped). Use for specific pages found via web_search — " +
|
||||
"e.g. to read a doc page or blog post in full when the search snippet was not enough.",
|
||||
"Fetch a URL and return readable text (HTML tags/scripts/styles stripped). Use for specific pages found via web_search.",
|
||||
schema,
|
||||
mutating: false,
|
||||
handler: async ({ url }) => {
|
||||
|
||||
@@ -45,8 +45,7 @@ function parseResults(html: string, limit: number): SearchResult[] {
|
||||
export const webSearchTool: ToolDef<z.infer<typeof schema>> = {
|
||||
name: "web_search",
|
||||
description:
|
||||
"Search the web via DuckDuckGo. Returns title, url, snippet. Use for info not in the local codebase — " +
|
||||
"e.g. an unfamiliar API, library docs, or an error message. Follow up with web_fetch on a specific result for full page text.",
|
||||
"Search the web via DuckDuckGo. Returns title, url, snippet. Use for info not in the local codebase.",
|
||||
schema,
|
||||
mutating: false,
|
||||
handler: async ({ query, max_results }) => {
|
||||
|
||||
@@ -0,0 +1,377 @@
|
||||
import { describe, expect, it, vi, beforeEach, afterEach } from "vitest";
|
||||
import { execa } from "execa";
|
||||
import { existsSync, mkdtempSync, rmSync, writeFileSync } from "node:fs";
|
||||
import { tmpdir } from "node:os";
|
||||
import path from "node:path";
|
||||
import { stripExports, extractJson, validateAgainstSchema, runWorkflow } from "./workflow.js";
|
||||
import type { SubAgentResult } from "./types.js";
|
||||
|
||||
describe("stripExports", () => {
|
||||
it("strips leading export keywords so module-style scripts run as plain scripts", () => {
|
||||
expect(stripExports("export const meta = { name: 'x' };\nconst y = 1;")).toBe("const meta = { name: 'x' };\nconst y = 1;");
|
||||
expect(stripExports("export function f() {}\nexport default 1;")).toBe("function f() {}\n1;");
|
||||
// Non-export lines are untouched.
|
||||
expect(stripExports("const a = 1;\n// export b\nconst c = 3;")).toBe("const a = 1;\n// export b\nconst c = 3;");
|
||||
});
|
||||
});
|
||||
|
||||
describe("extractJson", () => {
|
||||
it("parses plain JSON", () => {
|
||||
expect(extractJson('{"a":1}')).toEqual({ a: 1 });
|
||||
expect(extractJson("[1,2,3]")).toEqual([1, 2, 3]);
|
||||
});
|
||||
it("extracts JSON from a ```json fence", () => {
|
||||
expect(extractJson("Here you go:\n```json\n{\"a\":1}\n```\nthanks")).toEqual({ a: 1 });
|
||||
});
|
||||
it("extracts JSON from a fence with an arbitrary language tag (```javascript)", () => {
|
||||
expect(extractJson("```javascript\n{\"a\":1}\n```")).toEqual({ a: 1 });
|
||||
expect(extractJson("```ts\n{\"a\":1}\n```")).toEqual({ a: 1 });
|
||||
});
|
||||
it("strips trailing commas before } (a common local-model artifact)", () => {
|
||||
expect(extractJson('{"a":1,"b":2,}')).toEqual({ a: 1, b: 2 });
|
||||
expect(extractJson('```json\n{\n "a": 1,\n "b": 2,\n}\n```')).toEqual({ a: 1, b: 2 });
|
||||
});
|
||||
it("strips trailing commas before ] in arrays", () => {
|
||||
expect(extractJson("[1,2,3,]")).toEqual([1, 2, 3]);
|
||||
});
|
||||
it("extracts the first JSON object from surrounding prose", () => {
|
||||
expect(extractJson('The answer is {"a":1,"b":2} as shown.')).toEqual({ a: 1, b: 2 });
|
||||
});
|
||||
it("throws when no JSON is present", () => {
|
||||
expect(() => extractJson("no json here")).toThrow(/valid JSON/);
|
||||
});
|
||||
});
|
||||
|
||||
describe("validateAgainstSchema", () => {
|
||||
const sch = { type: "object" as const, required: ["a", "b"], properties: { a: { type: "number" }, b: { type: "string" } } };
|
||||
it("passes a valid object", () => {
|
||||
expect(() => validateAgainstSchema({ a: 1, b: "x" }, sch)).not.toThrow();
|
||||
});
|
||||
it("rejects a missing required field", () => {
|
||||
expect(() => validateAgainstSchema({ a: 1 }, sch)).toThrow(/missing required field: b/);
|
||||
});
|
||||
it("rejects a wrong property type", () => {
|
||||
expect(() => validateAgainstSchema({ a: "notnum", b: "x" }, sch)).toThrow(/field a: expected number/);
|
||||
});
|
||||
it("rejects a non-object when object is required", () => {
|
||||
expect(() => validateAgainstSchema([1, 2], sch)).toThrow(/expected a JSON object/);
|
||||
});
|
||||
it("validates array item types via items", () => {
|
||||
const arrSch = { type: "array", items: { type: "string" } };
|
||||
expect(() => validateAgainstSchema(["a", "b"], arrSch)).not.toThrow();
|
||||
expect(() => validateAgainstSchema(["a", 2], arrSch)).toThrow(/item\[1\]: expected string/);
|
||||
});
|
||||
it("validates enum membership", () => {
|
||||
const enumSch = { type: "string", enum: ["low", "med", "high"] };
|
||||
expect(() => validateAgainstSchema("med", enumSch)).not.toThrow();
|
||||
expect(() => validateAgainstSchema("nope", enumSch)).toThrow(/not in allowed enum/);
|
||||
});
|
||||
it("validates boolean and integer property types", () => {
|
||||
const s = { type: "object", properties: { ok: { type: "boolean" }, n: { type: "integer" } } };
|
||||
expect(() => validateAgainstSchema({ ok: true, n: 3 }, s)).not.toThrow();
|
||||
expect(() => validateAgainstSchema({ ok: "yes", n: 3 }, s)).toThrow(/field ok: expected boolean/);
|
||||
// integer must reject a non-integer number.
|
||||
expect(() => validateAgainstSchema({ ok: true, n: 1.5 }, s)).toThrow(/field n: expected integer/);
|
||||
});
|
||||
it("validates nested object properties", () => {
|
||||
const s = {
|
||||
type: "object",
|
||||
properties: { outer: { type: "object", required: ["inner"], properties: { inner: { type: "number" } } } },
|
||||
};
|
||||
expect(() => validateAgainstSchema({ outer: { inner: 5 } }, s)).not.toThrow();
|
||||
expect(() => validateAgainstSchema({ outer: {} }, s)).toThrow(/missing required field: inner/);
|
||||
expect(() => validateAgainstSchema({ outer: { inner: "x" } }, s)).toThrow(/field outer.inner: expected number/);
|
||||
});
|
||||
it("rejects additional properties when additionalProperties is false", () => {
|
||||
const s = { type: "object", properties: { a: { type: "number" } }, additionalProperties: false };
|
||||
expect(() => validateAgainstSchema({ a: 1 }, s)).not.toThrow();
|
||||
expect(() => validateAgainstSchema({ a: 1, extra: 2 }, s)).toThrow(/unexpected additional property: extra/);
|
||||
});
|
||||
});
|
||||
|
||||
function mockSubAgent(result: string): { runSubAgent: ReturnType<typeof vi.fn> } {
|
||||
return { runSubAgent: vi.fn(async (): Promise<SubAgentResult> => ({ agentId: "id", result, resumable: false })) };
|
||||
}
|
||||
|
||||
describe("runWorkflow", () => {
|
||||
it("runs a script that returns a value, calling agent() through ctx.runSubAgent", async () => {
|
||||
const ctx = mockSubAgent("the answer");
|
||||
const out = await runWorkflow(
|
||||
`const r = await agent("do something", { label: "worker" });\nreturn r;`,
|
||||
undefined,
|
||||
ctx,
|
||||
);
|
||||
expect(out).toBe("the answer");
|
||||
expect(ctx.runSubAgent).toHaveBeenCalledTimes(1);
|
||||
expect(ctx.runSubAgent.mock.calls[0]![0].description).toBe("worker");
|
||||
});
|
||||
|
||||
it("exposes args to the script", async () => {
|
||||
const ctx = mockSubAgent("ok");
|
||||
const out = await runWorkflow(
|
||||
`const items = args;\nconst r = await agent("process " + items.join(","));\nreturn r;`,
|
||||
["a", "b", "c"],
|
||||
ctx,
|
||||
);
|
||||
expect(out).toBe("ok");
|
||||
expect(ctx.runSubAgent.mock.calls[0]![0].prompt).toBe("process a,b,c");
|
||||
});
|
||||
|
||||
it("parallel() runs thunks concurrently and turns failures into null", async () => {
|
||||
const ctx = mockSubAgent("ok");
|
||||
const out = (await runWorkflow(
|
||||
`const results = await parallel([
|
||||
() => agent("task A").then(r => r + "!"),
|
||||
() => Promise.reject(new Error("boom")),
|
||||
() => agent("task C"),
|
||||
]);
|
||||
return results;`,
|
||||
undefined,
|
||||
ctx,
|
||||
)) as (string | null)[];
|
||||
expect(out).toHaveLength(3);
|
||||
expect(out[0]).toBe("ok!");
|
||||
expect(out[1]).toBeNull();
|
||||
expect(out[2]).toBe("ok");
|
||||
});
|
||||
|
||||
it("pipeline() runs each item through all stages, no barrier between stages", async () => {
|
||||
// A stage-2 item can finish before a slow stage-1 item — but for determinism in this test we
|
||||
// just assert each item passes through both stages in order and results land in input order.
|
||||
const ctx = mockSubAgent("ok");
|
||||
const out = (await runWorkflow(
|
||||
`const out = await pipeline(
|
||||
["a", "b", "c"],
|
||||
async (item) => item + "1",
|
||||
async (item) => item + "2",
|
||||
);
|
||||
return out;`,
|
||||
undefined,
|
||||
ctx,
|
||||
)) as string[];
|
||||
expect(out).toEqual(["a12", "b12", "c12"]);
|
||||
});
|
||||
|
||||
it("a pipeline stage that throws drops just that item to null", async () => {
|
||||
const ctx = mockSubAgent("ok");
|
||||
const out = (await runWorkflow(
|
||||
`const out = await pipeline(
|
||||
["a", "b", "c"],
|
||||
async (item) => { if (item === "b") throw new Error("nope"); return item + "1"; },
|
||||
async (item) => item + "2",
|
||||
);
|
||||
return out;`,
|
||||
undefined,
|
||||
ctx,
|
||||
)) as (string | null)[];
|
||||
expect(out).toEqual(["a12", null, "c12"]);
|
||||
});
|
||||
|
||||
it("agent() with a schema returns a parsed, validated object (retrying once on bad JSON)", async () => {
|
||||
// First call returns non-JSON; the retry (nudge prompt) returns valid JSON matching the schema.
|
||||
const calls: string[] = [];
|
||||
const ctx = {
|
||||
runSubAgent: vi.fn(async (task: { prompt: string }): Promise<SubAgentResult> => {
|
||||
calls.push(task.prompt);
|
||||
if (task.prompt.includes("not valid JSON")) {
|
||||
return { agentId: "id", result: '{"answer": 42}', resumable: false };
|
||||
}
|
||||
return { agentId: "id", result: "I think the answer is 42.", resumable: false };
|
||||
}),
|
||||
};
|
||||
const out = await runWorkflow(
|
||||
`const r = await agent("what is the answer", {
|
||||
schema: { type: "object", required: ["answer"], properties: { answer: { type: "number" } } },
|
||||
});
|
||||
return r;`,
|
||||
undefined,
|
||||
ctx,
|
||||
);
|
||||
expect(out).toEqual({ answer: 42 });
|
||||
expect(ctx.runSubAgent).toHaveBeenCalledTimes(2);
|
||||
// The retry carried a nudge.
|
||||
expect(calls[1]).toContain("not valid JSON");
|
||||
});
|
||||
|
||||
it("agent() with a schema throws if the retry still doesn't validate", async () => {
|
||||
const ctx = mockSubAgent("still not json at all");
|
||||
await expect(
|
||||
runWorkflow(
|
||||
`await agent("x", { schema: { type: "object", required: ["answer"] } });`,
|
||||
undefined,
|
||||
ctx,
|
||||
),
|
||||
).rejects.toThrow(/valid JSON/);
|
||||
});
|
||||
|
||||
it("agent() retry nudge includes the specific validation error from the first attempt", async () => {
|
||||
const calls: string[] = [];
|
||||
const ctx = {
|
||||
runSubAgent: vi.fn(async (task: { prompt: string }): Promise<SubAgentResult> => {
|
||||
calls.push(task.prompt);
|
||||
if (calls.length === 2) {
|
||||
// Retry returns valid JSON.
|
||||
return { agentId: "id", result: '{"answer": 42}', resumable: false };
|
||||
}
|
||||
// First call returns JSON with a wrong type (answer is a string, schema wants number).
|
||||
return { agentId: "id", result: '{"answer": "forty-two"}', resumable: false };
|
||||
}),
|
||||
};
|
||||
const out = await runWorkflow(
|
||||
`const r = await agent("what is the answer", {
|
||||
schema: { type: "object", required: ["answer"], properties: { answer: { type: "number" } } },
|
||||
});
|
||||
return r;`,
|
||||
undefined,
|
||||
ctx,
|
||||
);
|
||||
expect(out).toEqual({ answer: 42 });
|
||||
// The retry nudge quoted the first attempt's validation error (wrong type for `answer`).
|
||||
expect(calls[1]).toContain("not valid JSON");
|
||||
expect(calls[1]).toMatch(/answer.*expected number|expected number.*answer/);
|
||||
});
|
||||
|
||||
it("caps concurrency so a fan-out doesn't exceed the limit", async () => {
|
||||
process.env.LOCODE_WORKFLOW_CONCURRENCY = "2";
|
||||
try {
|
||||
let active = 0;
|
||||
let maxActive = 0;
|
||||
const ctx = {
|
||||
runSubAgent: vi.fn(async (): Promise<SubAgentResult> => {
|
||||
active++;
|
||||
maxActive = Math.max(maxActive, active);
|
||||
await new Promise((r) => setTimeout(r, 30));
|
||||
active--;
|
||||
return { agentId: "id", result: "ok", resumable: false };
|
||||
}),
|
||||
};
|
||||
await runWorkflow(
|
||||
`await parallel([
|
||||
() => agent("1"), () => agent("2"), () => agent("3"),
|
||||
() => agent("4"), () => agent("5"), () => agent("6"),
|
||||
]);`,
|
||||
undefined,
|
||||
ctx,
|
||||
);
|
||||
expect(ctx.runSubAgent).toHaveBeenCalledTimes(6);
|
||||
expect(maxActive).toBeLessThanOrEqual(2);
|
||||
} finally {
|
||||
delete process.env.LOCODE_WORKFLOW_CONCURRENCY;
|
||||
}
|
||||
});
|
||||
|
||||
it("log() and phase() surface notices through ctx.emitNotice", async () => {
|
||||
const notices: { text: string; isError?: boolean }[] = [];
|
||||
const ctx = {
|
||||
runSubAgent: vi.fn(async (): Promise<SubAgentResult> => ({ agentId: "id", result: "ok", resumable: false })),
|
||||
emitNotice: (text: string, isError?: boolean) => notices.push({ text, isError }),
|
||||
};
|
||||
await runWorkflow(
|
||||
`phase("Review");\nlog("halfway");\nconst r = await agent("x");\nlog("done");\nreturn r;`,
|
||||
undefined,
|
||||
ctx,
|
||||
);
|
||||
expect(notices).toContainEqual({ text: "▶ Review", isError: undefined });
|
||||
expect(notices.map((n) => n.text)).toEqual(["▶ Review", "halfway", "done"]);
|
||||
});
|
||||
|
||||
it("a thrown error fails the workflow", async () => {
|
||||
const ctx = mockSubAgent("ok");
|
||||
await expect(runWorkflow(`throw new Error("script broke");`, undefined, ctx)).rejects.toThrow("script broke");
|
||||
});
|
||||
|
||||
it("nested workflow() is rejected", async () => {
|
||||
const ctx = mockSubAgent("ok");
|
||||
await expect(runWorkflow(`await workflow();`, undefined, ctx)).rejects.toThrow(/cannot be nested/);
|
||||
});
|
||||
|
||||
it("resolves agentType to the specialist's toolset+addendum via overrides", async () => {
|
||||
const ctx = {
|
||||
runSubAgent: vi.fn(async (_task: unknown, overrides?: unknown): Promise<SubAgentResult> => {
|
||||
// Stash the overrides for assertion; return a plain result.
|
||||
(ctx as any).__overrides = overrides;
|
||||
return { agentId: "id", result: "ok", resumable: false };
|
||||
}),
|
||||
};
|
||||
await runWorkflow(`await agent("explore the repo", { agentType: "explore" });`, undefined, ctx);
|
||||
const overrides = (ctx as any).__overrides as { toolNames?: string[]; systemPromptAddendum?: string };
|
||||
expect(overrides?.toolNames).toContain("read_file");
|
||||
expect(overrides?.systemPromptAddendum).toContain("Explore agent");
|
||||
});
|
||||
});
|
||||
|
||||
describe("runWorkflow worktree isolation", () => {
|
||||
let repo: string;
|
||||
|
||||
beforeEach(() => {
|
||||
repo = mkdtempSync(path.join(tmpdir(), "locode-wf-wt-test-"));
|
||||
});
|
||||
|
||||
afterEach(async () => {
|
||||
await execa("git", ["worktree", "prune"], { cwd: repo, reject: false }).catch(() => {});
|
||||
try {
|
||||
rmSync(repo, { recursive: true, force: true });
|
||||
} catch {
|
||||
/* leave for the OS temp sweep */
|
||||
}
|
||||
});
|
||||
|
||||
async function gitInit(r: string): Promise<void> {
|
||||
await execa("git", ["init", "-q"], { cwd: r });
|
||||
await execa("git", ["config", "user.email", "t@t"], { cwd: r });
|
||||
await execa("git", ["config", "user.name", "t"], { cwd: r });
|
||||
writeFileSync(path.join(r, "README.md"), "hello\n");
|
||||
await execa("git", ["add", "."], { cwd: r });
|
||||
await execa("git", ["commit", "-q", "-m", "init"], { cwd: r });
|
||||
}
|
||||
|
||||
it("agent({isolation:'worktree'}) runs in a real worktree cwd distinct from the repo, cleaned up after", async () => {
|
||||
await gitInit(repo);
|
||||
let agentCwd: string | undefined;
|
||||
const ctx = {
|
||||
cwd: repo,
|
||||
runSubAgent: vi.fn(async (_task: unknown, overrides?: { cwd?: string }): Promise<SubAgentResult> => {
|
||||
agentCwd = overrides?.cwd;
|
||||
return { agentId: "id", result: "ok", resumable: false };
|
||||
}),
|
||||
};
|
||||
await runWorkflow(`await agent("edit things", { isolation: "worktree" });`, undefined, ctx);
|
||||
|
||||
// The agent ran in a throwaway worktree, not the shared repo cwd.
|
||||
expect(agentCwd).toBeDefined();
|
||||
expect(agentCwd).not.toBe(repo);
|
||||
expect(existsSync(agentCwd!)).toBe(false); // cleaned up after the agent finished
|
||||
});
|
||||
|
||||
it("agent({isolation:'worktree'}) on a non-git cwd falls back to the shared cwd (no isolation)", async () => {
|
||||
// repo exists but has no .git (gitInit not called) → createWorktree returns cwd:undefined.
|
||||
let agentCwd: string | undefined;
|
||||
const ctx = {
|
||||
cwd: repo,
|
||||
runSubAgent: vi.fn(async (_task: unknown, overrides?: { cwd?: string }): Promise<SubAgentResult> => {
|
||||
agentCwd = overrides?.cwd;
|
||||
return { agentId: "id", result: "ok", resumable: false };
|
||||
}),
|
||||
};
|
||||
await runWorkflow(`await agent("edit things", { isolation: "worktree" });`, undefined, ctx);
|
||||
// No git → no worktree → the agent runs in the shared cwd (undefined override = inherit parent).
|
||||
expect(agentCwd).toBeUndefined();
|
||||
});
|
||||
|
||||
it("agent({isolation:'worktree'}) still cleans up the worktree even when the agent throws", async () => {
|
||||
await gitInit(repo);
|
||||
let agentCwd: string | undefined;
|
||||
const ctx = {
|
||||
cwd: repo,
|
||||
runSubAgent: vi.fn(async (_task: unknown, overrides?: { cwd?: string }): Promise<SubAgentResult> => {
|
||||
agentCwd = overrides?.cwd;
|
||||
throw new Error("agent exploded");
|
||||
}),
|
||||
};
|
||||
await expect(
|
||||
runWorkflow(`await agent("edit things", { isolation: "worktree" });`, undefined, ctx),
|
||||
).rejects.toThrow("agent exploded");
|
||||
expect(agentCwd).toBeDefined();
|
||||
expect(existsSync(agentCwd!)).toBe(false); // the finally cleaned up despite the throw
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,358 @@
|
||||
import vm from "node:vm";
|
||||
import { z } from "zod";
|
||||
import { getAgentType } from "./agentTypes.js";
|
||||
import type { SubAgentOverrides, SubAgentResult, ToolDef } from "./types.js";
|
||||
import { createWorktree, type WorktreeIsolation } from "../utils/worktree.js";
|
||||
|
||||
/** Default cap on concurrent in-flight sub-agents inside a workflow. Local backends (Ollama, LM
|
||||
* Studio) typically serve a single model and can't truly parallelize many concurrent request
|
||||
* streams — an unbounded fan-out would queue a large burst and risk OOM/timeout. The cap keeps the
|
||||
* burst bounded; the model server serializes what it can't concurrentize. Override via
|
||||
* LOCODE_WORKFLOW_CONCURRENCY. */
|
||||
function workflowConcurrency(): number {
|
||||
const raw = Number(process.env.LOCODE_WORKFLOW_CONCURRENCY);
|
||||
if (Number.isFinite(raw) && raw >= 1) return Math.floor(raw);
|
||||
return 4;
|
||||
}
|
||||
|
||||
/** Wall-clock backstop for the whole workflow (the script itself is fast between agent calls; the
|
||||
* real time is in the agents, each already bounded by SUBAGENT_TIMEOUT_MS). This catches a script
|
||||
* that loops forever spawning agents. Override via LOCODE_WORKFLOW_TIMEOUT_MS. */
|
||||
function workflowTimeoutMs(): number {
|
||||
const raw = Number(process.env.LOCODE_WORKFLOW_TIMEOUT_MS);
|
||||
if (Number.isFinite(raw) && raw >= 1000) return Math.floor(raw);
|
||||
return 15 * 60 * 1000;
|
||||
}
|
||||
|
||||
/** A JSON-Schema-ish spec for forcing structured output from a sub-agent. Minimal validation only
|
||||
* (required keys, property/item types, enums, nested objects) — locode can't force a native tool
|
||||
* call for structured output the way Claude Code can, so this is best-effort: the sub-agent is told
|
||||
* to return JSON, we parse it, and retry once if it doesn't validate. */
|
||||
interface JsonSchemaProperty {
|
||||
type?: string;
|
||||
enum?: unknown[];
|
||||
items?: JsonSchemaProperty;
|
||||
properties?: Record<string, JsonSchemaProperty>;
|
||||
required?: string[];
|
||||
additionalProperties?: boolean;
|
||||
}
|
||||
|
||||
interface JsonSchema extends JsonSchemaProperty {}
|
||||
|
||||
const schema = z.object({
|
||||
script: z
|
||||
.string()
|
||||
.describe(
|
||||
"A self-contained JavaScript workflow script. It may begin with `const meta = { name, description, phases }` " +
|
||||
"(metadata only). Use the provided globals to orchestrate: agent(prompt, opts?) runs a sub-agent and returns its " +
|
||||
"answer (or a validated object when opts.schema is given); parallel([() => ..., () => ...]) runs thunks concurrently " +
|
||||
"(failures become null); pipeline(items, stage1, stage2, ...) runs each item through all stages with no barrier between " +
|
||||
"stages; phase(title) marks a progress group; log(message) surfaces a progress line to the user. Return a value to make " +
|
||||
"it the workflow's result. No filesystem or Node APIs; no Date.now/Math.random. A thrown error fails the workflow.",
|
||||
),
|
||||
args: z
|
||||
.any()
|
||||
.optional()
|
||||
.describe("Free-form value exposed to the script as the global `args` (pass arrays/objects, not a stringified string)."),
|
||||
});
|
||||
|
||||
/** Strips ES-module `export` keywords so a script written in Claude-Code's `export const meta` style
|
||||
* runs as plain script source inside the vm. Handles `export const/function/default` at line starts. */
|
||||
function stripExports(src: string): string {
|
||||
return src.replace(/^export\s+(default\s+)?/gm, "");
|
||||
}
|
||||
|
||||
/** Extracts the first balanced JSON value (object or array) from text that may have surrounding
|
||||
* prose/code fences — local models often wrap JSON in ```json … ``` (or ```javascript, ```ts, …)
|
||||
* and add commentary. Tolerates trailing commas, which local models emit frequently even though
|
||||
* they're invalid JSON. */
|
||||
function extractJson(text: string): unknown {
|
||||
const trimmed = text.trim();
|
||||
// Strip a surrounding ```lang ... ``` fence with any language tag (json, javascript, ts, …).
|
||||
// Local models label fences with whatever language they think the content is, so accept any tag.
|
||||
const fenced = trimmed.match(/```[a-zA-Z0-9+#]*\s*([\s\S]*?)```/);
|
||||
const candidate = (fenced ? fenced[1]! : trimmed).trim();
|
||||
// Local models frequently emit trailing commas before } or ] (invalid JSON). Strip them.
|
||||
const sanitized = stripTrailingCommas(candidate);
|
||||
try {
|
||||
return JSON.parse(sanitized);
|
||||
} catch {
|
||||
// Fall back to the first {...} or [...] span, also comma-sanitized.
|
||||
const span = candidate.match(/(\{[\s\S]*\}|\[[\s\S]*\])/);
|
||||
if (span) {
|
||||
try {
|
||||
return JSON.parse(stripTrailingCommas(span[1]!));
|
||||
} catch {
|
||||
/* fall through */
|
||||
}
|
||||
}
|
||||
}
|
||||
throw new Error("Sub-agent did not return valid JSON for structured output.");
|
||||
}
|
||||
|
||||
/** Removes trailing commas that immediately precede a closing } or ] — a common local-model
|
||||
* artifact that makes otherwise-valid JSON unparseable. */
|
||||
function stripTrailingCommas(s: string): string {
|
||||
return s.replace(/,\s*([}\]])/g, "$1");
|
||||
}
|
||||
|
||||
/** Minimal JSON-Schema validation: checks the top-level shape (object/array), required keys,
|
||||
* property types, array item types, enum membership, and nested object properties. Intentionally
|
||||
* not a full validator — just enough to catch an obviously-wrong shape and trigger the one retry
|
||||
* in agentFn. Error messages are kept stable (e.g. "missing required field: x") so the retry nudge
|
||||
* can quote them back to the sub-agent. */
|
||||
function validateAgainstSchema(value: unknown, schema: JsonSchema): void {
|
||||
// Top-level shape checks keep the stable, friendly messages callers (and the retry nudge) rely on.
|
||||
if (schema.type === "object") {
|
||||
if (typeof value !== "object" || value === null || Array.isArray(value)) {
|
||||
throw new Error("expected a JSON object");
|
||||
}
|
||||
} else if (schema.type === "array") {
|
||||
if (!Array.isArray(value)) throw new Error("expected a JSON array");
|
||||
}
|
||||
// Deeper checks (required keys, property types, items, enum, nested objects) share validateProperty.
|
||||
validateProperty(value, schema, "");
|
||||
}
|
||||
|
||||
/** Validates a value against a single property/schema descriptor. `path` is the dotted location used
|
||||
* in error messages ("" at the top level, "field a" / "field a.b" / "item[2]" beneath it). */
|
||||
function validateProperty(value: unknown, prop: JsonSchemaProperty, path: string): void {
|
||||
if (prop.type) checkType(value, prop.type, path);
|
||||
if (prop.enum && !prop.enum.includes(value)) {
|
||||
throw new Error(`${path || "value"}: value not in allowed enum`);
|
||||
}
|
||||
if (prop.type === "object" && typeof value === "object" && value !== null && !Array.isArray(value)) {
|
||||
const obj = value as Record<string, unknown>;
|
||||
for (const key of prop.required ?? []) {
|
||||
if (!(key in obj)) throw new Error(`${path ? path + ": " : ""}missing required field: ${key}`);
|
||||
}
|
||||
if (prop.properties) {
|
||||
for (const [key, sub] of Object.entries(prop.properties)) {
|
||||
if (key in obj) validateProperty(obj[key], sub, path ? `${path}.${key}` : `field ${key}`);
|
||||
}
|
||||
}
|
||||
if (prop.additionalProperties === false && prop.properties) {
|
||||
for (const key of Object.keys(obj)) {
|
||||
if (!(key in prop.properties)) throw new Error(`${path ? path + ": " : ""}unexpected additional property: ${key}`);
|
||||
}
|
||||
}
|
||||
}
|
||||
if (prop.type === "array" && Array.isArray(value) && prop.items) {
|
||||
value.forEach((item, i) => validateProperty(item, prop.items!, path ? `${path}[${i}]` : `item[${i}]`));
|
||||
}
|
||||
}
|
||||
|
||||
function checkType(v: unknown, type: string, field: string): void {
|
||||
const jsType = Array.isArray(v) ? "array" : v === null ? "null" : typeof v;
|
||||
// integer: must be a number AND a whole number. number: any number (incl. floats).
|
||||
if (type === "integer") {
|
||||
if (typeof v !== "number" || !Number.isInteger(v)) throw new Error(`field ${field}: expected integer`);
|
||||
return;
|
||||
}
|
||||
if (type === "number") {
|
||||
if (typeof v !== "number") throw new Error(`field ${field}: expected number`);
|
||||
return;
|
||||
}
|
||||
if (jsType !== type) throw new Error(`field ${field}: expected ${type}, got ${jsType}`);
|
||||
}
|
||||
|
||||
interface AgentOpts {
|
||||
label?: string;
|
||||
schema?: JsonSchema;
|
||||
toolNames?: string[];
|
||||
systemPromptAddendum?: string;
|
||||
cwd?: string;
|
||||
agentType?: string;
|
||||
/** Opt the agent into running inside a throwaway git worktree (detached at HEAD), so its file
|
||||
* writes can't collide with the main repo or with other concurrently-running agents. Use this for
|
||||
* writable agents (general-purpose, debugger, test-writer, or plugin agents) that you fan out in
|
||||
* parallel — their edits are discarded and only the returned answer matters. If the parent cwd
|
||||
* isn't a git repo, isolation is best-effort and the agent runs in the shared cwd. Don't use it
|
||||
* for read-only agents (explore, code-reviewer, planner) — they should see current uncommitted
|
||||
* state, so run them in the shared cwd. */
|
||||
isolation?: "worktree";
|
||||
}
|
||||
|
||||
/** Builds the SubAgentOverrides for an agent() call, resolving agentType to its toolset+addendum
|
||||
* unless the script explicitly supplies either. */
|
||||
function buildOverrides(opts: AgentOpts): SubAgentOverrides | undefined {
|
||||
if (opts.toolNames || opts.systemPromptAddendum) {
|
||||
return { toolNames: opts.toolNames, systemPromptAddendum: opts.systemPromptAddendum, cwd: opts.cwd };
|
||||
}
|
||||
if (opts.agentType && opts.agentType !== "general-purpose") {
|
||||
const spec = getAgentType(opts.agentType);
|
||||
return { toolNames: spec.toolNames, systemPromptAddendum: spec.systemPromptAddendum, cwd: opts.cwd };
|
||||
}
|
||||
return opts.cwd ? { cwd: opts.cwd } : undefined;
|
||||
}
|
||||
|
||||
/** Runs a workflow script with the orchestration globals in scope. Returns whatever the script
|
||||
* returns (or throws). agent/parallel/pipeline run sub-agents via ctx.runSubAgent, bounded by the
|
||||
* concurrency cap and the overall timeout. */
|
||||
async function runWorkflow(
|
||||
script: string,
|
||||
args: unknown,
|
||||
ctx: {
|
||||
cwd?: string;
|
||||
runSubAgent?: (task: { description: string; prompt: string }, overrides?: SubAgentOverrides) => Promise<SubAgentResult>;
|
||||
emitNotice?: (text: string, isError?: boolean) => void;
|
||||
},
|
||||
): Promise<unknown> {
|
||||
if (!ctx.runSubAgent) throw new Error("Sub-agents are not available in this context.");
|
||||
|
||||
const concurrency = workflowConcurrency();
|
||||
|
||||
// A bounded concurrency runner: at most `limit` thunks in flight at once. Used by parallel() and
|
||||
// pipeline() so a fan-out over many items doesn't dump a huge burst onto a local model server.
|
||||
async function runBounded<I, O>(limit: number, items: I[], fn: (item: I, index: number) => Promise<O>): Promise<O[]> {
|
||||
const results: O[] = new Array(items.length);
|
||||
let next = 0;
|
||||
async function worker(): Promise<void> {
|
||||
while (true) {
|
||||
const i = next++;
|
||||
if (i >= items.length) return;
|
||||
results[i] = await fn(items[i]!, i);
|
||||
}
|
||||
}
|
||||
const workers = Array.from({ length: Math.min(limit, items.length) }, () => worker());
|
||||
await Promise.all(workers);
|
||||
return results;
|
||||
}
|
||||
|
||||
async function agentFn(prompt: string, opts: AgentOpts = {}): Promise<unknown> {
|
||||
const description = opts.label ?? "workflow-agent";
|
||||
// Opt-in worktree isolation for writable parallel agents (see AgentOpts.isolation). Created here
|
||||
// and cleaned up in the finally below — runSubAgentTurn treats a cwd != parent.cwd as "isolated"
|
||||
// (edits discarded, not resumable) and appends the discard-notice itself. Best-effort: if the
|
||||
// parent cwd isn't a git repo, createWorktree returns cwd:undefined and we run in the shared cwd.
|
||||
let isolation: WorktreeIsolation | undefined;
|
||||
if (opts.isolation === "worktree") isolation = await createWorktree(ctx.cwd ?? ".");
|
||||
const worktreeCwd = isolation?.cwd;
|
||||
try {
|
||||
const overrides = buildOverrides({ ...opts, cwd: opts.cwd ?? worktreeCwd });
|
||||
const structuredAddendum = opts.schema
|
||||
? `\n\nReturn ONLY a JSON ${opts.schema.type ?? "object"} matching this schema (no prose, no code fences):\n${JSON.stringify(opts.schema)}`
|
||||
: undefined;
|
||||
const callOverrides = structuredAddendum
|
||||
? {
|
||||
...overrides,
|
||||
systemPromptAddendum: overrides?.systemPromptAddendum
|
||||
? `${overrides.systemPromptAddendum}\n${structuredAddendum}`
|
||||
: structuredAddendum,
|
||||
}
|
||||
: overrides;
|
||||
|
||||
const r1 = await ctx.runSubAgent!({ description, prompt }, callOverrides);
|
||||
if (!opts.schema) return r1.result;
|
||||
// Structured output: parse + validate, retry once with a targeted nudge if it fails. The nudge
|
||||
// includes the actual validation error so the sub-agent can correct the specific defect.
|
||||
let parsed: unknown;
|
||||
try {
|
||||
parsed = extractJson(r1.result);
|
||||
validateAgainstSchema(parsed, opts.schema);
|
||||
return parsed;
|
||||
} catch (err) {
|
||||
const why = (err as Error).message;
|
||||
const required = opts.schema.required ?? Object.keys(opts.schema.properties ?? {});
|
||||
const nudge = `${prompt}\n\nYour previous response was not valid JSON matching the schema (${why}). Return ONLY a JSON object with these fields: ${JSON.stringify(required)}. No prose, no code fences, no trailing commas.`;
|
||||
const r2 = await ctx.runSubAgent!({ description, prompt: nudge }, callOverrides);
|
||||
parsed = extractJson(r2.result);
|
||||
validateAgainstSchema(parsed, opts.schema);
|
||||
return parsed;
|
||||
}
|
||||
} finally {
|
||||
await isolation?.cleanup();
|
||||
}
|
||||
}
|
||||
|
||||
function parallelFn<T>(thunks: Array<() => Promise<T>>): Promise<Array<T | null>> {
|
||||
if (!Array.isArray(thunks)) throw new Error("parallel() expects an array of thunks");
|
||||
return runBounded<() => Promise<T>, T | null>(concurrency, thunks, async (thunk) => {
|
||||
try {
|
||||
return await thunk();
|
||||
} catch (err) {
|
||||
// A thunk that throws (or whose agent errors) resolves to null — the call itself never
|
||||
// rejects, so one failing branch doesn't abort the whole parallel batch.
|
||||
ctx.emitNotice?.(`parallel branch failed: ${(err as Error).message ?? String(err)}`, true);
|
||||
return null;
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
function pipelineFn<T, R>(items: T[], ...stages: Array<(prev: unknown, original: T, index: number) => Promise<unknown>>): Promise<Array<R | null>> {
|
||||
if (!Array.isArray(items)) throw new Error("pipeline() expects an array of items");
|
||||
if (stages.length === 0) throw new Error("pipeline() needs at least one stage");
|
||||
return runBounded<T, R | null>(concurrency, items, async (item, index) => {
|
||||
try {
|
||||
let val: unknown = item;
|
||||
for (const stage of stages) {
|
||||
val = await stage(val, item, index);
|
||||
}
|
||||
return val as R;
|
||||
} catch (err) {
|
||||
// A stage that throws drops just this item to null (skipping its remaining stages).
|
||||
ctx.emitNotice?.(`pipeline item ${index} failed: ${(err as Error).message ?? String(err)}`, true);
|
||||
return null;
|
||||
}
|
||||
});
|
||||
}
|
||||
|
||||
function phaseFn(title: string): void {
|
||||
ctx.emitNotice?.(`▶ ${title}`);
|
||||
}
|
||||
|
||||
function logFn(message: string): void {
|
||||
ctx.emitNotice?.(String(message));
|
||||
}
|
||||
|
||||
const sandbox = {
|
||||
agent: agentFn,
|
||||
parallel: parallelFn,
|
||||
pipeline: pipelineFn,
|
||||
phase: phaseFn,
|
||||
log: logFn,
|
||||
args,
|
||||
// Nested workflows are one level only (matches Claude Code); there's no child context to run in.
|
||||
workflow: () => {
|
||||
throw new Error("workflow() cannot be nested.");
|
||||
},
|
||||
};
|
||||
|
||||
const wrapped = `(async () => {\n${stripExports(script)}\n})()`;
|
||||
const context = vm.createContext(sandbox);
|
||||
let timeoutId: ReturnType<typeof setTimeout>;
|
||||
const timeout = new Promise<never>((_, reject) => {
|
||||
timeoutId = setTimeout(
|
||||
() => reject(new Error(`Workflow timed out after ${Math.round(workflowTimeoutMs() / 1000)}s.`)),
|
||||
workflowTimeoutMs(),
|
||||
);
|
||||
});
|
||||
try {
|
||||
const promise = vm.runInContext(wrapped, context, { filename: "workflow.js" }) as Promise<unknown>;
|
||||
return await Promise.race([promise, timeout]);
|
||||
} finally {
|
||||
clearTimeout(timeoutId!);
|
||||
}
|
||||
}
|
||||
|
||||
export const workflowTool: ToolDef<z.infer<typeof schema>> = {
|
||||
name: "workflow",
|
||||
description:
|
||||
"Run a multi-agent workflow from a self-contained JavaScript script that deterministically orchestrates sub-agents — " +
|
||||
"for being comprehensive (decompose and cover in parallel), confident (independent perspectives + adversarial checks " +
|
||||
"before committing), or scaling work one context can't hold (migrations, audits). Use it only when the user asks for " +
|
||||
"multi-agent orchestration (e.g. 'use a workflow', 'fan out agents'); a single agent() call inside is just a delegation. " +
|
||||
"The script globals are agent, parallel, pipeline, phase, log, args. agent(prompt, opts) accepts opts.isolation: 'worktree' " +
|
||||
"to run a writable agent in a throwaway git worktree so parallel writers don't collide (edits discarded; only the answer " +
|
||||
"matters) — use it for parallel general-purpose/debugger/test-writer agents, not for read-only explore/code-reviewer ones. " +
|
||||
"Runs headless; the return value is the result.",
|
||||
schema,
|
||||
mutating: false,
|
||||
handler: async (args, ctx) => {
|
||||
const result = await runWorkflow(args.script, args.args, ctx);
|
||||
return { result };
|
||||
},
|
||||
};
|
||||
|
||||
// Exported for tests.
|
||||
export { runWorkflow, extractJson, validateAgainstSchema, stripExports };
|
||||
@@ -0,0 +1,174 @@
|
||||
import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
|
||||
import { execa } from "execa";
|
||||
import { existsSync, mkdtempSync, rmSync, writeFileSync } from "node:fs";
|
||||
import path from "node:path";
|
||||
import os from "node:os";
|
||||
import { enterWorktreeTool, exitWorktreeTool } from "./worktreeSession.js";
|
||||
import { removeInteractiveWorktree } from "../utils/worktree.js";
|
||||
import type { ToolContext } from "./types.js";
|
||||
|
||||
// These tests shell out to real git for the create/remove paths, and use a mocked ctx to observe
|
||||
// the setCwd/setWorktree/getWorktree wiring without a full Session.
|
||||
|
||||
async function gitInit(repo: string): Promise<void> {
|
||||
await execa("git", ["init", "-q"], { cwd: repo });
|
||||
await execa("git", ["config", "user.email", "t@t"], { cwd: repo });
|
||||
await execa("git", ["config", "user.name", "t"], { cwd: repo });
|
||||
writeFileSync(path.join(repo, "README.md"), "hello\n");
|
||||
await execa("git", ["add", "."], { cwd: repo });
|
||||
await execa("git", ["commit", "-q", "-m", "init"], { cwd: repo });
|
||||
}
|
||||
|
||||
// A mock ctx that records setCwd calls and tracks the worktree over a shared object, mirroring how
|
||||
// gateAndRun wires these over the Session. `cwd` is the current working dir (the repo, or the
|
||||
// worktree once switched).
|
||||
function ctxFor(cwd: string): { ctx: ToolContext; state: { cwd: string; worktree?: { dir: string; branch: string; originalCwd: string } } } {
|
||||
const state: { cwd: string; worktree?: { dir: string; branch: string; originalCwd: string } } = { cwd };
|
||||
const ctx: ToolContext = {
|
||||
cwd,
|
||||
setCwd: (newCwd) => {
|
||||
state.cwd = newCwd;
|
||||
},
|
||||
getWorktree: () => state.worktree,
|
||||
setWorktree: (wt) => {
|
||||
state.worktree = wt;
|
||||
},
|
||||
};
|
||||
return { ctx, state };
|
||||
}
|
||||
|
||||
describe("enter_worktree / exit_worktree tool wiring", () => {
|
||||
let repo: string;
|
||||
|
||||
beforeEach(() => {
|
||||
repo = mkdtempSync(path.join(os.tmpdir(), "locode-wt-tool-test-"));
|
||||
});
|
||||
|
||||
afterEach(async () => {
|
||||
if (existsSync(repo)) {
|
||||
await execa("git", ["worktree", "prune"], { cwd: repo, reject: false }).catch(() => {});
|
||||
try {
|
||||
rmSync(repo, { recursive: true, force: true });
|
||||
} catch {
|
||||
/* leave for the OS temp sweep */
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
it("enter_worktree throws when the ctx doesn't support worktree sessions", async () => {
|
||||
const ctx: ToolContext = { cwd: repo };
|
||||
await expect(enterWorktreeTool.handler({}, ctx)).rejects.toThrow(/not available/i);
|
||||
});
|
||||
|
||||
it("enter_worktree refuses when already in a worktree session", async () => {
|
||||
await gitInit(repo);
|
||||
const { ctx, state } = ctxFor(repo);
|
||||
state.worktree = { dir: "/tmp/prev", branch: "locode-wt-prev", originalCwd: repo };
|
||||
const result = (await enterWorktreeTool.handler({ name: "second" }, ctx)) as { error: string };
|
||||
expect(result.error).toMatch(/already in a worktree/i);
|
||||
});
|
||||
|
||||
it("enter_worktree returns an error (without switching) for a non-git cwd", async () => {
|
||||
// repo dir exists but no git init.
|
||||
const { ctx, state } = ctxFor(repo);
|
||||
const result = (await enterWorktreeTool.handler({ name: "x" }, ctx)) as { error: string };
|
||||
expect(result.error).toMatch(/Not a git repository/i);
|
||||
expect(state.worktree).toBeUndefined();
|
||||
expect(state.cwd).toBe(repo);
|
||||
});
|
||||
|
||||
it("enter_worktree creates a worktree, switches cwd, and records the worktree", async () => {
|
||||
await gitInit(repo);
|
||||
const { ctx, state } = ctxFor(repo);
|
||||
const result = (await enterWorktreeTool.handler({ name: "feature" }, ctx)) as {
|
||||
dir: string;
|
||||
branch: string;
|
||||
message: string;
|
||||
};
|
||||
expect(result.branch).toBe("locode-wt-feature");
|
||||
expect(existsSync(result.dir)).toBe(true);
|
||||
// The tool switched the cwd into the worktree and recorded the tracking (with the original cwd).
|
||||
expect(state.cwd).toBe(result.dir);
|
||||
expect(state.worktree).toEqual({ dir: result.dir, branch: "locode-wt-feature", originalCwd: repo });
|
||||
// Clean up the created worktree+branch so afterEach's repo removal is clean.
|
||||
await removeInteractiveWorktree(repo, result.dir, result.branch);
|
||||
});
|
||||
|
||||
it("exit_worktree returns an error when not in a worktree session", async () => {
|
||||
const { ctx } = ctxFor(repo);
|
||||
const result = (await exitWorktreeTool.handler({ action: "keep" }, ctx)) as { error: string };
|
||||
expect(result.error).toMatch(/not in a worktree/i);
|
||||
});
|
||||
|
||||
it("exit_worktree(keep) restores the original cwd and clears the worktree tracking", async () => {
|
||||
await gitInit(repo);
|
||||
const { ctx, state } = ctxFor(repo);
|
||||
const enter = (await enterWorktreeTool.handler({ name: "keepme" }, ctx)) as { dir: string; branch: string };
|
||||
expect(state.cwd).toBe(enter.dir);
|
||||
const result = (await exitWorktreeTool.handler({ action: "keep" }, ctx)) as {
|
||||
action: string;
|
||||
restoredCwd: string;
|
||||
};
|
||||
expect(result.action).toBe("keep");
|
||||
expect(result.restoredCwd).toBe(repo);
|
||||
expect(state.cwd).toBe(repo);
|
||||
expect(state.worktree).toBeUndefined();
|
||||
// keep leaves the worktree dir + branch in place.
|
||||
expect(existsSync(enter.dir)).toBe(true);
|
||||
// Clean up the kept worktree so afterEach can remove the repo.
|
||||
await removeInteractiveWorktree(repo, enter.dir, enter.branch);
|
||||
});
|
||||
|
||||
it("exit_worktree(remove) refuses a dirty worktree without discardChanges (no removal, no restore)", async () => {
|
||||
await gitInit(repo);
|
||||
const { ctx, state } = ctxFor(repo);
|
||||
const enter = (await enterWorktreeTool.handler({ name: "dirty" }, ctx)) as { dir: string; branch: string };
|
||||
writeFileSync(path.join(enter.dir, "README.md"), "changed\n"); // uncommitted change
|
||||
const result = (await exitWorktreeTool.handler({ action: "remove" }, ctx)) as { error: string };
|
||||
expect(result.error).toMatch(/uncommitted changes/i);
|
||||
// Refusal leaves everything in place: still in the worktree, worktree still tracked, dir still exists.
|
||||
expect(state.cwd).toBe(enter.dir);
|
||||
expect(state.worktree).toBeDefined();
|
||||
expect(existsSync(enter.dir)).toBe(true);
|
||||
// Now opt in to discard → removes + restores.
|
||||
const ok = (await exitWorktreeTool.handler({ action: "remove", discardChanges: true }, ctx)) as {
|
||||
action: string;
|
||||
restoredCwd: string;
|
||||
};
|
||||
expect(ok.action).toBe("remove");
|
||||
expect(ok.restoredCwd).toBe(repo);
|
||||
expect(state.cwd).toBe(repo);
|
||||
expect(state.worktree).toBeUndefined();
|
||||
expect(existsSync(enter.dir)).toBe(false);
|
||||
});
|
||||
|
||||
it("exit_worktree(remove) on a clean worktree removes + restores without discardChanges", async () => {
|
||||
await gitInit(repo);
|
||||
const { ctx, state } = ctxFor(repo);
|
||||
const enter = (await enterWorktreeTool.handler({ name: "clean" }, ctx)) as { dir: string; branch: string };
|
||||
const result = (await exitWorktreeTool.handler({ action: "remove" }, ctx)) as {
|
||||
action: string;
|
||||
restoredCwd: string;
|
||||
};
|
||||
expect(result.action).toBe("remove");
|
||||
expect(state.cwd).toBe(repo);
|
||||
expect(state.worktree).toBeUndefined();
|
||||
expect(existsSync(enter.dir)).toBe(false);
|
||||
});
|
||||
});
|
||||
|
||||
describe("worktree tool flags + schema", () => {
|
||||
it("enter_worktree is non-mutating; exit_worktree is mutating (destructive on remove)", () => {
|
||||
expect(enterWorktreeTool.mutating).toBe(false);
|
||||
expect(exitWorktreeTool.mutating).toBe(true);
|
||||
});
|
||||
it("exit_worktree action is required and must be keep|remove", () => {
|
||||
expect(() => exitWorktreeTool.schema.parse({})).toThrow();
|
||||
expect(() => exitWorktreeTool.schema.parse({ action: "other" })).toThrow();
|
||||
expect(() => exitWorktreeTool.schema.parse({ action: "keep" })).not.toThrow();
|
||||
});
|
||||
it("enter_worktree name is optional", () => {
|
||||
expect(() => enterWorktreeTool.schema.parse({})).not.toThrow();
|
||||
expect(() => enterWorktreeTool.schema.parse({ name: "feature" })).not.toThrow();
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,135 @@
|
||||
import { z } from "zod";
|
||||
import type { ToolDef } from "./types.js";
|
||||
import {
|
||||
createInteractiveWorktree,
|
||||
hasUncommittedChanges,
|
||||
removeInteractiveWorktree,
|
||||
} from "../utils/worktree.js";
|
||||
|
||||
const enterSchema = z.object({
|
||||
name: z
|
||||
.string()
|
||||
.optional()
|
||||
.describe(
|
||||
"Optional name for the worktree and its branch (locode-wt-<name>). Letters, digits, dot, " +
|
||||
"underscore, dash, starting alphanumeric, max 64 chars. Omit to auto-generate. Must be unique — " +
|
||||
"re-using an existing branch name returns an error.",
|
||||
),
|
||||
});
|
||||
|
||||
const exitSchema = z.object({
|
||||
action: z
|
||||
.enum(["keep", "remove"])
|
||||
.describe(
|
||||
"'keep' leaves the worktree directory and branch in place (restored to the original working " +
|
||||
"directory; the branch is preserved in git). 'remove' deletes the worktree directory AND the " +
|
||||
"branch (irreversible).",
|
||||
),
|
||||
discardChanges: z
|
||||
.boolean()
|
||||
.optional()
|
||||
.describe(
|
||||
"Only matters for action 'remove'. If the worktree has uncommitted changes, removal is refused " +
|
||||
"unless this is true (the changes are then discarded along with the worktree and branch).",
|
||||
),
|
||||
});
|
||||
|
||||
/** Enters an interactive git worktree session: creates a new branch `locode-wt-<name>` at HEAD in a
|
||||
* throwaway directory, switches the session's working directory into it, and remembers the original
|
||||
* cwd so `exit_worktree` can restore it. While in the worktree, all file tools operate there (the
|
||||
* workspace root becomes the worktree), so experiments can't touch the user's uncommitted work in
|
||||
* the main repo. Non-mutating: it creates an isolated copy, not an edit to user files. Refuses if
|
||||
* already in a worktree (exit first) or the cwd isn't a git repo. */
|
||||
export const enterWorktreeTool: ToolDef<z.infer<typeof enterSchema>> = {
|
||||
name: "enter_worktree",
|
||||
description:
|
||||
"Create an isolated git worktree on a new branch (locode-wt-<name>) at the current HEAD and switch the " +
|
||||
"session's working directory into it. Use it to try changes without touching the main working tree — " +
|
||||
"the worktree starts from the last commit, so uncommitted changes in the main repo don't carry over. " +
|
||||
"While inside, every file tool (read_file, edit_file, grep, bash, etc.) operates in the worktree. Leave " +
|
||||
"with exit_worktree (keep preserves the branch; remove discards it). Refuses if already in a worktree " +
|
||||
"(exit first) or the cwd isn't a git repo. Only the main session should use this (not sub-agents).",
|
||||
schema: enterSchema,
|
||||
mutating: false,
|
||||
handler: async (args, ctx) => {
|
||||
if (!ctx.setCwd || !ctx.setWorktree || !ctx.getWorktree) {
|
||||
throw new Error("Worktree sessions are not available in this context.");
|
||||
}
|
||||
if (ctx.getWorktree()) {
|
||||
return { error: "Already in a worktree session. Use exit_worktree (keep or remove) before entering another." };
|
||||
}
|
||||
const parentCwd = ctx.cwd;
|
||||
try {
|
||||
const wt = await createInteractiveWorktree(parentCwd, args.name);
|
||||
ctx.setWorktree({ dir: wt.dir, branch: wt.branch, originalCwd: parentCwd });
|
||||
ctx.setCwd(wt.dir);
|
||||
return {
|
||||
dir: wt.dir,
|
||||
branch: wt.branch,
|
||||
message:
|
||||
`Switched into worktree at ${wt.dir} on branch ${wt.branch}. File tools now operate there. ` +
|
||||
`Use exit_worktree to leave (keep the branch, or remove it).`,
|
||||
};
|
||||
} catch (err) {
|
||||
return { error: (err as Error).message };
|
||||
}
|
||||
},
|
||||
};
|
||||
|
||||
/** Leaves the active worktree session, restoring the session's working directory to the original cwd.
|
||||
* `action: "keep"` leaves the worktree dir + branch in place (the branch persists in git; the temp dir
|
||||
* remains until the process/OS reclaims it). `action: "remove"` deletes the worktree dir AND the
|
||||
* branch — refused if the worktree has uncommitted changes unless `discardChanges: true`. Mutating
|
||||
* (a remove is destructive), so the user is asked to confirm. Refuses if not in a worktree session. */
|
||||
export const exitWorktreeTool: ToolDef<z.infer<typeof exitSchema>> = {
|
||||
name: "exit_worktree",
|
||||
description:
|
||||
"Leave the active worktree session, restoring the working directory to where it was before enter_worktree. " +
|
||||
"action 'keep' preserves the worktree directory and branch (the branch stays in git — recover the work via " +
|
||||
"git checkout/worktree add later); action 'remove' deletes the worktree directory AND the branch. A remove " +
|
||||
"is refused when the worktree has uncommitted changes unless discardChanges is true. Use this only after " +
|
||||
"enter_worktree; it returns an error if you're not in a worktree session.",
|
||||
schema: exitSchema,
|
||||
mutating: true,
|
||||
handler: async (args, ctx) => {
|
||||
if (!ctx.setCwd || !ctx.setWorktree || !ctx.getWorktree) {
|
||||
throw new Error("Worktree sessions are not available in this context.");
|
||||
}
|
||||
const wt = ctx.getWorktree();
|
||||
if (!wt) {
|
||||
return { error: "Not in a worktree session — nothing to exit." };
|
||||
}
|
||||
if (args.action === "remove") {
|
||||
// Refuse to silently destroy uncommitted work; the model must opt in via discardChanges.
|
||||
let dirty = false;
|
||||
try {
|
||||
dirty = await hasUncommittedChanges(wt.dir);
|
||||
} catch {
|
||||
dirty = false;
|
||||
}
|
||||
if (dirty && !args.discardChanges) {
|
||||
return {
|
||||
error:
|
||||
`Worktree "${wt.branch}" has uncommitted changes. Re-run with discardChanges: true to discard them ` +
|
||||
`along with the worktree and branch, or use action: "keep" to preserve them.`,
|
||||
};
|
||||
}
|
||||
try {
|
||||
await removeInteractiveWorktree(wt.originalCwd, wt.dir, wt.branch);
|
||||
} catch (err) {
|
||||
return { error: `Failed to remove worktree: ${(err as Error).message}` };
|
||||
}
|
||||
}
|
||||
// Restore the session cwd for both keep and remove, and clear the tracking.
|
||||
ctx.setCwd(wt.originalCwd);
|
||||
ctx.setWorktree(undefined);
|
||||
return {
|
||||
action: args.action,
|
||||
restoredCwd: wt.originalCwd,
|
||||
message:
|
||||
args.action === "keep"
|
||||
? `Left worktree ${wt.dir} in place (branch ${wt.branch} is preserved in git). Restored working directory to ${wt.originalCwd}.`
|
||||
: `Removed worktree ${wt.dir} and deleted branch ${wt.branch}. Restored working directory to ${wt.originalCwd}.`,
|
||||
};
|
||||
},
|
||||
};
|
||||
+29
-38
@@ -1,45 +1,36 @@
|
||||
import { mkdirSync, mkdtempSync, readFileSync, rmSync } from "node:fs";
|
||||
import os from "node:os";
|
||||
import { mkdtemp, readFile, rm } from "node:fs/promises";
|
||||
import { tmpdir } from "node:os";
|
||||
import path from "node:path";
|
||||
import { afterEach, beforeEach, describe, expect, it } from "vitest";
|
||||
import { describe, expect, it } from "vitest";
|
||||
import { writeFileTool } from "./writeFile.js";
|
||||
import type { ToolContext } from "./types.js";
|
||||
|
||||
describe("writeFile tool — path containment", () => {
|
||||
let cwd: string;
|
||||
let ctx: ToolContext;
|
||||
|
||||
beforeEach(() => {
|
||||
cwd = mkdtempSync(path.join(os.tmpdir(), "locode-writefile-"));
|
||||
ctx = { cwd };
|
||||
describe("writeFileTool", () => {
|
||||
it("rejects paths that escape the working directory", async () => {
|
||||
const cwd = await mkdtemp(path.join(tmpdir(), "locode-writefile-"));
|
||||
try {
|
||||
await expect(
|
||||
writeFileTool.handler({ path: "../outside.txt", content: "x" }, { cwd }),
|
||||
).rejects.toThrow("Path resolves outside the working directory");
|
||||
await expect(
|
||||
writeFileTool.handler({ path: "sub/../../outside.txt", content: "x" }, { cwd }),
|
||||
).rejects.toThrow("Path resolves outside the working directory");
|
||||
await expect(
|
||||
writeFileTool.handler({ path: "/etc/passwd", content: "x" }, { cwd }),
|
||||
).rejects.toThrow("Path resolves outside the working directory");
|
||||
} finally {
|
||||
await rm(cwd, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
|
||||
afterEach(() => {
|
||||
rmSync(cwd, { recursive: true, force: true });
|
||||
});
|
||||
|
||||
it("writes a file inside the working directory", async () => {
|
||||
const result = (await writeFileTool.handler({ path: "note.txt", content: "hi" }, ctx)) as { path: string };
|
||||
expect(readFileSync(result.path, "utf-8")).toBe("hi");
|
||||
});
|
||||
|
||||
it("refuses to write outside the working directory via ../ traversal", async () => {
|
||||
await expect(writeFileTool.handler({ path: "../escape.txt", content: "oops" }, ctx)).rejects.toThrow(/outside the working directory/);
|
||||
});
|
||||
|
||||
it("refuses to write to an absolute path outside the working directory", async () => {
|
||||
const outside = path.join(os.tmpdir(), "locode-outside-target.txt");
|
||||
await expect(writeFileTool.handler({ path: outside, content: "oops" }, ctx)).rejects.toThrow(/outside the working directory/);
|
||||
});
|
||||
|
||||
it("preview reports the block instead of showing a diff", async () => {
|
||||
const preview = await writeFileTool.preview!({ path: "../escape.txt", content: "oops" }, ctx);
|
||||
expect(preview).toMatch(/outside the working directory/);
|
||||
});
|
||||
|
||||
it("still applies even when the escaping subdirectory already exists", async () => {
|
||||
// Sanity check that the guard runs before mkdir/write, not after.
|
||||
mkdirSync(path.join(cwd, "sub"), { recursive: true });
|
||||
await expect(writeFileTool.handler({ path: "sub/../../escape.txt", content: "oops" }, ctx)).rejects.toThrow(/outside the working directory/);
|
||||
it("writes files inside the working directory", async () => {
|
||||
const cwd = await mkdtemp(path.join(tmpdir(), "locode-writefile-"));
|
||||
try {
|
||||
const result = (await writeFileTool.handler({ path: "nested/file.txt", content: "hello" }, { cwd })) as { path: string };
|
||||
expect(result.path).toBe(path.join(cwd, "nested/file.txt"));
|
||||
const content = await readFile(path.join(cwd, "nested/file.txt"), "utf-8");
|
||||
expect(content).toBe("hello");
|
||||
} finally {
|
||||
await rm(cwd, { recursive: true, force: true });
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
+16
-22
@@ -1,9 +1,8 @@
|
||||
import { createPatch } from "diff";
|
||||
import { randomUUID } from "node:crypto";
|
||||
import { mkdir, readFile as fsReadFile, rename, unlink, writeFile as fsWriteFile } from "node:fs/promises";
|
||||
import { mkdir, readFile as fsReadFile, writeFile as fsWriteFile } from "node:fs/promises";
|
||||
import path from "node:path";
|
||||
import { z } from "zod";
|
||||
import { resolveWithinCwd } from "./pathGuard.js";
|
||||
import { assertWithinWorkspace } from "../utils/path.js";
|
||||
import type { ToolDef } from "./types.js";
|
||||
|
||||
const schema = z.object({
|
||||
@@ -14,24 +13,21 @@ const schema = z.object({
|
||||
async function readExisting(resolved: string): Promise<string | null> {
|
||||
try {
|
||||
return await fsReadFile(resolved, "utf-8");
|
||||
} catch (err: any) {
|
||||
if (err?.code === "ENOENT") return null;
|
||||
throw err;
|
||||
} catch {
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
export const writeFileTool: ToolDef<z.infer<typeof schema>> = {
|
||||
name: "write_file",
|
||||
description: "Create or overwrite a file with the given content. Use for new files or full rewrites. For small changes to an existing file, prefer edit_file instead.",
|
||||
description:
|
||||
"Create or overwrite a file with the given content. Use this for new files or when rewriting " +
|
||||
"most of a file; prefer edit_file for small, targeted changes.",
|
||||
schema,
|
||||
mutating: true,
|
||||
preview: async ({ path: filePath, content }, ctx) => {
|
||||
let resolved: string;
|
||||
try {
|
||||
resolved = resolveWithinCwd(ctx.cwd, filePath);
|
||||
} catch (err) {
|
||||
return (err as Error).message;
|
||||
}
|
||||
const resolved = path.resolve(ctx.cwd, filePath);
|
||||
assertWithinWorkspace(resolved, ctx.cwd, filePath);
|
||||
const existing = await readExisting(resolved);
|
||||
if (existing === null) {
|
||||
return `Create new file ${resolved} (${content.length} chars)`;
|
||||
@@ -39,16 +35,14 @@ export const writeFileTool: ToolDef<z.infer<typeof schema>> = {
|
||||
return createPatch(resolved, existing, content, "", "");
|
||||
},
|
||||
handler: async ({ path: filePath, content }, ctx) => {
|
||||
const resolved = resolveWithinCwd(ctx.cwd, filePath);
|
||||
const resolved = path.resolve(ctx.cwd, filePath);
|
||||
assertWithinWorkspace(resolved, ctx.cwd, filePath);
|
||||
const previousContent = await readExisting(resolved);
|
||||
await mkdir(path.dirname(resolved), { recursive: true });
|
||||
const tmpPath = resolved + ".tmp-" + randomUUID();
|
||||
try {
|
||||
await fsWriteFile(tmpPath, content, "utf-8");
|
||||
await rename(tmpPath, resolved);
|
||||
} catch (err) {
|
||||
try { await unlink(tmpPath); } catch {}
|
||||
throw err;
|
||||
await fsWriteFile(resolved, content, "utf-8");
|
||||
if (ctx.setLastEdit) {
|
||||
ctx.setLastEdit({ path: filePath, previousContent: previousContent ?? "" });
|
||||
}
|
||||
return { path: resolved, bytesWritten: Buffer.byteLength(content, "utf-8") };
|
||||
},
|
||||
};
|
||||
};
|
||||
|
||||
+565
-465
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,54 @@
|
||||
import { describe, expect, it } from "vitest";
|
||||
import { getGraphemeBoundaries, nextGraphemeBoundary } from "./ChatInput.js";
|
||||
|
||||
describe("getGraphemeBoundaries", () => {
|
||||
it("returns [0, length] for an empty string", () => {
|
||||
expect(getGraphemeBoundaries("")).toEqual([0]);
|
||||
});
|
||||
|
||||
it("returns boundaries for plain ASCII", () => {
|
||||
expect(getGraphemeBoundaries("abc")).toEqual([0, 1, 2, 3]);
|
||||
});
|
||||
|
||||
it("treats surrogate pairs (emoji/Hangul) as single graphemes", () => {
|
||||
// "a👍b" — thumbs up is a surrogate pair (2 JS indices, 1 displayed cell).
|
||||
const boundaries = getGraphemeBoundaries("a👍b");
|
||||
expect(boundaries).toEqual([0, 1, 3, 4]);
|
||||
});
|
||||
|
||||
it("treats ZWJ emoji sequences as single graphemes", () => {
|
||||
// "👨👩👧👦" is a family emoji made of multiple code points joined with ZWJs.
|
||||
const str = "👨👩👧👦";
|
||||
const boundaries = getGraphemeBoundaries(str);
|
||||
expect(boundaries).toHaveLength(2);
|
||||
expect(boundaries).toContain(0);
|
||||
expect(boundaries).toContain(str.length);
|
||||
});
|
||||
});
|
||||
|
||||
describe("nextGraphemeBoundary", () => {
|
||||
it("moves right past a surrogate-pair emoji", () => {
|
||||
// "a👍b", cursor after "a" (index 1) should jump to index 3 (after emoji).
|
||||
expect(nextGraphemeBoundary("a👍b", 1, 1)).toBe(3);
|
||||
});
|
||||
|
||||
it("moves left past a surrogate-pair emoji", () => {
|
||||
// cursor at index 3 (after emoji) should jump back to index 1 (before emoji).
|
||||
expect(nextGraphemeBoundary("a👍b", 3, -1)).toBe(1);
|
||||
});
|
||||
|
||||
it("does not move past the start or end", () => {
|
||||
expect(nextGraphemeBoundary("ab", 0, -1)).toBe(0);
|
||||
expect(nextGraphemeBoundary("ab", 2, 1)).toBe(2);
|
||||
});
|
||||
|
||||
it("snaps an invalid offset to the next boundary when moving right", () => {
|
||||
// index 2 is inside the emoji surrogate pair.
|
||||
expect(nextGraphemeBoundary("a👍b", 2, 1)).toBe(3);
|
||||
});
|
||||
|
||||
it("snaps an invalid offset to the previous boundary when moving left", () => {
|
||||
// index 2 is inside the emoji surrogate pair.
|
||||
expect(nextGraphemeBoundary("a👍b", 2, -1)).toBe(1);
|
||||
});
|
||||
});
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user