Compare commits

1 Commits
Author SHA1 Message Date
kimandClaude f347c150a4 v0.6.0: sub-agent orchestration, plan mode, memory, and the Claude-Code feature gap
Closes the gap to Claude Code across four upgrade rounds (all verified
green: typecheck, 388 tests, build 313 KB).

Sub-agent orchestration:
- Resumable sub-agents (send_message) + MaxIterations "continue" path
- Configurable multi-depth sub-agent nesting
- Parallel fan-out for pure agent/agent__* batches, each writable delegation
  in its own throwaway git worktree (no file collisions)
- Named specialist fleet: explore, code-reviewer, planner, debugger,
  test-writer (read-only types skip worktree isolation)
- Named teammates: agent tool `name` arg + list_teammates + send_message
  by name (synchronous, addressable handle — not true async background)

Plan mode:
- True in-turn pause/resume on approval (no synthetic proceed turn)
- exit_plan_mode tool + side-by-side diff rendering (/diff)

Persistent memory:
- Typed file-per-fact memory under config dir, index folded into the
  system prompt (bounded), memory + memory_write tools, /memory command

Toolset + UI:
- multi_edit, notebook_edit, web_search/web_fetch, workflow orchestration
  primitive (agent/parallel/pipeline/phase/log + worktree isolation +
  best-effort structured output with recursive schema validation)
- Structured task system (task_create/list/get/update with dependency
  graph + ownership) replacing flat todo_write
- Cron/scheduled tasks (cron_create/list/delete, schedule_wakeup) with
  durable persistence and idle-gated ticking
- Interactive worktree session (enter_worktree/exit_worktree) with a
  live-cwd UI switch that survives /model changes
- AskUserQuestion tool + QuestionPrompt UI
- Permission rule granularity + user/project settings merge
- Color diffs, sub-agent activity panel, /undo, dashboard refinements

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-16 11:35:05 +09:00
130 changed files with 8644 additions and 10379 deletions
-2095
View File
File diff suppressed because one or more lines are too long
+9 -17
View File
@@ -87,12 +87,11 @@ locode is a full-screen terminal app built with [Ink](https://github.com/vadimde
- **Ctrl+O**: print the full text of the last `/compact` (or auto-compact) summary. The collapsed notice you see right after compacting only shows a one-line hint — press Ctrl+O any time afterward to print the whole thing.
- **Ctrl+B**: while a `bash` command is running, detaches it into the background and returns control to you immediately — the turn continues with a `bash_output`-checkable job id instead of waiting for the command to finish. A notice appears in the transcript once the backgrounded command actually completes. Only `bash` supports this today. The model can kill a still-running backgrounded job with `bash_kill`; any jobs still running when locode itself exits are killed too, so they don't outlive the process as orphans.
- **Ctrl+F**: open and focus a file panel docked to the right of the chat (hidden by default); press again to close it. It has two tabs — **Files**, the project's collapsible file tree (directories in cyan, same `node_modules`/`.git`/`dist` exclusions as `@` mentions), and **Activity**, the files `read_file`/`write_file`/`edit_file` have touched so far this session, most recent first, with a status glyph (`·` read, `+` written, `~` edited) and a repeat count. **Ctrl+G** switches between the two tabs. While the panel is focused, `↑`/`↓` move the selection (auto-scrolling to keep it in view), `↵`/`←`/`→` expand or collapse the selected folder, and **Esc** hands keyboard focus back to the chat input without closing the panel — typing is disabled while the panel has focus, so the same arrow key doesn't simultaneously recall chat history.
- **Backends**: `--backend ollama` (default) or `--backend lmstudio`, or `--base-url <url>` for anything else that speaks the same API.
- **Tools**: `read_file`, `list_files`, `grep`, `definition`, `references`, `diagnostics`, `web_search`, `web_fetch`, `git_status`, `bash_output`, `todo_write`, `task_create`, `task_list`, `task_get`, `task_update` run automatically. `write_file`, `edit_file`, `multi_edit`, `bash`, `bash_kill`, and `git_commit` show a diff/preview in a bordered box and ask you to pick Yes / Yes-always-this-session / No with the arrow keys before running.
- **Tools**: `read_file`, `list_files`, `grep`, `web_search`, `web_fetch`, `git_status`, `bash_output`, `todo_write` run automatically. `write_file`, `edit_file`, `bash`, `bash_kill`, and `git_commit` show a diff/preview in a bordered box and ask you to pick Yes / Yes-always-this-session / No with the arrow keys before running.
- **Permission modes**: `default` (ask before every mutating tool), `plan` (research only — every mutating tool is blocked outright, no prompt; the model is expected to describe what it would do in its final answer instead), `auto-edit` (file edits auto-approved, `bash`/`git_commit` still ask), `auto-accept` (everything auto-approved — use with care). Cycle with `Shift+Tab` or set directly with `/perm <mode>`.
- **Structured tasks**: for multi-step work the model can call `task_create`/`task_list`/`task_get`/`task_update` to track units of work with a dependency graph (`blocks`/`blockedBy`), ownership (`owner`), status, and free-form metadata — created incrementally rather than replaced wholesale. The older flat `todo_write` live checklist (`☐`/`◐`/`☑`) remains for simpler cases.
- **Task checklists**: for multi-step work the model can call `todo_write` to show a live checklist (`☐`/`◐`/`☑`) in the transcript instead of silently working through a list you can't see progress on.
- **Project instructions**: a `CLAUDE.md` (or `AGENTS.md`) file in the project root is automatically read at session start and folded into the system prompt — put repo-specific conventions there and every session picks them up without being told.
- **Images**: `read_file` returns image files (png, jpg, jpeg, gif, webp, bmp — up to 5MB) as actual image content instead of trying to decode them as text, so vision-capable models can see them when the model itself calls the tool. To attach a file or image to your own message directly, use `/import <path> [caption]`.
- **`@` file mentions**: type `@` in the chat input to open a fuzzy file picker (searches the whole project, skipping `node_modules`/`.git`/`dist`) — keep typing to filter, `↑`/`↓` to navigate, `Tab` (or `Enter`) to insert the highlighted path. Any `@path` left in your message when you hit `Enter` for real is resolved against disk and attached to that message (text inlined, images attached as image content) — a stray `@` that isn't an actual file (e.g. an email address) is left as plain text.
@@ -101,12 +100,10 @@ locode is a full-screen terminal app built with [Ink](https://github.com/vadimde
- **MCP servers**: locode connects to any [MCP](https://modelcontextprotocol.io) servers configured via `locode mcp add` or a project's `.mcp.json` (stdio and remote/streamable-HTTP transports), and adds their tools to every session, namespaced as `mcp__<server>__<tool>`. Every MCP tool is treated as mutating (confirmation required on every call) regardless of what it reports — the MCP `readOnlyHint` annotation is advisory and could be wrong (or set by a malicious server specifically to skip confirmation), so locode never trusts it. One misconfigured server doesn't block the others — check `/mcp` for per-server connection status.
- **Claude Code plugins**: `locode plugin add <path-or-git-url>` installs a Claude Code-compatible plugin — locode reads its `.claude-plugin/plugin.json`, then loads it all directly: MCP servers (its `.mcp.json` or manifest `mcpServers`, merged in like any other MCP server), slash commands (`commands/*.md` — frontmatter `description`/`argument-hint`, body is a template expanded with `$ARGUMENTS`/`$1..$9` and submitted as your message), agents (`agents/*.md` — the body becomes a sub-agent's system prompt, exposed as a callable tool named `agent__<plugin>__<agent>`; a `tools:` frontmatter list restricts what it can use, with Claude Code's built-in tool names — Read, Grep, Edit, etc. — automatically mapped to locode's equivalents), hooks (`hooks/hooks.json`, see below), and skills (`skills/*/SKILL.md`, see below). Check `/plugins` for what's loaded.
- **Skills**: named instructions the model loads on demand rather than a hook or a sub-agent — every installed skill (`skills/<name>/SKILL.md`) is exposed through one shared `skill` tool, whose own description lists every skill's name and "use this when..." blurb so the model knows when to call it. You can also invoke one directly with `/<skill-name> [request]`, which skips the model's own judgment and submits the skill's instructions (plus your request, if any) as the turn. Check `/skills` for what's installed.
- **Hooks**: shell commands that fire on session lifecycle events — `SessionStart`, `UserPromptSubmit`, `PreToolUse`, `PostToolUse`, `PermissionRequest`, `SubagentStart`, `SubagentStop`, `CwdChanged`, `FileChanged`, `ConfigChange`, `Stop`, `SessionEnd`. Configured the same way MCP servers are — plugin-bundled (`hooks/hooks.json`), user-level (`locode hooks path`, hand-edited), and project-level (`.locode/hooks.json`) all merge together, every hook from every source runs. A hook receives a JSON payload on stdin (`session_id`, `cwd`, `hook_event_name`, plus event-specific fields like `prompt` or `tool_name`/`tool_input`); exit 0 allows (stdout becomes injected context for `SessionStart`/`UserPromptSubmit`), exit 2 blocks (stderr is the reason shown), anything else is a non-blocking warning. `PreToolUse`, `UserPromptSubmit`, and `PermissionRequest` can block; the rest are fire-and-forget. `PreToolUse` fires before permission modes apply, so a hook's block can't be bypassed by auto-accept. `command` and `http` hook types run an external process/request; a `prompt` hook just injects its static `message` as additional context instead. Command/http hooks can opt into structured JSON output via `outputSchema: "json"`. `SessionEnd` fires on every exit path (including Ctrl+C and external `SIGTERM`/`SIGHUP`). Check `/hooks` for what's configured.
- **Hooks**: shell commands that fire on session lifecycle events — `SessionStart`, `UserPromptSubmit`, `PreToolUse`, `PostToolUse`, `PermissionRequest`, `SubagentStart`, `SubagentStop`, `CwdChanged`, `FileChanged`, `ConfigChange`, `Stop`, `SessionEnd`. Configured the same way MCP servers are — plugin-bundled (`hooks/hooks.json`), user-level (`locode hooks path`, hand-edited), and project-level (`.locode/hooks.json`) all merge together, every hook from every source runs. A hook receives a JSON payload on stdin (`session_id`, `cwd`, `hook_event_name`, plus event-specific fields like `prompt` or `tool_name`/`tool_input`); exit 0 allows (stdout becomes injected context for `SessionStart`/`UserPromptSubmit`), exit 2 blocks (stderr is the reason shown), anything else is a non-blocking warning. `PreToolUse`, `UserPromptSubmit`, and `PermissionRequest` can block; the rest are fire-and-forget. `PreToolUse` fires before permission modes apply, so a hook's block can't be bypassed by auto-accept. `command` and `http` hook types are supported (`prompt` is declared in the config format but not yet executed); command hooks can opt into structured JSON output via `outputSchema: "json"`. `SessionEnd` fires on every exit path (including Ctrl+C and external `SIGTERM`/`SIGHUP`). Check `/hooks` for what's configured.
- **Tool-calling mode**: on connect, locode probes whether the model reliably uses native OpenAI-style function calling. If not, it switches to a prompt-based fallback mode where the model is instructed to emit tool calls as fenced ` ```tool_call ``` ` JSON blocks, which locode parses itself. The result is cached per backend+model so future sessions skip the probe. Override with `--tool-mode native|fallback|auto` or the in-session `/mode` command.
- **Context tracking & compaction**: the status bar shows `ctx NN%` — context window usage, from real `usage.prompt_tokens` when the backend reports it (requested via `stream_options.include_usage`), or a `~`-prefixed char-based estimate otherwise. The window size itself is auto-detected (Ollama's `/api/show`, then LM Studio's `/api/v0/models`) and cached per backend+model; falls back to a configurable default (`locode config set contextWindow <n>`, or `$LOCODE_CONTEXT_WINDOW`) if neither responds. At 85% usage, locode automatically asks the model to summarize the conversation and replaces the history with that summary (a notice tells you when this happens) — or trigger it yourself anytime with `/compact`.
- **Code intelligence (LSP)**: `definition`, `references`, and `diagnostics` use a real language server (LSP) for go-to-definition, find-all-references, and type/syntax error checks — the same engine an editor's Problems panel uses, more precise than `grep`. A server is lazily started per language on first use and reused for the whole session: `typescript-language-server` (TypeScript/JavaScript), `pyright-langserver` (Python), `gopls` (Go), `rust-analyzer` (Rust), and `clangd` (C/C++ — one clangd covers both). The relevant server binary must be on your PATH; if it isn't, the tool returns a clear "install X" error. After any edit, locode syncs the file to the live server so a subsequent `diagnostics` call reflects the change (it waits for the server to publish fresh diagnostics rather than reading a stale snapshot). Add or override servers with `locode config set lspServers` (see Config).
Note: even models with genuine native tool-calling support occasionally emit a tool call as plain text instead of a real structured call — this is model sampling variance, not a bug. If a turn seems to "describe" a tool call instead of running it, just ask again or try `/mode fallback`.
## Slash commands
@@ -115,7 +112,6 @@ Note: even models with genuine native tool-calling support occasionally emit a t
/model <name> switch the model used for the current backend
/backend <name> switch backend (ollama | lmstudio), keeps current model
/mode <name> view or force tool-call mode (native | fallback)
/mouse [on|off] toggle mouse tracking (on by default; hold Shift+click/drag for native text selection)
/perm [mode] cycle or set permission mode (default | plan | auto-edit | auto-accept)
/status show current model, backend, tool-call mode, and cwd
/dashboard show session stats: token I/O, elapsed/model time, turns, tool calls
@@ -127,7 +123,7 @@ Note: even models with genuine native tool-calling support occasionally emit a t
/hooks show configured hooks per lifecycle event
/skills show installed skills; /<skill-name> [request] invokes one directly
/compact summarize the conversation now to free up context
/export [file] save the conversation as markdown (or /export json [file] for a full JSON dump incl. tool calls/results)
/export [file] save the conversation as markdown — opens an editable filename prompt (default: locode-export-<timestamp>.md)
/import <path> [caption] attach a local file or image to your next message
/clear clear conversation history
/help show this help
@@ -147,11 +143,6 @@ locode config set autoCompactThreshold 0.85 # fraction of context window at whi
locode config set requestTimeoutMs 300000 # per-request timeout in ms (default 180000); raise this if
# your backend queues requests behind a concurrency limit
# (e.g. Ollama's OLLAMA_NUM_PARALLEL) under multi-session load
locode config set lspServers '{"java":{"command":"jdtls","extensions":[".java"]}}' # add a language server
# (JSON object keyed by language id; built-in ids
# override command/args, new ids add support and
# require extensions). Built-ins: typescript,
# python, go, rust, c (C/C++ share clangd).
locode config get
locode config path
```
@@ -160,12 +151,13 @@ locode config path
- Requires a real interactive terminal (TTY) — you can't pipe input into it or run it from a non-interactive script.
- Native tool-calling reliability varies by model and is non-deterministic even for capable models (see above).
- No OS-level sandboxing (no container/VM isolation) — mutating tools operate on the real filesystem/shell with the permissions of the user running `locode`. Only approve commands you understand. Two lightweight guardrails run unconditionally regardless of permission mode (including `auto-accept`), as a safety floor rather than a full sandbox: `write_file`/`edit_file`/`bash`'s `cwd` override can't target a path outside the working directory (`../` traversal, an absolute path elsewhere, or — on Windows — a different drive all refuse), and `bash` refuses a short list of unambiguously catastrophic commands (wiping the filesystem root or home directory, a fork bomb, formatting/wiping a whole drive, writing raw data to a block device) before they'd ever run. Neither guard stops a model from doing damage confined to *within* the project directory, or running something merely inadvisable — see `src/tools/pathGuard.ts` and `src/tools/bashGuard.ts`.
- In-app scrollback: mouse wheel scrolls the conversation view when mouse tracking is on (the default). PageUp/PageDown also scroll a page at a time. Hold Shift+click/drag for native terminal text selection and copy (when mouse tracking is on). Scrolling back up unpins the view from the latest message; scrolling back to the bottom (or sending a new message) re-pins it so new messages auto-scroll into view. Toggle mouse tracking with `/mouse on|off`.
- No sandboxing beyond the confirmation prompts — mutating tools operate on the real filesystem/shell with the permissions of the user running `locode`. Only approve commands you understand.
- Session resume replays prior user/assistant text so you can see it, but it doesn't re-display prior tool-call/tool-result lines from before the resume (the model still has that history — it's just not re-rendered).
- No in-app scrollback — once a message scrolls off the top of the window it's gone until you resize the terminal taller (the conversation itself is still intact and sent to the model; this only affects what you can visually re-read).
- Windows shell quoting for the `bash` tool has only had light testing; behavior may differ from Unix shells for complex quoting.
- `git_commit` covers add/commit/create_branch/checkout/push/reset/stash/merge/rebase/delete_branch. Use `bash` for anything beyond that.
- MCP tool results support text, image, audio, and resource content blocks. Images are returned in the same shape as `read_file` so vision-capable models can see them; audio and binary resources are summarized. Remote (HTTP) MCP servers support static headers (e.g. a bearer token) but not OAuth flows.
- Compaction (`/compact` or automatic) replaces history with a model-generated prose summary — it costs one extra model call and loses tool-call/tool-result detail (the model's own account of what happened survives; the raw record doesn't). The auto-compact threshold defaults to 85% and is configurable via `locode config set autoCompactThreshold` or `LOCODE_AUTO_COMPACT_THRESHOLD`.
- Plugin support (`locode plugin add`) now covers every part of a plugin: MCP servers, slash commands, agents, hooks, and skills. A slash command's `allowed-tools` frontmatter restricts that one invocation's toolset (same tool-name translation as an agent's `tools:` — see `/plugins`); a skill invoked directly via `/<skill-name>` isn't restricted this way, since skills have no `allowed-tools` field of their own. Duplicate MCP server names across sources are now detected and surfaced in `/mcp` — project-level wins over user-level wins over plugin-level. Duplicate slash-command and skill names are also surfaced in `/plugins` and `/skills`.
- Plugin support (`locode plugin add`) now covers every part of a plugin: MCP servers, slash commands, agents, hooks, and skills. A plugin's `allowed-tools` restriction on a command isn't enforced (the expanded prompt just runs as a normal turn with the full toolset). Duplicate MCP server names across sources are now detected and surfaced in `/mcp` — project-level wins over user-level wins over plugin-level. Duplicate slash-command and skill names are also surfaced in `/plugins` and `/skills`.
- Skills are exposed as one shared `skill` tool rather than one tool per skill — if two plugins install a skill with the same name, the first plugin in load order wins and the collision is shown in `/skills`. Skills can bundle sibling `references/*.md` files that are included when the skill is invoked.
- Hooks cover the most useful subset of Claude Code's lifecycle events: `SessionStart`, `UserPromptSubmit`, `PreToolUse`, `PostToolUse`, `PermissionRequest`, `SubagentStart`, `SubagentStop`, `CwdChanged`, `FileChanged`, `ConfigChange`, `Stop`, `SessionEnd`. `CwdChanged` is declared in the config format but not yet fired anywhere — a cwd never changes mid-session in locode today, so configuring it is a no-op for now. All three hook types are supported: `command` and `http` run an external process/request, and `prompt` just injects its static `message` as additional context (the same way a command/http hook's stdout does) — it has no process to fail, so it can't block an event the way a command hook's exit code 2 can. Command/http hooks can opt into structured JSON output via `outputSchema: "json"`. `PreToolUse`, `UserPromptSubmit`, and `PermissionRequest` can block; the rest are fire-and-forget. `SessionEnd` fires on every exit path (including Ctrl+C and external `SIGTERM`/`SIGHUP`), so it doesn't always have a real session id to report. All hooks for an event run in parallel with no defined ordering, and every configured hook always runs — there's no way to disable one without editing the file it came from.
- Hooks cover the most useful subset of Claude Code's lifecycle events: `SessionStart`, `UserPromptSubmit`, `PreToolUse`, `PostToolUse`, `PermissionRequest`, `SubagentStart`, `SubagentStop`, `CwdChanged`, `FileChanged`, `ConfigChange`, `Stop`, `SessionEnd`. `CwdChanged` is declared in the config format but not yet fired anywhere — a cwd never changes mid-session in locode today, so configuring it is a no-op for now. `command` hooks and `http` hooks are supported; `prompt` hooks are declared in the config format but not yet executed. Command hooks can opt into structured JSON output via `outputSchema: "json"`. `PreToolUse`, `UserPromptSubmit`, and `PermissionRequest` can block; the rest are fire-and-forget. `SessionEnd` fires on every exit path (including Ctrl+C and external `SIGTERM`/`SIGHUP`), so it doesn't always have a real session id to report. All hooks for an event run in parallel with no defined ordering, and every configured hook always runs — there's no way to disable one without editing the file it came from.
-91
View File
@@ -1,91 +0,0 @@
# locode 설계 메모
## 프로젝트 개요
**locode** — Claude Code의 설계 철학을 가져와 구현한 에이전트 코딩 CLI. TypeScript + Ink(React-for-CLI) 기반. 백엔드는 OpenAI 호환 `/v1/chat/completions` 엔드포인트 사용 (Ollama, LM Studio, 클라우드 API 모두 지원).
- 핵심 철학: "신뢰할 수 없고 느리고 비전/툴콜 지원이 불확실한 모델"이라는 현실에 맞춰 모든 가정을 비관적으로 재단. 로컬/클라우드 자동 감지로 프롬프트 분기.
- Claude Code 플러그인 포맷을 직접 소비하는 하위호환 브리지 (`.claude-plugin/plugin.json`, commands/agents/skills/hooks/MCP)
---
## 아키텍처 결정 이유
### Normal Screen + `<Static>` (alternate screen 폐지)
이전에는 alternate screen buffer + 인앱 가상 스크롤 + 마우스 트래킹을 직접 구현했으나 완전히 폐지. 이유:
- **코드 복잡도**: 마우스 SGR-1006 파싱, 선택 영역, 스크롤 상태 관리가 App.tsx의 절반을 차지
- **터미널 호환성**: alternate screen은 SSH, tmux, Windows Terminal 등에서 파편화 심함
- **버그 발생률**: 마우스 크래시, 선택 텍스트 깨짐 등 이슈가 끝없이 발생
- **현재 방식**: Ink `<Static>`으로 완료된 히스토리를 한 번만 렌더링 → 터미널 스크롤백의 영구 부분. 재렌더링 없음. 마우스는 터미널 네이티브에 의존.
### 로컬/클라우드 모델 감지: `isSmallLocalModel()`
- **이전**: `isLocalBackendURL(baseURL)` — URL이 localhost면 무조건 "로컬 모델"
- **문제**: Ollama가 클라우드 라우팅 모델(`glm-5.2:cloud`, `qwen3.5:397b-cloud`)도 같은 localhost에서 서비스함. 이 모델들은 컨텍스트 128K~1M이고 툴콜도 안정적인데 로컬용 보수적 프롬프트가 적용됨.
- **해결**: `isSmallLocalModel(baseURL, model)` = `isLocalBackendURL(baseURL) && !isCloudRoutedModelName(model)` — 모델명의 `:cloud`/`:-cloud` 태그로 구분.
- **기본값**: `isLocal`의 기본값을 `true`에서 `false`(클라우드)로 변경. 명시적 지정이 없으면 보수적이 아닌 기본 프롬프트 사용.
### 컨텍스트 윈도우 기본값 분리
- **이전**: `DEFAULT_CONTEXT_WINDOW = 8192` (단일, 로컬 모델 기준)
- **현재**: `DEFAULT_CONTEXT_WINDOW_LOCAL = 8192`, `DEFAULT_CONTEXT_WINDOW_CLOUD = 131072` — 클라우드/Ollama 클라우드 라우팅 모델은 128K~1M 컨텍스트를 가지므로 8192는 과도하게 보수적.
### 번인레이트(🔥) 계산
- **이전**: `outputTokens / 세션 경과 시간` — 사용자가 방치하면 번인레이트가 0에 수렴해서 의미 없음
- **현재**: `outputTokens / modelTimeMs` — 실제 모델 응답 시간으로 계산. "이 모델이 얼마나 빠르게 토큰을 뿜는가"를 정확히 반영.
### 파일 인코딩: CRLF/LF 혼재
- Windows 환경에서는 CRLF, Unix에서는 LF가 섞여 있음. `edit_file`/`multi_edit`은 매칭 전 LF로 정규화하고, 쓰기 전 원래 EOL을 복원. 이것 없이는 Windows에서 거의 모든 edit_file이 실패함.
---
## 핵심 파일 맵
- `src/ui/ink/index.tsx` — 진입점. alternate screen 없이 Ink render. cleanup 시 flush + 종료.
- `src/ui/ink/App.tsx` — 메인 UI 컴포넌트. `<Static>` + 라이브 영역. 상태: starting→connecting→loading-models→model-select/session-select→input.
- `src/ui/ink/ChatInput.tsx` — 커스텀 multiline 입력. Shift+Enter 줄바꿈, bracket paste, @멘션 fuzzy picker, IME 커서.
- `src/agent/loop.ts` — 메인 에이전트 루프. 턴/스트리밍/툴콜/컴팩션/서브에이전트/병렬 툴 배치/반복 루프 감지.
- `src/agent/session.ts` — Session 객체, 통계, 상태, mutation gate.
- `src/agent/systemPrompt.ts` — 시스템 프롬프트 빌더 (`isSmallLocalModel` 기반 로컬/클라우드 분기).
- `src/config/defaults.ts` — 모든 기본값. `isSmallLocalModel()`, `isCloudRoutedModelName()`, `DEFAULT_CONTEXT_WINDOW_LOCAL/CLOUD` 등.
- `src/toolcalling/` — native 어댑터, fallback 파서/프롬프트, partialJson 복구, resolve (Ollama 빈키 복구 포함).
---
## 설정 기본값
| 설정 | 기본값 | 비고 |
|---|---|---|
| `DEFAULT_CONTEXT_WINDOW_LOCAL` | 8192 | 작은 로컬 모델 폴백 |
| `DEFAULT_CONTEXT_WINDOW_CLOUD` | 131072 | 클라우드/클라우드 라우팅 폴백 |
| `DEFAULT_MAX_ITERATIONS` | **300** | 50→100→300 상향 |
| `DEFAULT_MAX_OUTPUT_TOKENS` | **131072** | 128K. GLM 등 1M 컨텍스트 모델 대응 |
| `DEFAULT_MAX_RETRIES` | 0 | SDK 지수 백오프 |
| `DEFAULT_AUTO_COMPACT_THRESHOLD` | 0.85 | |
| `DEFAULT_REQUEST_TIMEOUT_MS` | 180,000 | 3분 |
| `DEFAULT_SUBAGENT_TIMEOUT_MS` | 600,000 | 10분 |
| `MAX_EMPTY_RESPONSE_RETRIES` | **3** | 1→3 상향 |
| `MAX_SUBAGENT_DEPTH` | 1 | 서브에이전트 중첩 금지 |
| `MAX_PRESERVED_TAIL_MESSAGES` | 8 | 컴팩션 시 보존 |
| `MAX_PRESERVED_TAIL_FRACTION` | 0.3 | 컴팩션 시 보존 비율 |
| `MAX_RETAINED_IMAGES` | 2 | 히스토리 이미지 보존 |
---
## 트러블슈팅 힌트
- **"자꾸 에러"**: 주요 원인은 max_tokens 잘림 → malformed 툴콜. 동적 max_tokens + CRLF 제어문자 이스케이프로 해결됨.
- **CRLF edit_file 매칭 버그**: LF 정규화 공간에서 매칭, 쓰기 전 원래 EOL 복원으로 해결됨.
- **"Paused after N steps"**: `locode config set maxIterations <number>` (기본 300)
- **클라우드 모델 빈 응답**: `MAX_EMPTY_RESPONSE_RETRIES=3`으로 재시도
- **로컬/클라우드 프롬프트 분기**: `isSmallLocalModel(baseURL, model)` — Ollama 클라우드 라우팅 모델(`:cloud` 태그)은 localhost여도 클라우드 프롬프트 사용
- **반복 루프 감지**: `detectRepetitionLoop()` — 스트리밍 텍스트에서 짧은 반복 패턴 감지 시 중단
- **번인레이트(🔥)**: `outputTokens ÷ modelTimeMs` 기준 (세션 경과 시간이 아닌 실제 모델 응답 시간)
- **IPv6 localhost**: `isLocalBackendURL()`은 `[::1]` 형식(WHATWG URL 직렬화)도 인식
---
## 의존성 제약
- **marked는 15에 고정**. `marked-terminal@7.3.0`의 peer가 `marked >=1 <16`이고 marked-terminal 업데이트가 없음. marked 16+로 올리려면 marked-terminal을 교체하거나 peer를 강제해야 함.
- **typescript는 7.x (네이티브 컴파일러 포팅)**. 프로젝트는 `tsc` CLI로 타입체크만 하고 programmatic API를 안 씀 → 네이티브 포트로 안전하게 이전. 빌드는 tsup/esbuild라 tsc와 무관. 플랫폼별 `@typescript/typescript-*` 바이너리가 optional dep으로 붙음.
- **wrap-ansi는 10.x (ink와 동일)**. `renderMarkdown`이 히스토리 출력을 터미널 폭으로 하드랩할 때 사용 — ink 내부 래핑과 같은 string-width v8을 공유해야 폭 계산이 어긋나지 않음. v10은 타입 내장(앰비언트 선언 불필요).
- `allowScripts`에 트리에 실제로 존재하는 esbuild 버전을 모두 나열해야 함 (현재 `0.27.2` = tsup/vite, `0.28.1` = tsx). 빠지면 postinstall 경고.
## TODO
- LSP 실서버 통합 테스트 (실제 tsserver/pyright 띄워서 검증)
- task store 영속화 (세션에 task 저장)
+632 -1393
View File
File diff suppressed because it is too large Load Diff
+18 -19
View File
@@ -1,6 +1,6 @@
{
"name": "locode",
"version": "0.5.2",
"version": "0.6.0",
"description": "Agentic coding CLI for local models served via Ollama and LM Studio",
"type": "module",
"bin": {
@@ -10,7 +10,7 @@
"dist"
],
"engines": {
"node": ">=22.12.0"
"node": ">=22"
},
"scripts": {
"build": "tsup",
@@ -20,39 +20,38 @@
"prepublishOnly": "npm run build"
},
"dependencies": {
"@modelcontextprotocol/sdk": "^1.30.0",
"@anthropic-ai/claude-code": "^2.1.233",
"@modelcontextprotocol/sdk": "^1.29.0",
"@vscode/ripgrep": "^1.18.0",
"commander": "^15.0.0",
"commander": "^13.0.0",
"diff": "^9.0.0",
"env-paths": "^4.0.0",
"execa": "^10.0.1",
"execa": "^9.6.1",
"fast-glob": "^3.3.3",
"ink": "^7.1.1",
"ink": "^7.1.0",
"ink-select-input": "^6.2.0",
"ink-spinner": "^5.0.0",
"ink-text-input": "^6.0.0",
"marked": "^15.0.12",
"marked-terminal": "^7.3.0",
"openai": "^7.13.0",
"react": "^19.3.0",
"string-width": "^8.2.2",
"openai": "^6.45.0",
"react": "^19.2.7",
"string-width": "^8.2.1",
"tree-kill": "^1.2.2",
"vscode-languageserver-protocol": "^3.18.3",
"vscode-uri": "^3.2.0",
"wrap-ansi": "^10.0.1",
"zod": "^4.6.1"
"zod": "^4.4.3"
},
"devDependencies": {
"@types/marked-terminal": "^6.1.1",
"@types/node": "^26.5.1",
"@types/react": "^19.3.0",
"@types/node": "^22.10.0",
"@types/react": "^19.2.17",
"tsup": "^8.3.0",
"tsx": "^4.23.13",
"typescript": "^7.0.2",
"vitest": "^5.0.0"
"tsx": "^4.19.0",
"typescript": "^5.7.0",
"vitest": "^3.0.0"
},
"allowScripts": {
"esbuild@0.27.2": true,
"@anthropic-ai/claude-code@2.1.233": true,
"esbuild@0.27.7": true,
"esbuild@0.28.1": true
}
}
+15 -15
View File
@@ -1,33 +1,33 @@
import type { TodoItem } from "../tools/types.js";
import type { TaskSummary } from "../tools/task.js";
export type AgentEvent =
| { type: "text_delta"; delta: string }
| { type: "text_done"; fullText: string }
| { type: "thinking_delta"; delta: string }
| { type: "thinking_done"; fullThinking: string }
/** Throw away any text streamed so far this turn without committing it as an assistant message —
* emitted when a partially-streamed native tool-call turn turns out to have malformed args and is
* retried non-streaming, so the UI doesn't carry the stale partial into the retry's output. */
| { type: "stream_discard" }
/** `name`/`args` are only populated for an actually-resolved tool call (not the "unknown tool"
* error path) — the file panel's Activity tab (App.tsx) uses them to track which files a
* read_file/write_file/edit_file call touched, without having to re-parse the display `label`. */
| { type: "tool_call"; label: string; name?: string; args?: unknown }
/** `name`/`result` mirror `tool_call`'s — only populated when a tool actually ran (not a
* hook-blocked/denied/unknown-tool result), for the same file-panel tracking purpose. */
| { type: "tool_result"; summary: string; isError: boolean; name?: string; result?: unknown }
| { type: "tool_call"; label: string }
| { type: "tool_result"; summary: string; isError: boolean; diff?: string }
/** The model finished a turn in plan mode with a prose plan. Rendered as a distinct `plan`
* HistoryItem (not a plain assistant message) and followed by an approve/reject prompt built
* on the same confirm primitive tool permissions use — see maybePresentPlan in agent/loop.ts. */
| { type: "plan_presented"; text: string }
/** The user approved the presented plan. loop.ts has already exited plan mode; the UI drives
* implementation by injecting a "proceed" follow-up turn (see App.tsx submitTurn). */
| { type: "plan_approved"; text: string }
/** A sub-agent's tool call or result, forwarded to the parent so its work is visible while it
* runs headless. Routed to the dedicated sub-agent panel below the input (not the main
* scrollback) — see App.tsx. */
* runs headless. Rendered inline in the main scrollback (tagged with the sub-agent's
* description) — see App.tsx and HistoryItemView. */
| { type: "subagent"; description: string; line: SubagentLine }
/** A hook (see hooks/runner.ts) blocked something or failed non-fatally — surfaced as a notice. */
| { type: "hook_notice"; text: string; isError: boolean }
/** A general informational notice from the loop itself (not tied to a hook) — e.g. a mid-turn
* auto-compaction. Surfaced the same way as hook_notice. */
| { type: "notice"; text: string; isError: boolean }
/** The `todo_write` tool replaced the session's task checklist — carries the full new list so
* the UI can render it as a standalone checklist item rather than raw JSON tool output. */
| { type: "todos_update"; todos: TodoItem[] };
/** A task_create/task_update mutation changed the session's task store — carries a snapshot so
* the UI can render the current checklist as a standalone item rather than raw JSON tool output. */
| { type: "tasks_update"; tasks: TaskSummary[] };
export type SubagentLine =
| { kind: "call"; label: string }
+1217 -467
View File
File diff suppressed because it is too large Load Diff
+674 -651
View File
File diff suppressed because it is too large Load Diff
-229
View File
@@ -1,229 +0,0 @@
import { describe, expect, it, vi } from "vitest";
import { z } from "zod";
import { runTurn } from "./loop.js";
import { agentTool } from "../tools/agentTool.js";
import { createSession } from "./session.js";
import { buildToolSet } from "../tools/toolset.js";
import type { ToolDef } from "../tools/types.js";
/** A one-shot streaming response: yields `chunk` once, then ends. */
function oneShotStream(chunk: any) {
let yielded = false;
return {
[Symbol.asyncIterator]: () => ({
next: async () => {
if (yielded) return { done: true, value: undefined };
yielded = true;
return { done: false, value: chunk };
},
}),
};
}
function textChunk(text: string) {
return { choices: [{ delta: { content: text }, finish_reason: "stop" }] };
}
function toolCallChunk(name: string, args: string, id = "call_0") {
return {
choices: [
{
delta: { tool_calls: [{ index: 0, id, function: { name, arguments: args } }] },
finish_reason: "tool_calls",
},
],
};
}
describe("agent tool / parallel sub-agents", () => {
it("runs a `tasks` batch and returns each result independently", async () => {
const agentArgs = JSON.stringify({
description: "parallel research",
tasks: [
{ description: "task A", prompt: "do A" },
{ description: "task B", prompt: "do B" },
],
});
// create() is called: 1 parent agent-call, then sub-agent A, then sub-agent B, then parent final.
let call = 0;
const fakeClient = {
chat: {
completions: {
create: vi.fn(async () => {
call++;
if (call === 1) return oneShotStream(toolCallChunk("agent", agentArgs));
if (call === 2) return oneShotStream(textChunk("result-A"));
if (call === 3) return oneShotStream(textChunk("result-B"));
return oneShotStream(textChunk("all done"));
}),
},
},
} as any;
const session = createSession(fakeClient, "test-model", process.cwd(), async () => "once", "native", [agentTool]);
const result = await runTurn(session, "research A and B in parallel", () => {});
expect(result).toBe("all done");
// The agent tool's result must carry both sub-agent answers as a `results` array.
const toolResultMsg = session.messages.find(
(m) => m.role === "tool" && typeof (m as any).content === "string" && (m as any).content.includes("results"),
) as any;
expect(toolResultMsg).toBeTruthy();
const parsed = JSON.parse(toolResultMsg.content);
expect(parsed.results).toHaveLength(2);
const byDesc = Object.fromEntries(parsed.results.map((r: any) => [r.description, r]));
expect(byDesc["task A"].result).toBe("result-A");
expect(byDesc["task B"].result).toBe("result-B");
});
it("a failed sub-agent surfaces as its own error without discarding sibling results", async () => {
const agentArgs = JSON.stringify({
description: "mixed batch",
tasks: [
{ description: "ok", prompt: "succeed" },
{ description: "boom", prompt: "fail" },
],
});
let call = 0;
const fakeClient = {
chat: {
completions: {
create: vi.fn(async () => {
call++;
if (call === 1) return oneShotStream(toolCallChunk("agent", agentArgs));
if (call === 2) return oneShotStream(textChunk("ok-result"));
if (call === 3) throw new Error("sub-agent boom failed");
return oneShotStream(textChunk("done"));
}),
},
},
} as any;
const session = createSession(fakeClient, "test-model", process.cwd(), async () => "once", "native", [agentTool]);
await runTurn(session, "run mixed batch", () => {});
const toolResultMsg = session.messages.find(
(m) => m.role === "tool" && typeof (m as any).content === "string" && (m as any).content.includes("results"),
) as any;
expect(toolResultMsg).toBeTruthy();
const parsed = JSON.parse(toolResultMsg.content);
const byDesc = Object.fromEntries(parsed.results.map((r: any) => [r.description, r]));
expect(byDesc.ok.result).toBe("ok-result");
expect(byDesc.boom.error).toBeTruthy();
expect(byDesc.boom.error).toMatch(/boom/);
});
});
describe("mutation gate / serialization", () => {
it("serializes mutating tool calls so two never run concurrently", async () => {
let inFlight = 0;
let maxOverlap = 0;
const slowEdit: ToolDef = {
name: "slow_edit",
description: "slow edit",
schema: z.object({}),
mutating: true,
handler: async () => {
inFlight++;
maxOverlap = Math.max(maxOverlap, inFlight);
await new Promise((r) => setTimeout(r, 20));
inFlight--;
return { ok: true };
},
};
function twoEditsChunk() {
return {
choices: [
{
delta: {
tool_calls: [
{ index: 0, id: "c0", function: { name: "slow_edit", arguments: "{}" } },
{ index: 1, id: "c1", function: { name: "slow_edit", arguments: "{}" } },
],
},
finish_reason: "tool_calls",
},
],
};
}
let call = 0;
const fakeClient = {
chat: {
completions: {
create: vi.fn(async () => {
call++;
const chunk = call === 1 ? twoEditsChunk() : textChunk("done");
return oneShotStream(chunk);
}),
},
},
} as any;
const session = createSession(fakeClient, "m", process.cwd(), async () => "once", "native", [slowEdit]);
session.maxIterations = 5;
await runTurn(session, "two edits", () => {});
expect(maxOverlap).toBe(1);
});
it("lets read-only tools run concurrently (gate only blocks mutating)", async () => {
let inFlight = 0;
let maxOverlap = 0;
const fastRead: ToolDef = {
name: "fast_read",
description: "fast read",
schema: z.object({}),
mutating: false,
handler: async () => {
inFlight++;
maxOverlap = Math.max(maxOverlap, inFlight);
await new Promise((r) => setTimeout(r, 20));
inFlight--;
return { ok: true };
},
};
function threeReadsChunk() {
return {
choices: [
{
delta: {
tool_calls: [
{ index: 0, id: "c0", function: { name: "fast_read", arguments: "{}" } },
{ index: 1, id: "c1", function: { name: "fast_read", arguments: "{}" } },
{ index: 2, id: "c2", function: { name: "fast_read", arguments: "{}" } },
],
},
finish_reason: "tool_calls",
},
],
};
}
let call = 0;
const fakeClient = {
chat: {
completions: {
create: vi.fn(async () => {
call++;
const chunk = call === 1 ? threeReadsChunk() : textChunk("done");
return oneShotStream(chunk);
}),
},
},
} as any;
const session = createSession(fakeClient, "m", process.cwd(), async () => "once", "native", [fastRead]);
session.maxIterations = 5;
await runTurn(session, "three reads", () => {});
// All three reads are read-only → runToolBatch runs them concurrently → they overlap.
expect(maxOverlap).toBe(3);
});
});
-98
View File
@@ -1,98 +0,0 @@
import { describe, it, expect } from "vitest";
import { undoLastTurn, resetSession } from "./session.js";
import type { ChatCompletionMessageParam } from "openai/resources/chat/completions";
function makeSession(messages: ChatCompletionMessageParam[]) {
return {
messages,
lastContextTokens: 0,
lastContextTokensIsEstimate: true,
} as any;
}
describe("undoLastTurn", () => {
it("returns 0 when there are no user messages", () => {
const session = makeSession([
{ role: "system", content: "You are helpful." },
]);
expect(undoLastTurn(session)).toBe(0);
});
it("removes a single user turn at the end", () => {
const session = makeSession([
{ role: "system", content: "You are helpful." },
{ role: "user", content: "Hello" },
{ role: "assistant", content: "Hi there!" },
]);
const removed = undoLastTurn(session);
expect(removed).toBe(2);
expect(session.messages.length).toBe(1);
expect(session.messages[0]!.role).toBe("system");
});
it("removes user + assistant + tool results together", () => {
const session = makeSession([
{ role: "system", content: "You are helpful." },
{ role: "user", content: "Read the file" },
{ role: "assistant", content: "", tool_calls: [{ id: "tc1", type: "function", function: { name: "read_file", arguments: "{}" } }] } as any,
{ role: "tool", content: "file contents here", tool_call_id: "tc1" } as any,
{ role: "assistant", content: "The file contains..." },
]);
const removed = undoLastTurn(session);
expect(removed).toBe(4);
expect(session.messages.length).toBe(1);
});
it("only removes the last turn, keeping earlier turns", () => {
const session = makeSession([
{ role: "system", content: "You are helpful." },
{ role: "user", content: "First question" },
{ role: "assistant", content: "First answer" },
{ role: "user", content: "Second question" },
{ role: "assistant", content: "Second answer" },
]);
const removed = undoLastTurn(session);
expect(removed).toBe(2);
expect(session.messages.length).toBe(3);
expect((session.messages[2] as any).content).toBe("First answer");
});
it("handles consecutive user messages (removing only the last one)", () => {
const session = makeSession([
{ role: "system", content: "You are helpful." },
{ role: "user", content: "Message 1" },
{ role: "user", content: "Message 2" },
]);
const removed = undoLastTurn(session);
expect(removed).toBe(1);
expect(session.messages.length).toBe(2);
expect((session.messages[1] as any).content).toBe("Message 1");
});
it("updates context token tracking after undo", () => {
const session = makeSession([
{ role: "system", content: "You are helpful." },
{ role: "user", content: "Hello" },
{ role: "assistant", content: "Hi!" },
]);
session.lastContextTokens = 5000;
session.lastContextTokensIsEstimate = false;
undoLastTurn(session);
expect(session.lastContextTokensIsEstimate).toBe(true);
// lastContextTokens should be recalculated (smaller than before)
expect(session.lastContextTokens).toBeLessThan(5000);
});
});
describe("resetSession", () => {
it("clears all messages except system prompt", () => {
const session = makeSession([
{ role: "system", content: "You are helpful." },
{ role: "user", content: "Hello" },
{ role: "assistant", content: "Hi!" },
]);
resetSession(session);
expect(session.messages.length).toBe(1);
expect(session.messages[0]!.role).toBe("system");
});
});
+132 -83
View File
@@ -2,15 +2,16 @@ import { randomUUID } from "node:crypto";
import type OpenAI from "openai";
import type { ChatCompletionMessageParam } from "openai/resources/chat/completions";
import type { ToolCallMode } from "../backend/capabilityProbe.js";
import { DEFAULT_AUTO_COMPACT_THRESHOLD, DEFAULT_CONTEXT_WINDOW_CLOUD, DEFAULT_CONTEXT_WINDOW_LOCAL, DEFAULT_MAX_ITERATIONS } from "../config/defaults.js";
import { DEFAULT_AUTO_COMPACT_THRESHOLD, DEFAULT_CONTEXT_WINDOW, DEFAULT_MAX_ITERATIONS, DEFAULT_SUBAGENT_MAX_DEPTH, DEFAULT_SUBAGENT_MAX_ITERATIONS } from "../config/defaults.js";
import type { SessionRecord } from "../persistence/sessionStore.js";
import { PermissionManager } from "../permissions/permissionManager.js";
import type { ConfirmFn } from "../permissions/types.js";
import { TOOLS } from "../tools/index.js";
import { buildToolSet, type ToolSet } from "../tools/toolset.js";
import type { TodoItem, ToolDef } from "../tools/types.js";
import type { AskQuestionSpec, AskQuestionAnswer, ToolDef } from "../tools/types.js";
import { TaskStore } from "../tools/task.js";
import type { TaskStoreSnapshot } from "../tools/task.js";
import type { CronStore } from "../scheduler/cron.js";
import type { PermissionMode, PermissionRule } from "../permissions/types.js";
import { estimateTokens } from "../utils/tokens.js";
import { buildSystemPrompt } from "./systemPrompt.js";
@@ -47,12 +48,34 @@ export interface Session {
mode: ToolCallMode;
messages: ChatCompletionMessageParam[];
maxIterations: number;
/** Max tool calls in a single sub-agent turn launched from this session. Intentionally smaller
* than maxIterations (sub-agents run one focused task, one round, no auto-continue) so a runaway
* sub-agent fails fast and surfaces a "split the task" hint instead of burning a large budget. */
subagentMaxIterations: number;
/** Max nesting depth for sub-agents launched from this session. The main session is depth 0;
* a sub-agent it spawns is depth 1, and so on. A sub-agent at the cap has `agent` excluded from
* its toolset (with an explicit depth-check backstop in agent/loop.ts) so it can't delegate
* further. Configurable so it can be lowered in tests without touching env/config. */
subagentMaxDepth: number;
permissions: PermissionManager;
confirm: ConfirmFn;
/** Optional structured-question callback wired by the UI so the `ask_user_question` tool can
* prompt the user with multiple-choice options. Absent in headless/non-UI contexts (sub-agents),
* in which case the tool returns a clear "can't ask" error instead of hanging. */
askQuestion?: (questions: AskQuestionSpec[]) => Promise<AskQuestionAnswer[]>;
/** Promise-chain lock serializing calls to `confirm` across concurrent sub-agents. When a parent
* turn fans out multiple `agent` delegations in parallel (see runBatch in agent/loop.ts), each
* sub-agent shares this same mutex *holder* (subSession.confirmMutex = parent.confirmMutex, by
* reference) so their mutating-tool confirmation prompts queue one at a time instead of racing
* for the UI's single PendingPermission slot. The holder wraps a `chain` promise that
* withConfirmLock reassigns on each confirm; sharing the holder (rather than the promise itself)
* keeps every concurrent sub-agent queued on the same lock even as the chain advances. */
confirmMutex: { chain: Promise<void> };
/** Local tools plus any dynamically-discovered ones (currently: MCP) available for this session. */
toolset: ToolSet;
/** 0 for a normal session; incremented for each level of sub-agent nesting (capped at
* MAX_SUBAGENT_DEPTH in agent/loop.ts, independent of the toolset already excluding `agent`). */
* session.subagentMaxDepth in agent/loop.ts, independent of the toolset already excluding `agent`
* for sub-agents at the depth cap). */
subAgentDepth: number;
/** The model's context window in tokens — auto-detected where possible (backend/contextWindow.ts),
* otherwise a configured/hardcoded fallback (see contextWindowIsEstimate). */
@@ -85,29 +108,48 @@ export interface Session {
* project's CLAUDE.md/AGENTS.md conventions survive everything that regenerates messages[0].
* Null when neither file exists. */
projectInstructions: string | null;
/** Current task checklist shown to the user via the `todo_write` tool — session-scoped state
* since checklist items are a snapshot of progress, not part of the model-visible conversation. */
todos: TodoItem[];
/** Structured task store backing the task_create/list/get/update tools — session-scoped,
* in-memory (not persisted), and independent of the flat `todos` checklist. */
/** The user's personal memory file (config dir / memory.md), folded into the system prompt
* alongside projectInstructions so learned preferences/feedback survive across sessions and repos.
* Null when the file is absent or empty. See utils/userMemory.ts. */
userMemory: string | null;
/** The session's structured task store (dependency graph + ownership), surfaced to the model via
* the task_create/list/get/update tools and to the UI via `tasks_update` events. Session-scoped —
* tasks are a progress snapshot, not part of the model-visible conversation, so not persisted. */
taskStore: TaskStore;
/** A serialized-mutation gate shared by EVERY tool call in this session — including sub-agents
* spawned in parallel via the `agent` tool's `tasks` array. Parallel sub-agents share their
* parent's session, so without a lock two of them could simultaneously call a mutating tool,
* race on the single React `permission` slot (makeConfirmFn), and interleave filesystem writes.
* This gate lets read-only tools run concurrently (matching runToolBatch) while serializing
* mutating ones: each mutating call awaits the previous one before it even prompts, so
* permission prompts stay one-at-a-time and edits can't overlap. The chain is per-session,
* so a top-level turn and its sub-agents all funnel through the same queue. Initialized as a
* resolved promise so the first caller doesn't wait on anything. */
mutationGate: Promise<void>;
/** True when this is a small/less-capable local model, not just a local *backend* — Ollama's
* cloud-routed models (e.g. "glm-5.2:cloud") share a localhost endpoint with genuinely local
* ones, so this is more than a baseURL check. Used to tailor the system prompt — small local
* models need extra guidance about their limitations; everything else gets a leaner prompt
* without self-fulfilling "you may produce empty responses" framing. Derived by
* isSmallLocalModel(baseURL, model); recomputed on /model and /backend switches (see App.tsx). */
isLocal: boolean;
/** The session's cron/wakeup scheduler, set by the App after creating the session (the App owns its
* lifecycle: starts the tick with enqueue/isIdle callbacks, stops it on unmount). Absent in
* non-UI contexts. Used by the cron_create/list/delete and schedule_wakeup tools. */
cronStore?: CronStore;
/** Tracks the most recent file mutation (write_file or edit_file) so the user can roll it back
* with the /undo slash command. The path is stored as the user-supplied path rather than a
* resolved absolute path, so the undo re-uses the same relative path logic as the original edit. */
lastEdit: { path: string; previousContent: string } | null;
/** Resumable sub-agent sessions keyed by their agentId, so the parent can continue one with a
* follow-up message via the `send_message` tool (ctx.resumeSubAgent). Only sub-agents that ran in
* the shared cwd are stored here — worktree-isolated parallel agents' cwd is cleaned up after the
* batch, so they're fire-and-forget and never added. Cleared on resetSession. */
subAgentSessions: Map<string, Session>;
/** Named teammates: maps a teammate name (supplied via the `agent` tool's `name` arg) to its
* agentId, so `send_message` can address it by name and `list_teammates` can roster it. Backed by
* the same resumable sub-agent sessions as subAgentSessions — this is just a name → agentId index
* over them. Independent per session (a sub-agent gets its own map for nested teammates). Cleared
* on resetSession alongside subAgentSessions. */
namedAgents: Map<string, string>;
/** The active interactive worktree session, set by `enter_worktree` and cleared by `exit_worktree`.
* While set, `cwd` points at the worktree dir (an isolated checkout on its own branch) and file
* tools operate there; `originalCwd` is restored on exit. Undefined when not in a worktree session.
* Not persisted — worktree sessions don't survive an app restart (the branch stays in git, so the
* work itself isn't lost; re-enter via git if needed). */
worktree?: { dir: string; branch: string; originalCwd: string };
/** App-provided callback invoked when the session's cwd changes (currently only via
* `enter_worktree`/`exit_worktree`), so the UI can update its live cwd display, /undo resolution,
* git-info refresh, and @mention resolution. Absent in non-UI contexts (a headless sub-agent
* switching its own cwd just has no UI to notify). */
onCwdChange?: (newCwd: string) => void;
/** App-provided callback invoked when the worktree session is entered/exited, so the App can keep a
* ref that survives a /model switch (which recreates the session) and re-attach the worktree
* tracking to the new session. Absent in non-UI contexts. */
onWorktreeChange?: (worktree: { dir: string; branch: string; originalCwd: string } | null) => void;
}
export function createSession(
@@ -117,18 +159,19 @@ export function createSession(
confirm: ConfirmFn,
mode: ToolCallMode,
tools: ToolDef[] = TOOLS,
contextWindow?: number,
contextWindow: number = DEFAULT_CONTEXT_WINDOW,
contextWindowIsEstimate: boolean = true,
maxIterations: number = DEFAULT_MAX_ITERATIONS,
autoCompactThreshold: number = DEFAULT_AUTO_COMPACT_THRESHOLD,
projectInstructions: string | null = null,
isLocal?: boolean,
subagentMaxIterations: number = DEFAULT_SUBAGENT_MAX_ITERATIONS,
userMemory: string | null = null,
subagentMaxDepth: number = DEFAULT_SUBAGENT_MAX_DEPTH,
permissionRules: PermissionRule[] = [],
): Session {
const resolvedIsLocal = isLocal ?? false;
const resolvedContextWindow = contextWindow ?? (resolvedIsLocal ? DEFAULT_CONTEXT_WINDOW_LOCAL : DEFAULT_CONTEXT_WINDOW_CLOUD);
const toolset = buildToolSet(tools);
const messages: ChatCompletionMessageParam[] = [
{ role: "system", content: buildSystemPrompt(toolset.tools, mode, projectInstructions, resolvedIsLocal) },
{ role: "system", content: buildSystemPrompt(toolset.tools, mode, projectInstructions, userMemory) },
];
return {
id: randomUUID(),
@@ -137,14 +180,16 @@ export function createSession(
model,
cwd,
mode,
isLocal: resolvedIsLocal,
messages,
maxIterations,
permissions: new PermissionManager(),
subagentMaxIterations,
subagentMaxDepth,
permissions: new PermissionManager(permissionRules),
confirm,
confirmMutex: { chain: Promise.resolve() },
toolset,
subAgentDepth: 0,
contextWindow: resolvedContextWindow,
contextWindow,
contextWindowIsEstimate,
lastContextTokens: estimateTokens(messages),
lastContextTokensIsEstimate: true,
@@ -153,9 +198,11 @@ export function createSession(
activeBackground: null,
mutationCommitLength: null,
projectInstructions,
todos: [],
userMemory,
taskStore: new TaskStore(),
mutationGate: Promise.resolve(),
lastEdit: null,
subAgentSessions: new Map(),
namedAgents: new Map(),
};
}
@@ -167,35 +214,45 @@ export function createSessionFromRecord(
cwd: string,
confirm: ConfirmFn,
tools: ToolDef[] = TOOLS,
contextWindow?: number,
contextWindow: number = DEFAULT_CONTEXT_WINDOW,
contextWindowIsEstimate: boolean = true,
maxIterations: number = DEFAULT_MAX_ITERATIONS,
autoCompactThreshold: number = DEFAULT_AUTO_COMPACT_THRESHOLD,
projectInstructions: string | null = null,
isLocal?: boolean,
subagentMaxIterations: number = DEFAULT_SUBAGENT_MAX_ITERATIONS,
userMemory: string | null = null,
subagentMaxDepth: number = DEFAULT_SUBAGENT_MAX_DEPTH,
permissionRules: PermissionRule[] = [],
): Session {
const resolvedIsLocal = isLocal ?? false;
const resolvedContextWindow = contextWindow ?? (resolvedIsLocal ? DEFAULT_CONTEXT_WINDOW_LOCAL : DEFAULT_CONTEXT_WINDOW_CLOUD);
const toolset = buildToolSet(tools);
// Build the system prompt with the *restored* permission mode (not the default) so a resumed
// plan-mode session gets plan instructions in its prompt from the first request, rather than
// only learning it's in plan mode from tool-rejection errors.
const messages: ChatCompletionMessageParam[] = [
{ role: "system", content: buildSystemPrompt(toolset.tools, record.mode, projectInstructions, resolvedIsLocal) },
{ role: "system", content: buildSystemPrompt(toolset.tools, record.mode, projectInstructions, userMemory, record.permissionMode) },
...record.messages,
];
const session: Session = {
const permissions = new PermissionManager(permissionRules);
// Restore the saved permission mode so plan/auto-edit/auto-accept survive a resume instead of
// always resetting to default. Older saved sessions omit the field → default.
if (record.permissionMode) permissions.setMode(record.permissionMode);
return {
id: record.id,
createdAt: record.createdAt,
client,
model: record.model,
cwd,
mode: record.mode,
isLocal: resolvedIsLocal,
messages,
maxIterations,
permissions: new PermissionManager(),
subagentMaxIterations,
subagentMaxDepth,
permissions,
confirm,
confirmMutex: { chain: Promise.resolve() },
toolset,
subAgentDepth: 0,
contextWindow: resolvedContextWindow,
contextWindow,
contextWindowIsEstimate,
lastContextTokens: estimateTokens(messages),
lastContextTokensIsEstimate: true,
@@ -204,19 +261,12 @@ export function createSessionFromRecord(
activeBackground: null,
mutationCommitLength: null,
projectInstructions,
todos: [],
taskStore: TaskStore.fromJSON(record.tasks ?? { seq: 0, tasks: [] }),
mutationGate: Promise.resolve(),
userMemory,
taskStore: new TaskStore(),
lastEdit: null,
subAgentSessions: new Map(),
namedAgents: new Map(),
};
// Restore session-allowed tools from the saved record, so /perm approvals survive resume.
if (record.allowedTools) {
for (const toolName of record.allowedTools) {
session.permissions.allowForSession(toolName);
}
}
return session;
}
export function toSessionRecord(session: Session, baseURL: string): SessionRecord {
@@ -228,9 +278,8 @@ export function toSessionRecord(session: Session, baseURL: string): SessionRecor
baseURL,
model: session.model,
mode: session.mode,
permissionMode: session.permissions.getMode(),
messages: session.messages.slice(1),
allowedTools: session.permissions.listAllowed(),
tasks: session.taskStore.toJSON(),
};
}
@@ -238,32 +287,32 @@ export function resetSession(session: Session): void {
session.messages = [session.messages[0] as ChatCompletionMessageParam];
session.lastContextTokens = estimateTokens(session.messages);
session.lastContextTokensIsEstimate = true;
}
/** Removes the last complete user turn (the user message + all subsequent assistant/tool messages
* up to the next user message or the end of history). Returns the number of messages removed,
* or 0 if there's no user message to undo (only the system prompt remains). This is a soft undo \u2014
* filesystem changes from tool calls are NOT rolled back, but the model will no longer see the
* removed context, so it won't repeat those actions. */
export function undoLastTurn(session: Session): number {
// Walk backwards from the end to find the last user message.
let lastUserIdx = -1;
for (let i = session.messages.length - 1; i >= 1; i--) {
if (session.messages[i]!.role === "user") {
lastUserIdx = i;
break;
}
}
if (lastUserIdx === -1) return 0; // No user messages to undo.
const removed = session.messages.length - lastUserIdx;
session.messages.length = lastUserIdx;
session.lastContextTokens = estimateTokens(session.messages);
session.lastContextTokensIsEstimate = true;
return removed;
// Drop any resumable sub-agent sessions — their context references the old conversation and would
// be stale after a /clear.
session.subAgentSessions.clear();
// Drop the name → agentId index over those sessions too.
session.namedAgents.clear();
}
export function setMode(session: Session, mode: ToolCallMode): void {
session.mode = mode;
session.messages[0] = { role: "system", content: buildSystemPrompt(session.toolset.tools, mode, session.projectInstructions, session.isLocal) };
// Re-apply the current permission mode so plan-mode instructions survive a tool-call-mode switch
// (native ↔ fallback) rather than being dropped from the rebuilt prompt.
session.messages[0] = {
role: "system",
content: buildSystemPrompt(session.toolset.tools, mode, session.projectInstructions, session.userMemory, session.permissions.getMode()),
};
}
/** Switch the session's permission mode and rebuild the system prompt so the model is told about
* the new mode (e.g. entering plan mode injects the plan-only research instructions). This is the
* permission-mode counterpart to setMode (which handles tool-call mode). UI paths that change the
* permission mode (/perm, Shift+Tab) should call this instead of `session.permissions.setMode` alone,
* which would leave the prompt stale. */
export function setPermissionMode(session: Session, mode: PermissionMode): void {
session.permissions.setMode(mode);
session.messages[0] = {
role: "system",
content: buildSystemPrompt(session.toolset.tools, session.mode, session.projectInstructions, session.userMemory, mode),
};
}
+54 -120
View File
@@ -1,131 +1,65 @@
import { describe, it, expect } from "vitest";
import { describe, expect, it } from "vitest";
import { z } from "zod";
import { buildSystemPrompt, } from "./systemPrompt.js";
import { createSession, setMode, setPermissionMode } from "./session.js";
import { buildToolSet } from "../tools/toolset.js";
import { FALLBACK_TOOL_INSTRUCTIONS } from "../toolcalling/fallbackPrompt.js";
import type { ToolDef } from "../tools/types.js";
import { buildSystemPrompt } from "./systemPrompt.js";
const dummyTool = (name: string, mutating: boolean): ToolDef => ({
name,
description: `Tool ${name} for testing`,
const dummyTool: ToolDef = {
name: "noop",
description: "does nothing",
schema: z.object({}),
mutating,
handler: async () => null,
mutating: false,
handler: async () => ({ ok: true }),
};
const toolset = buildToolSet([dummyTool]);
function systemPromptOf(session: { messages: { content?: unknown }[] }): string {
return String(session.messages[0]!.content);
}
describe("buildSystemPrompt plan-mode injection", () => {
it("omits plan instructions in the default mode", () => {
const prompt = buildSystemPrompt(toolset.tools, "native");
expect(prompt).not.toContain("Plan mode is ACTIVE");
});
it("injects plan instructions when permissionMode is 'plan'", () => {
const prompt = buildSystemPrompt(toolset.tools, "native", null, null, "plan");
expect(prompt).toContain("Plan mode is ACTIVE");
expect(prompt).toContain("present a concrete implementation plan");
});
it("does not inject plan instructions for auto-edit/auto-accept", () => {
expect(buildSystemPrompt(toolset.tools, "native", null, null, "auto-edit")).not.toContain("Plan mode is ACTIVE");
expect(buildSystemPrompt(toolset.tools, "native", null, null, "auto-accept")).not.toContain("Plan mode is ACTIVE");
});
});
describe("buildSystemPrompt", () => {
const readTool = dummyTool("read_file", false);
const writeTool = dummyTool("write_file", true);
describe("setPermissionMode / setMode prompt rebuild", () => {
// A fake OpenAI client — these tests never make requests, createSession just needs a client.
const fakeClient = {} as any;
it("includes tool names grouped by category", () => {
const prompt = buildSystemPrompt([readTool, writeTool], "native");
expect(prompt).toContain("read_file");
expect(prompt).toContain("write_file");
expect(prompt).toContain("Read-only");
expect(prompt).toContain("Mutating");
it("setPermissionMode('plan') rebuilds the system prompt with plan instructions, and back to default removes them", () => {
const session = createSession(fakeClient, "m", process.cwd(), async () => "once", "native", [dummyTool]);
expect(systemPromptOf(session)).not.toContain("Plan mode is ACTIVE");
setPermissionMode(session, "plan");
expect(systemPromptOf(session)).toContain("Plan mode is ACTIVE");
setPermissionMode(session, "default");
expect(systemPromptOf(session)).not.toContain("Plan mode is ACTIVE");
});
it("includes core principles for local models", () => {
const prompt = buildSystemPrompt([readTool], "native", null, true);
expect(prompt).toContain("Inspect before answering");
expect(prompt).toContain("Prefer small, targeted edits");
expect(prompt).toContain("Recovery over retry");
expect(prompt).toContain("Respect confirmation");
});
it("includes core principles for cloud models", () => {
const prompt = buildSystemPrompt([readTool], "native", null, false);
expect(prompt).toContain("Inspect before answering");
expect(prompt).toContain("Prefer small, targeted edits");
expect(prompt).toContain("Recovery over retry");
expect(prompt).toContain("Respect confirmation");
});
it("includes tool usage guide", () => {
const prompt = buildSystemPrompt([readTool], "native");
expect(prompt).toContain("read_file");
expect(prompt).toContain("edit_file");
expect(prompt).toContain("bash");
expect(prompt).toContain("agent");
});
it("includes local model guidance when isLocal is true", () => {
const prompt = buildSystemPrompt([readTool], "native", null, true);
expect(prompt).toContain("Working with local models");
expect(prompt).toContain("Tool-call formatting can be unreliable");
expect(prompt).toContain("Empty or malformed responses can happen");
});
it("excludes local model guidance when isLocal is false (cloud)", () => {
const prompt = buildSystemPrompt([readTool], "native", null, false);
expect(prompt).not.toContain("Working with local models");
expect(prompt).not.toContain("Tool-call formatting can be unreliable");
expect(prompt).not.toContain("Empty or malformed responses can happen");
});
it("defaults to cloud prompt when isLocal is not specified", () => {
const prompt = buildSystemPrompt([readTool], "native");
expect(prompt).toContain("coding assistant");
expect(prompt).not.toContain("Working with local models");
});
it("includes safety guidelines for both local and cloud", () => {
const localPrompt = buildSystemPrompt([readTool], "native", null, true);
const cloudPrompt = buildSystemPrompt([readTool], "native", null, false);
expect(localPrompt).toContain(".git");
expect(localPrompt).toContain("destructive");
expect(cloudPrompt).toContain(".git");
expect(cloudPrompt).toContain("destructive");
});
it("includes fallback instructions when mode is fallback", () => {
const prompt = buildSystemPrompt([readTool], "fallback");
expect(prompt).toContain("tool_call");
expect(prompt.toLowerCase()).toContain("fallback");
});
it("includes native mode instructions when mode is native (local)", () => {
const prompt = buildSystemPrompt([readTool], "native", null, true);
expect(prompt).toContain("native tool-call mode");
});
it("includes native mode instructions when mode is native (cloud)", () => {
const prompt = buildSystemPrompt([readTool], "native", null, false);
expect(prompt).toContain("native tool-call mode");
});
it("appends project instructions", () => {
const prompt = buildSystemPrompt([readTool], "native", "Always use TypeScript strict mode.");
expect(prompt).toContain("Always use TypeScript strict mode.");
// Project instructions should be at the end
const idx = prompt.indexOf("Always use TypeScript strict mode.");
const safetyIdx = prompt.indexOf("## Safety");
expect(idx).toBeGreaterThan(safetyIdx);
});
it("works without project instructions", () => {
const prompt = buildSystemPrompt([readTool], "native", null);
expect(prompt).not.toContain("Project instructions");
});
it("handles empty tool list", () => {
const prompt = buildSystemPrompt([], "native");
expect(prompt).toContain("Available tools");
expect(prompt).toContain("Core principles");
});
it("handles all read-only tools", () => {
const tools = [dummyTool("read_file", false), dummyTool("grep", false), dummyTool("definition", false)];
const prompt = buildSystemPrompt(tools, "native");
// Tool list section should only have Read-only
const toolSection = prompt.split("## Core principles")[0];
expect(toolSection).toContain("Read-only: read_file, grep, definition");
expect(toolSection).not.toContain("Mutating");
});
it("handles all mutating tools", () => {
const tools = [dummyTool("write_file", true), dummyTool("edit_file", true)];
const prompt = buildSystemPrompt(tools, "native");
const toolSection = prompt.split("## Core principles")[0];
expect(toolSection).toContain("Mutating (requires confirmation): write_file, edit_file");
expect(toolSection).not.toContain("Read-only");
it("setMode (tool-call mode switch) preserves plan instructions while plan mode is active", () => {
const session = createSession(fakeClient, "m", process.cwd(), async () => "once", "native", [dummyTool]);
setPermissionMode(session, "plan");
// Switch tool-call mode to fallback while still in plan mode — the rebuilt prompt must carry
// BOTH the fallback tool instructions and the plan instructions.
setMode(session, "fallback");
const prompt = systemPromptOf(session);
expect(prompt).toContain(FALLBACK_TOOL_INSTRUCTIONS);
expect(prompt).toContain("Plan mode is ACTIVE");
});
});
+72 -124
View File
@@ -1,144 +1,92 @@
import type { ToolCallMode } from "../backend/capabilityProbe.js";
import type { PermissionMode } from "../permissions/types.js";
import { FALLBACK_TOOL_INSTRUCTIONS } from "../toolcalling/fallbackPrompt.js";
import type { ToolDef } from "../tools/types.js";
function formatToolList(tools: ToolDef[]): string {
const readWrite = new Map<string, string[]>();
for (const t of tools) {
const category = t.mutating ? "Mutating (requires confirmation)" : "Read-only";
const list = readWrite.get(category) ?? [];
list.push(t.name);
readWrite.set(category, list);
}
const parts: string[] = [];
for (const [category, names] of readWrite) {
parts.push(`${category}: ${names.join(", ")}`);
}
return parts.join("\n");
}
/** Injected into the system prompt only while the session is in plan mode. Tells the model it must
* research read-only and present a plan rather than attempt changes (which would be blocked at the
* tool gate anyway). Without this, the model only learns it's in plan mode from tool-rejection
* errors, after it has already tried (and failed) to mutate. */
const PLAN_INSTRUCTIONS = `Plan mode is ACTIVE. In this mode:
- Do NOT call any mutating tool (write_file, edit_file, multi_edit, notebook_edit, git_commit, or a bash command that changes state). They are blocked and will return an error — that is expected.
- Explore read-only first: use grep, list_files, read_file, and git_status to fully understand the request and the code it touches.
- Then present a concrete implementation plan: the files you would change, the approach for each, and the key edits. Do not make the changes yet. Call the \`exit_plan_mode\` tool with the plan once you're ready for the user to approve it — on approval plan mode ends and you implement in this same turn; on rejection, refine and call it again. (If you can't call tools, present the plan as prose instead.)
- Keep the plan focused and actionable so the user can review it.
- Only ask a clarifying question if the request is still genuinely ambiguous after you've explored.`;
export function buildSystemPrompt(tools: ToolDef[], mode: ToolCallMode, projectInstructions?: string | null, isLocal?: boolean): string {
const toolList = formatToolList(tools);
export function buildSystemPrompt(
tools: ToolDef[],
mode: ToolCallMode,
projectInstructions?: string | null,
userMemory?: string | null,
permissionMode?: PermissionMode,
): string {
const toolList = tools.map((t) => `- ${t.name}: ${t.description}`).join("\n");
// Auto-detect: if not explicitly specified, use the cloud prompt by default.
// Callers (App.tsx) always pass isSmallLocalModel(baseURL, model) explicitly, so this
// default only affects tests or edge cases without a baseURL/model.
const useLocal = isLocal ?? false;
const base = useLocal
? buildLocalPrompt(toolList, mode)
: buildCloudPrompt(toolList, mode);
return projectInstructions ? `${base}\n\n${projectInstructions}` : base;
}
/** System prompt for local models (Ollama / LM Studio) — includes extra guidance about their
* limitations (unreliable tool-call formatting, occasional empty/malformed responses).
* Cloud models get a leaner prompt (buildCloudPrompt) that omits these assumptions. */
function buildLocalPrompt(toolList: string, mode: ToolCallMode): string {
return `You are a helpful coding assistant with access to tools for exploring and editing a codebase on the user's machine.
## Available tools
const base = `You are a capable local coding assistant with access to tools on the user's machine.
Available tools:
${toolList}
## Core principles
Core workflow:
1. Understand the user's goal before acting. Ask clarifying questions if the request is ambiguous or could destroy data.
2. Explore the codebase efficiently: use grep to locate symbols/patterns, list_files to understand structure, and read_file only on the files or page ranges you actually need.
3. Read files before editing them. Make the smallest change that solves the problem.
4. Test your assumptions when possible (run typecheck/tests, read related code, verify file contents).
5. Respond in plain text once you have enough information. Be concise; avoid restating obvious context.
1. **Inspect before answering.** Never guess file contents, function signatures, or directory structures — use read_file, list_files, grep, or definition to verify. Stale assumptions are worse than an extra tool call.
Editing guidelines:
- Prefer edit_file for small, targeted changes. Include enough surrounding context in old_string to make the match unique.
- Use write_file for new files or when you are replacing most of a file's content.
- Never invent file contents you haven't read; if unsure, read the file first.
- When edit_file fails with "old_string not found", re-read the file and try again with a more precise match.
2. **Prefer small, targeted edits.** Use edit_file (or multi_edit for several changes in one file) for surgical changes. Use write_file only for new files or full rewrites. edit_file requires old_string to match exactly — copy the exact text from the file (read it first), including indentation and blank lines.
Task tracking:
- For non-trivial multi-step work (3+ steps), create tasks with task_create so progress is visible. Mark a task in_progress when you start it and completed when done.
- Express dependencies with addBlocks/addBlockedBy (task ids) when one step must finish before another can start; check a task's blockedBy via task_get before starting it.
- Set owner when a sub-agent or teammate will claim a specific task. Keep task subjects short and imperative.
3. **One tool call per response in fallback mode.** If you are in fallback mode (see below), call at most one tool per response and wait for the result before proceeding. In native mode you may call multiple read-only tools in parallel.
Named teammates:
- When you'll send a sub-agent several messages across the conversation, give the 'agent' call a 'name' (e.g. "researcher", "implementer") to create a named teammate. Then continue it with send_message using that 'name' instead of tracking its agentId, and use list_teammates to see your roster.
- A named teammate runs in the shared working directory (a single, non-parallel delegation) and is resumable; parallel (worktree-isolated) delegations are fire-and-forget and ignore the name.
- Teammate names must be unique per session — re-using an existing name returns an error (don't clobber a teammate in use); address the existing one via send_message instead.
- Note: a named teammate still runs synchronously within your turn (it blocks until it answers). It is a stable, re-addressable handle, not a truly background process.
4. **Preserve existing style.** Match the surrounding code's indentation, naming conventions, quotes, and formatting. Don't reformat code outside the change scope.
Worktree sessions:
- To try changes without touching the main working tree, call enter_worktree (optionally with a name). It creates an isolated git worktree on a new branch at the current HEAD and switches your working directory into it — every file tool then operates there, and the worktree starts from the last commit so the user's uncommitted changes aren't carried over.
- When done, call exit_worktree. Use action "keep" to preserve the work on its branch (recoverable later via git), or "remove" to discard it entirely (deletes the worktree and the branch). A remove is refused if the worktree has uncommitted changes unless you pass discardChanges: true.
- You can only be in one worktree session at a time — exit before entering another. Don't switch models mid-worktree-session. /worktree shows the current worktree session.
5. **Keep answers concise.** When you have enough information, respond in plain text — don't pad with pleasantries or restated context. Code explanations should be brief and focused on the "why", not the "what" (the code already says what).
Bash guidelines:
- Destructive commands (rm, git push, git reset --hard, etc.) require explicit user confirmation via the tool's confirmation prompt.
- For long-running commands, increase timeout_ms or press Ctrl+B while the command is running to background it; then use bash_output with the returned jobId.
- Prefer git_status and git_commit for git work rather than raw git commands in bash.
6. **Recovery over retry.** If a tool call fails (edit_file "not found", bash non-zero exit, etc.), read the file or check the error output before retrying — don't repeat the same call. If edit_file suggests a closest match, use that text exactly.
Tool-use discipline:
- Call at most one tool at a time. Exception: when a single response delegates several independent sub-tasks to sub-agents, you may emit multiple 'agent' or 'agent__*' calls together — they run in parallel and their results come back in order. Writable parallel delegations (general-purpose, debugger, test-writer, or plugin agents) each run in an isolated throwaway git worktree at the last commit, so their file changes are discarded and only the returned answer matters. Read-only parallel delegations (explore, code-reviewer, planner) run in the shared working directory so they see current uncommitted state. Use parallel batches for research/review/planning; for implementation that must persist, delegate a single (sequential) agent call. Do not batch any other tool combinations. A single (non-parallel) delegation runs in the shared working directory, persists its edits, and is resumable — its result includes an agentId you can pass to send_message to continue it (refine the answer, ask a follow-up, or resume one that ran out of budget) without re-delegating from scratch. Parallel/worktree-isolated delegations are fire-and-forget and don't return an agentId.
- Mutating tools (write_file, edit_file, bash, git_commit) require user confirmation unless the user has changed the permission mode or chosen "Yes, and don't ask again this session".
- Use grep first when searching across many files; do not read_file dozens of files blindly.
- If a tool returns an error or empty result, adapt: refine your grep pattern, check the path, or ask the user.
- The session auto-continues a long tool chain internally up to a large per-turn budget. If it ever pauses because the budget was exhausted, briefly report progress and the user can continue with another message.
- For deterministic multi-agent orchestration the user explicitly asks for ('use a workflow', 'fan out agents'), use the 'workflow' tool with a JS script that calls agent/parallel/pipeline/phase/log. Don't invoke it for ordinary single delegations — that's just the 'agent' tool. When fanning out writable agents (general-purpose, debugger, test-writer, or plugin agents) in parallel, pass each agent() call opts.isolation: 'worktree' so each runs in its own throwaway git worktree and their file writes can't collide (edits are discarded; only the returned answer matters). Read-only agents (explore, code-reviewer, planner) should NOT use isolation — they run in the shared cwd to see current uncommitted state.
7. **Respect confirmation.** Mutating tools (write_file, edit_file, multi_edit, notebook_edit, bash, git_commit) require user confirmation — you will see a permission prompt. Plan your edits so the user sees a clear, concise preview.
Slash commands the user can type:
- /undo — roll back your most recent write_file or edit_file to its previous content.
- /summary — ask you to summarize the conversation without replacing history.
## Tool usage guide
When to ask the user:
- The request is ambiguous or underspecified.
- A change would delete or overwrite significant user data.
- You are about to push commits, force-delete branches, or run commands with side effects outside the project.
- You cannot complete the task with the available tools or information.
- For a decision that is genuinely the user's to make and that you can't resolve from the code or a sensible default, call the 'ask_user_question' tool with a short multiple-choice question (1-4 questions, 2-4 options each). Don't offload decisions you could make yourself — explore and pick a reasonable default first, and only ask when the choice truly changes what you do next.
- **read_file**: Start here. Use offset/limit for large files. Always read before editing.
- **list_files**: Explore directory structure. Supports glob patterns like "src/**/*.ts".
- **grep**: Search file contents. Prefer over read_file when you know what you're looking for.
- **definition / references / diagnostics**: LSP-powered code intelligence. Use definition to find where a symbol is declared, references for all usages, diagnostics for type errors.
- **edit_file**: For small changes to existing files. old_string must match exactly — include enough surrounding context to be unique. On mismatch, the tool suggests the closest similar text.
- **multi_edit**: Apply several edits to the same file in one call. Each edit sees the result of previous edits, so adjust old_string for context shifts.
- **write_file**: For new files or complete rewrites. Overwrites the entire file — use with care.
- **bash**: Run shell commands. Prefer targeted tools (grep, definition) over broad shell commands when possible. Use timeout_ms for long-running commands. Background with Ctrl+B for very long commands.
- **git_status / git_commit**: Inspect repo state and commit changes. Always check status before committing.
- **web_search / web_fetch**: Look up information not in the local codebase. For API docs, error messages, or unfamiliar libraries.
- **agent**: Delegate a sub-task to a focused sub-agent. Good for researching many files in parallel. Sub-agents cannot spawn further sub-agents.
- **task_create / task_list / task_get / task_update**: Track structured work items with dependencies. Use for multi-step tasks (3+ steps) so progress is visible.
- **todo_write**: Simple checklist for progress tracking. Good for linear step-by-step work.
Personal memory:
- A lightweight index of saved memory facts is included below when present; each line is 'name (type) — description'. Call the 'memory' tool with a name to read that fact's full body when its hook looks relevant to the current task.
- 'memory_write' saves (action='write') or removes (action='delete') a typed fact: pick a kebab-case 'name', a one-line 'description' (the recall hook), a 'type' (user/feedback/project/reference), and the 'content' body. Proactively save durable facts the user states — preferences, working-style feedback, corrections worth remembering. Do not save transient per-task notes.`;
## Working with local models
- **Tool-call formatting can be unreliable.** If you're in fallback mode, follow the tool_call format strictly. If native mode produces errors, the system will automatically retry with fallback parsing.
- **Empty or malformed responses can happen.** The system retries automatically, but if you see repeated failures, simplify your request.
- **Output length may be limited.** For large file generations, prefer edit_file over write_file when possible — it uses fewer output tokens.
## Fallback mode
${mode === "fallback" ? FALLBACK_TOOL_INSTRUCTIONS : "You are in native tool-call mode. Call tools using the standard function-calling format. You may call multiple read-only tools in parallel, but mutating tools are always run sequentially."}
## Safety
- Do not modify .git directories or other version-control internals.
- Do not delete large sections of code without clear justification and user confirmation.
- When running bash commands, prefer read-only inspections (ls, cat, git status) over destructive operations (rm, git reset --hard).
- If unsure about a destructive action, ask the user first rather than proceeding.`;
const withToolMode = mode === "fallback" ? `${base}\n\n${FALLBACK_TOOL_INSTRUCTIONS}` : base;
const withPlan = permissionMode === "plan" ? `${withToolMode}\n\n${PLAN_INSTRUCTIONS}` : withToolMode;
const withProject = projectInstructions ? `${withPlan}\n\n${projectInstructions}` : withPlan;
return userMemory ? `${withProject}\n\n${userMemory}` : withProject;
}
/** System prompt for cloud models (large context window, reliable tool calls, no local-model quirks).
* Leaner than the local prompt — skips the "Working with local models" section entirely and uses
* a more direct tone, since cloud models don't need hand-holding about their own limitations. */
function buildCloudPrompt(toolList: string, mode: ToolCallMode): string {
return `You are a coding assistant with access to tools for exploring and editing a codebase on the user's machine.
## Available tools
${toolList}
## Core principles
1. **Inspect before answering.** Never guess file contents, function signatures, or directory structures — use read_file, list_files, grep, or definition to verify. Stale assumptions are worse than an extra tool call.
2. **Prefer small, targeted edits.** Use edit_file (or multi_edit for several changes in one file) for surgical changes. Use write_file only for new files or full rewrites. edit_file requires old_string to match exactly — copy the exact text from the file (read it first), including indentation and blank lines.
3. **Preserve existing style.** Match the surrounding code's indentation, naming conventions, quotes, and formatting. Don't reformat code outside the change scope.
4. **Keep answers concise.** When you have enough information, respond in plain text — don't pad with pleasantries or restated context. Code explanations should be brief and focused on the "why", not the "what" (the code already says what).
5. **Recovery over retry.** If a tool call fails (edit_file "not found", bash non-zero exit, etc.), read the file or check the error output before retrying — don't repeat the same call. If edit_file suggests a closest match, use that text exactly.
6. **Respect confirmation.** Mutating tools (write_file, edit_file, multi_edit, notebook_edit, bash, git_commit) require user confirmation — you will see a permission prompt. Plan your edits so the user sees a clear, concise preview.
## Tool usage guide
- **read_file**: Start here. Use offset/limit for large files. Always read before editing.
- **list_files**: Explore directory structure. Supports glob patterns like "src/**/*.ts".
- **grep**: Search file contents. Prefer over read_file when you know what you're looking for.
- **definition / references / diagnostics**: LSP-powered code intelligence. Use definition to find where a symbol is declared, references for all usages, diagnostics for type errors.
- **edit_file**: For small changes to existing files. old_string must match exactly — include enough surrounding context to be unique. On mismatch, the tool suggests the closest similar text.
- **multi_edit**: Apply several edits to the same file in one call. Each edit sees the result of previous edits, so adjust old_string for context shifts.
- **write_file**: For new files or complete rewrites. Overwrites the entire file — use with care.
- **bash**: Run shell commands. Prefer targeted tools (grep, definition) over broad shell commands when possible. Use timeout_ms for long-running commands. Background with Ctrl+B for very long commands.
- **git_status / git_commit**: Inspect repo state and commit changes. Always check status before committing.
- **web_search / web_fetch**: Look up information not in the local codebase. For API docs, error messages, or unfamiliar libraries.
- **agent**: Delegate a sub-task to a focused sub-agent. Good for researching many files in parallel. Sub-agents cannot spawn further sub-agents.
- **task_create / task_list / task_get / task_update**: Track structured work items with dependencies. Use for multi-step tasks (3+ steps) so progress is visible.
- **todo_write**: Simple checklist for progress tracking. Good for linear step-by-step work.
## ${mode === "fallback" ? "Fallback mode" : "Tool calling"}
${mode === "fallback" ? FALLBACK_TOOL_INSTRUCTIONS : "You are in native tool-call mode. Call tools using the standard function-calling format. You may call multiple read-only tools in parallel, but mutating tools are always run sequentially."}
## Safety
- Do not modify .git directories or other version-control internals.
- Do not delete large sections of code without clear justification and user confirmation.
- When running bash commands, prefer read-only inspections (ls, cat, git status) over destructive operations (rm, git reset --hard).
- If unsure about a destructive action, ask the user first rather than proceeding.`;
}
+4 -40
View File
@@ -7,7 +7,7 @@ import type { ToolCallMode } from "./capabilityProbe.js";
const paths = envPaths("locode", { suffix: "" });
const cacheFile = path.join(paths.config, "model-capabilities.json");
type Cache = Record<string, { mode: ToolCallMode; cachedAt?: number }>;
type Cache = Record<string, ToolCallMode>;
function keyFor(baseURL: string, model: string): string {
return `${baseURL}::${model}`;
@@ -23,19 +23,7 @@ function load(): Cache {
// Always re-read from disk if possible, so concurrent processes' writes aren't overwritten.
if (existsSync(cacheFile)) {
try {
const raw = JSON.parse(readFileSync(cacheFile, "utf-8")) as Record<string, unknown>;
// Migrate legacy format: bare string values ("native" | "fallback") become { mode, cachedAt }.
const cache: Cache = {};
for (const [k, v] of Object.entries(raw)) {
if (typeof v === "string" && (v === "native" || v === "fallback")) {
// Legacy entry — no cachedAt, so it can be re-validated on next probe.
cache[k] = { mode: v };
} else if (typeof v === "object" && v !== null && "mode" in v) {
cache[k] = v as Cache[string];
}
// Silently drop unrecognized entries.
}
memoryCache = cache;
memoryCache = JSON.parse(readFileSync(cacheFile, "utf-8")) as Cache;
return memoryCache;
} catch {
memoryCache = {};
@@ -53,36 +41,12 @@ function save(cache: Cache): void {
void writeFileAtomic(cacheFile, JSON.stringify(cache, null, 2));
}
/** How long a cached tool-call mode detection stays fresh before locode re-probes. A model's
* tool-call capability rarely changes, but a transient probe failure (network timeout, 5xx)
* can leave a stale "fallback" entry that permanently disables native tool calls. The TTL
* ensures periodic re-validation. Set to 0 via LOCODE_CAPABILITY_CACHE_TTL_DAYS=0 to force
* a re-probe every session. */
const DEFAULT_CACHE_TTL_DAYS = 30;
function resolveCacheTtlDays(): number {
const envValue = Number(process.env.LOCODE_CAPABILITY_CACHE_TTL_DAYS);
if (Number.isFinite(envValue) && envValue >= 0 && envValue <= 365) return envValue;
return DEFAULT_CACHE_TTL_DAYS;
}
function isStale(entry: { cachedAt?: number }, ttlDays: number): boolean {
if (ttlDays <= 0) return true;
if (typeof entry.cachedAt !== "number") return true; // legacy entry — always re-probe
const ageMs = Date.now() - entry.cachedAt;
return ageMs > ttlDays * 24 * 60 * 60 * 1000;
}
export function getCachedMode(baseURL: string, model: string): ToolCallMode | undefined {
const entry = load()[keyFor(baseURL, model)];
if (entry === undefined) return undefined;
// Treat stale or legacy entries as a miss so the backend is re-probed.
if (isStale(entry, resolveCacheTtlDays())) return undefined;
return entry.mode;
return load()[keyFor(baseURL, model)];
}
export function setCachedMode(baseURL: string, model: string, mode: ToolCallMode): void {
const cache = load();
cache[keyFor(baseURL, model)] = { mode, cachedAt: Date.now() };
cache[keyFor(baseURL, model)] = mode;
save(cache);
}
+2 -6
View File
@@ -1,5 +1,5 @@
import OpenAI from "openai";
import { resolveMaxRetries, resolveRequestTimeoutMs } from "../config/config.js";
import { resolveRequestTimeoutMs } from "../config/config.js";
import type { AppConfig } from "../config/types.js";
export function makeClient(cfg: AppConfig): OpenAI {
@@ -15,10 +15,6 @@ export function makeClient(cfg: AppConfig): OpenAI {
// requests behind a concurrency limit (e.g. Ollama's OLLAMA_NUM_PARALLEL) can legitimately take
// longer than the 180s default to even start serving a request under contention.
timeout: resolveRequestTimeoutMs(),
// Configurable retries on transient failures (connection errors, 429, 5xx) with exponential
// backoff. Defaults to 0 (fail immediately) to preserve the old behavior, since a local
// backend's slow response usually means the model is stuck rather than a transient blip — but
// raise via `maxRetries` / LOCODE_MAX_RETRIES for setups with occasional connection drops.
maxRetries: resolveMaxRetries(),
maxRetries: 0,
});
}
+4 -9
View File
@@ -61,14 +61,9 @@ async function detectLmStudioContextWindow(baseURL: string, model: string): Prom
* assume which one is actually running behind an OpenAI-compatible baseURL. Returns null (rather
* than guessing) if neither responds usefully — callers should fall back to a configured default. */
export async function detectContextWindow(baseURL: string, model: string): Promise<number | null> {
// Try both backends in parallel to halve detection latency.
const [ollama, lmStudio] = await Promise.allSettled([
detectOllamaContextWindow(baseURL, model),
detectLmStudioContextWindow(baseURL, model),
]);
if (ollama.status === "fulfilled" && ollama.value !== null) return ollama.value;
if (lmStudio.status === "fulfilled" && lmStudio.value !== null) return lmStudio.value;
return null;
const ollama = await detectOllamaContextWindow(baseURL, model);
if (ollama !== null) return ollama;
return detectLmStudioContextWindow(baseURL, model);
}
export interface ResolvedContextWindow {
@@ -90,5 +85,5 @@ export async function resolveContextWindow(baseURL: string, model: string): Prom
return { value: detected, isEstimate: false };
}
return { value: resolveContextWindowDefault(baseURL, model), isEstimate: true };
return { value: resolveContextWindowDefault(), isEstimate: true };
}
+3 -44
View File
@@ -4,8 +4,6 @@ import envPaths from "env-paths";
import { afterEach, beforeEach, describe, expect, it } from "vitest";
import { getCachedContextWindow } from "./contextWindowCache.js";
const KEY = "http://localhost:11434/v1::ttl-model";
const cacheFile = path.join(envPaths("locode", { suffix: "" }).config, "context-windows.json");
describe("contextWindowCache", () => {
@@ -21,18 +19,16 @@ describe("contextWindowCache", () => {
} else if (existsSync(cacheFile)) {
rmSync(cacheFile);
}
delete process.env.LOCODE_CONTEXT_WINDOW_CACHE_TTL_DAYS;
});
it("reads a well-formed cached entry", () => {
// Written directly (rather than via setCachedContextWindow, whose write is fire-and-forget
// async and would race this file-backed cache's always-read-from-disk load()) so the test is
// deterministic and can't leak a pending write past its own afterEach cleanup. Includes a
// fresh cachedAt so the TTL check treats it as current.
// deterministic and can't leak a pending write past its own afterEach cleanup.
mkdirSync(path.dirname(cacheFile), { recursive: true });
writeFileSync(cacheFile, JSON.stringify({ "http://localhost:11434/v1::test-model": { value: 32768, isEstimate: false, cachedAt: Date.now() } }), "utf-8");
writeFileSync(cacheFile, JSON.stringify({ "http://localhost:11434/v1::test-model": { value: 32768, isEstimate: false } }), "utf-8");
expect(getCachedContextWindow("http://localhost:11434/v1", "test-model")).toEqual({ value: 32768, isEstimate: false, cachedAt: expect.any(Number) });
expect(getCachedContextWindow("http://localhost:11434/v1", "test-model")).toEqual({ value: 32768, isEstimate: false });
});
it("treats a legacy bare-number cache entry as a miss instead of returning {value: undefined}", () => {
@@ -46,41 +42,4 @@ describe("contextWindowCache", () => {
const result = getCachedContextWindow("http://localhost:11434/v1", "legacy-model");
expect(result).toBeUndefined();
});
it("treats an entry without cachedAt as expired (re-detect)", () => {
mkdirSync(path.dirname(cacheFile), { recursive: true });
writeFileSync(cacheFile, JSON.stringify({ [KEY]: { value: 32768, isEstimate: false } }), "utf-8");
// No cachedAt field — legacy entry from before TTL was added; should be a miss.
expect(getCachedContextWindow("http://localhost:11434/v1", "ttl-model")).toBeUndefined();
});
it("treats a fresh entry (recent cachedAt) as a hit", () => {
mkdirSync(path.dirname(cacheFile), { recursive: true });
writeFileSync(cacheFile, JSON.stringify({ [KEY]: { value: 32768, isEstimate: false, cachedAt: Date.now() } }), "utf-8");
expect(getCachedContextWindow("http://localhost:11434/v1", "ttl-model")).toEqual({ value: 32768, isEstimate: false, cachedAt: expect.any(Number) });
});
it("treats an old entry (cachedAt beyond TTL) as a miss", () => {
mkdirSync(path.dirname(cacheFile), { recursive: true });
// 30 days ago, default TTL is 7 days — stale.
const old = Date.now() - 30 * 24 * 60 * 60 * 1000;
writeFileSync(cacheFile, JSON.stringify({ [KEY]: { value: 32768, isEstimate: false, cachedAt: old } }), "utf-8");
expect(getCachedContextWindow("http://localhost:11434/v1", "ttl-model")).toBeUndefined();
});
it("respects a configured TTL of 0 (always re-detect)", () => {
process.env.LOCODE_CONTEXT_WINDOW_CACHE_TTL_DAYS = "0";
mkdirSync(path.dirname(cacheFile), { recursive: true });
writeFileSync(cacheFile, JSON.stringify({ [KEY]: { value: 32768, isEstimate: false, cachedAt: Date.now() } }), "utf-8");
expect(getCachedContextWindow("http://localhost:11434/v1", "ttl-model")).toBeUndefined();
});
it("respects a longer configured TTL", () => {
process.env.LOCODE_CONTEXT_WINDOW_CACHE_TTL_DAYS = "365";
mkdirSync(path.dirname(cacheFile), { recursive: true });
// 30 days ago, but TTL is now 365 days — fresh.
const old = Date.now() - 30 * 24 * 60 * 60 * 1000;
writeFileSync(cacheFile, JSON.stringify({ [KEY]: { value: 32768, isEstimate: false, cachedAt: old } }), "utf-8");
expect(getCachedContextWindow("http://localhost:11434/v1", "ttl-model")).toBeDefined();
});
});
+2 -28
View File
@@ -6,23 +6,9 @@ import { writeFileAtomic } from "../utils/writeFileAtomic.js";
const paths = envPaths("locode", { suffix: "" });
const cacheFile = path.join(paths.config, "context-windows.json");
/** How long a cached context-window detection stays fresh before locode re-detects it. A model's
* context window rarely changes, but a backend can be reconfigured (quantization swapped, a
* different model loaded under the same id, Ollama's `num_ctx` raised) — a TTL avoids pinning a
* stale value forever. Set to 0 to disable caching (re-detect every session). */
const DEFAULT_CACHE_TTL_DAYS = 7;
export function resolveCacheTtlDays(): number {
const envValue = Number(process.env.LOCODE_CONTEXT_WINDOW_CACHE_TTL_DAYS);
if (Number.isFinite(envValue) && envValue >= 0 && envValue <= 365) return envValue;
return DEFAULT_CACHE_TTL_DAYS;
}
export interface CachedContextWindow {
value: number;
isEstimate: boolean;
/** Unix epoch ms when this entry was cached. Absent on legacy entries (treated as expired). */
cachedAt?: number;
}
type Cache = Record<string, CachedContextWindow>;
@@ -53,17 +39,7 @@ function save(cache: Cache): void {
void writeFileAtomic(cacheFile, JSON.stringify(cache, null, 2));
}
/** Returns true when the entry is stale given the configured TTL. A TTL of 0 means "always
* re-detect", so every entry is stale; a missing `cachedAt` (legacy entry) is also stale. */
function isStale(entry: CachedContextWindow, ttlDays: number): boolean {
if (ttlDays <= 0) return true;
if (typeof entry.cachedAt !== "number") return true;
const ageMs = Date.now() - entry.cachedAt;
return ageMs > ttlDays * 24 * 60 * 60 * 1000;
}
export function getCachedContextWindow(baseURL: string, model: string): CachedContextWindow | undefined {
const ttlDays = resolveCacheTtlDays();
const entry = load()[keyFor(baseURL, model)];
if (entry === undefined) return undefined;
// An older locode version cached a bare number instead of { value, isEstimate }. Treat that
@@ -74,13 +50,11 @@ export function getCachedContextWindow(baseURL: string, model: string): CachedCo
if (typeof entry !== "object" || entry === null || typeof (entry as CachedContextWindow).value !== "number") {
return undefined;
}
// Expired entries are treated as a miss so the backend is re-queried and the entry refreshed.
if (isStale(entry, ttlDays)) return undefined;
return entry;
}
export function setCachedContextWindow(baseURL: string, model: string, contextWindow: CachedContextWindow): void {
const cache = load();
cache[keyFor(baseURL, model)] = { ...contextWindow, cachedAt: Date.now() };
cache[keyFor(baseURL, model)] = contextWindow;
save(cache);
}
}
+18 -41
View File
@@ -2,14 +2,13 @@ import { execa } from "execa";
import { existsSync, mkdirSync } from "node:fs";
import path from "node:path";
import { Command } from "commander";
import pkg from "../package.json" with { type: "json" };
import { makeClient } from "./backend/client.js";
import { ConfigError, resolveBackendConfig, resolveModel } from "./config/config.js";
import { configFilePath, loadStoredConfig, saveStoredConfig, type StoredConfig } from "./config/store.js";
import { loadMergedHooks, userHooksFilePath } from "./hooks/config.js";
import { loadMergedServers, removeUserServer, saveUserServer, userMcpFilePath } from "./mcp/config.js";
import { isHttpServerConfig } from "./mcp/types.js";
import { deleteSession, listSessions, loadSession, mostRecentSessionId, sessionsDir } from "./persistence/sessionStore.js";
import { deleteSession, listSessions, mostRecentSessionId, sessionsDir } from "./persistence/sessionStore.js";
import { addInstalledPlugin, loadInstalledPlugins, pluginsDir, removeInstalledPlugin } from "./plugins/config.js";
import { loadPlugin } from "./plugins/loader.js";
import { runInkApp } from "./ui/ink/index.js";
@@ -25,7 +24,7 @@ function addBackendOptions(cmd: Command): Command {
program
.name("locode")
.description("Agentic coding CLI for local models via Ollama and LM Studio")
.version(pkg.version);
.version("0.3.1");
addBackendOptions(program)
.option("-m, --model <name>", "model name as known to the backend")
@@ -115,54 +114,44 @@ configCmd
configCmd
.command("set <key> <value>")
.description(
"Persist a config value (backend, model, baseUrl, contextWindow, maxOutputTokens, maxIterations, autoCompactThreshold, requestTimeoutMs, subagentTimeoutMs, maxRetries, lspServers)",
)
.description("Persist a config value (backend, model, baseUrl, contextWindow, maxIterations, subagentMaxIterations, subagentMaxDepth, autoCompactThreshold, requestTimeoutMs, subagentTimeoutMs)")
.action((key: string, value: string) => {
if (
key !== "backend" &&
key !== "model" &&
key !== "baseUrl" &&
key !== "contextWindow" &&
key !== "maxOutputTokens" &&
key !== "maxIterations" &&
key !== "subagentMaxIterations" &&
key !== "subagentMaxDepth" &&
key !== "autoCompactThreshold" &&
key !== "requestTimeoutMs" &&
key !== "subagentTimeoutMs" &&
key !== "maxRetries" &&
key !== "lspServers"
key !== "subagentTimeoutMs"
) {
console.error(
`Unknown config key "${key}". Valid keys: backend, model, baseUrl, contextWindow, maxOutputTokens, maxIterations, autoCompactThreshold, requestTimeoutMs, subagentTimeoutMs, maxRetries, lspServers`,
`Unknown config key "${key}". Valid keys: backend, model, baseUrl, contextWindow, maxIterations, subagentMaxIterations, subagentMaxDepth, autoCompactThreshold, requestTimeoutMs, subagentTimeoutMs`,
);
process.exit(1);
}
const stored = loadStoredConfig();
if (key === "lspServers") {
// lspServers is a JSON object: { "<languageId>": { "command": "...", "args": [...], "extensions": [...] } }
let parsed: unknown;
try {
parsed = JSON.parse(value);
} catch {
console.error(`lspServers must be a JSON object, got invalid JSON: ${value}`);
process.exit(1);
}
if (typeof parsed !== "object" || parsed === null || Array.isArray(parsed)) {
console.error(`lspServers must be a JSON object keyed by language id, got: ${value}`);
process.exit(1);
}
stored.lspServers = parsed as Record<string, { command: string; args?: string[]; extensions?: string[] }>;
} else if (key === "contextWindow" || key === "maxIterations") {
if (key === "contextWindow" || key === "maxIterations") {
const n = Number(value);
if (!Number.isFinite(n) || n <= 0) {
console.error(`${key} must be a positive number, got "${value}".`);
process.exit(1);
}
stored[key] = n;
} else if (key === "maxOutputTokens") {
} else if (key === "subagentMaxIterations") {
const n = Number(value);
if (!Number.isFinite(n) || n < 256 || n > 1_000_000) {
console.error(`maxOutputTokens must be between 256 and 1000000, got "${value}".`);
if (!Number.isFinite(n) || n < 1 || n > 1000) {
console.error(`subagentMaxIterations must be between 1 and 1000, got "${value}".`);
process.exit(1);
}
stored[key] = n;
} else if (key === "subagentMaxDepth") {
const n = Number(value);
if (!Number.isFinite(n) || n < 0 || n > 10) {
console.error(`subagentMaxDepth must be between 0 and 10, got "${value}".`);
process.exit(1);
}
stored[key] = n;
@@ -187,13 +176,6 @@ configCmd
process.exit(1);
}
stored[key] = n;
} else if (key === "maxRetries") {
const n = Number(value);
if (!Number.isFinite(n) || n < 0 || n > 10) {
console.error(`maxRetries must be between 0 and 10, got "${value}".`);
process.exit(1);
}
stored[key] = n;
} else {
stored[key] = value;
}
@@ -230,11 +212,6 @@ sessionsCmd
.action((id: string) => {
if (deleteSession(id)) {
console.log(`Deleted session ${id}.`);
} else if (loadSession(id)) {
// The file exists but deleteSession() couldn't actually remove it (e.g. locked by another
// process) — a different situation from "no such session", so say so distinctly.
console.error(`Could not delete session "${id}" — the file may be in use by another process.`);
process.exit(1);
} else {
console.error(`No saved session found with id "${id}".`);
process.exit(1);
-60
View File
@@ -1,60 +0,0 @@
import { afterEach, beforeEach, describe, expect, it } from "vitest";
import { configureLanguageSpecs, _resetSpecsForTests, _specsForTests } from "./lspManager.js";
// configureLanguageSpecs mutates the module's LANGUAGE_SPECS (forward-only by design — production
// applies it once at startup). Tests restore the built-in defaults via _resetSpecsForTests so they
// stay independent, then assert the merged spec list through _specsForTests (no server spawned).
function specFor(ext: string) {
const specs = _specsForTests();
return specs.find((s) => s.extensions.includes(ext)) ?? null;
}
describe("configureLanguageSpecs", () => {
beforeEach(() => _resetSpecsForTests());
afterEach(() => _resetSpecsForTests());
it("leaves the built-in specs untouched for an empty override", () => {
configureLanguageSpecs({});
expect(_specsForTests().map((s) => s.languageId)).toEqual([
"typescript",
"python",
"go",
"rust",
"c",
]);
// C and C++ share one clangd spec (no separate "cpp" entry).
expect(specFor(".cpp")?.languageId).toBe("c");
expect(specFor(".h")?.languageId).toBe("c");
});
it("adds a brand-new language with extensions", () => {
configureLanguageSpecs({ java: { command: "jdtls", extensions: [".java"] } });
expect(specFor(".java")?.command).toBe("jdtls");
expect(specFor(".java")?.languageId).toBe("java");
});
it("ignores a new-language entry without extensions (can't route files to it)", () => {
configureLanguageSpecs({ ruby: { command: "solargraph" } });
expect(specFor(".rb")).toBeNull();
});
it("overrides a built-in server's command and args", () => {
configureLanguageSpecs({ typescript: { command: "my-tsserver", args: ["--stdio"] } });
expect(specFor(".ts")?.command).toBe("my-tsserver");
expect(specFor(".ts")?.args).toEqual(["--stdio"]);
});
it("keeps a built-in's extensions when an override omits them", () => {
configureLanguageSpecs({ python: { command: "basedpyright", args: ["--stdio"] } });
expect(specFor(".py")?.command).toBe("basedpyright");
expect(specFor(".pyi")?.languageId).toBe("python"); // extensions unchanged
});
it("rewrites a built-in language's extensions when provided", () => {
configureLanguageSpecs({ go: { command: "gopls", args: ["serve"], extensions: [".rs"] } });
// .rs now routes to "go", not "rust".
expect(specFor(".rs")?.languageId).toBe("go");
expect(specFor(".go")).toBeNull(); // .go no longer claimed by go
});
});
-452
View File
@@ -1,452 +0,0 @@
import { spawn, type ChildProcess } from "node:child_process";
import path from "node:path";
import { readFile as fsReadFile } from "node:fs/promises";
import {
createProtocolConnection,
DidChangeTextDocumentNotification,
DidOpenTextDocumentNotification,
DefinitionRequest,
ReferencesRequest,
type ProtocolConnection,
type TextDocumentIdentifier,
type Position,
type Location,
type Diagnostic,
} from "vscode-languageserver-protocol";
import { StreamMessageReader, StreamMessageWriter } from 'vscode-languageserver-protocol/node';
import { URI } from "vscode-uri";
import type { MarkupContent } from "vscode-languageserver-protocol";
/** LSP diagnostic messages can be either a plain string or a { kind, value } MarkupContent object.
* locode's tool surface deals in plain strings, so flatten either form to text. */
function messageToString(message: string | MarkupContent): string {
if (typeof message === "string") return message;
return message?.value ?? "";
}
/** A connected LSP server for one language, plus its child process so we can clean it up. */
interface LspHandle {
connection: ProtocolConnection;
child: ChildProcess;
languageId: string;
/** Open documents we've already sent didOpen for, so we send didChange (not didOpen) on edits. */
openDocs: Set<string>;
}
/** Maps a file extension to a language id (the LSP "languageId" string) and the server command to
* spawn for it. Only one server per language is ever spawned (lazy, on first use). A missing entry
* means locode has no built-in mapping — the user can still point a server at it via config in a
* future extension. The command is resolved on the PATH; if it isn't installed the spawn fails and
* the tool returns a clear "install X" error rather than a silent no-op. */
interface LanguageSpec {
languageId: string;
extensions: string[];
/** The server command (no args). Must be on PATH. */
command: string;
/** Args passed to the server command. */
args?: string[];
}
// The built-in language→server mappings. Mutable so `configureLanguageSpecs` can merge in user
// overrides/additions from config (see config.ts `lspServers`). One clangd spec covers both C
// and C++ — clangd handles both, and merging avoids spawning a second clangd for a mixed C/C++
// project (two servers keyed by separate languageIds would each index the same headers twice).
let LANGUAGE_SPECS: LanguageSpec[] = [
// TypeScript / JavaScript — `typescript-language-server` wraps tsserver and speaks LSP. The most
// common local-model codebase shape, so it's the first one locode wires up.
{
languageId: "typescript",
extensions: [".ts", ".tsx", ".mts", ".cts", ".js", ".jsx", ".mjs", ".cjs"],
command: "typescript-language-server",
args: ["--stdio"],
},
{
languageId: "python",
extensions: [".py", ".pyi"],
command: "pyright-langserver",
args: ["--stdio"],
},
{
languageId: "go",
extensions: [".go"],
command: "gopls",
args: ["serve"],
},
{
languageId: "rust",
extensions: [".rs"],
command: "rust-analyzer",
},
// C and C++ share clangd. The languageId is "c" (clangd treats .cpp/.hpp the same way);
// all C/C++ extensions route to the single clangd process.
{
languageId: "c",
extensions: [".c", ".h", ".cpp", ".cc", ".cxx", ".hpp", ".hh", ".hxx"],
command: "clangd",
},
];
/** Merge user-configured LSP server entries (from `locode config set lspServers`) into the
* built-in specs. An entry keyed by a built-in languageId overrides that spec's command/args
* and, if `extensions` is provided, which file extensions route to it. An entry keyed by a new
* languageId (e.g. "java", "ruby") adds a brand-new mapping — it MUST supply `extensions` so
* files can be routed to it. Call once at startup; idempotent against the built-in list.
*
* Entries missing a `command` are ignored (a server we can't spawn is useless), and entries for
* new ids without `extensions` are ignored too (no way to route files to them). */
export function configureLanguageSpecs(overrides: Record<string, { command: string; args?: string[]; extensions?: string[] }>): void {
const merged: LanguageSpec[] = LANGUAGE_SPECS.map((spec) => {
const ov = overrides[spec.languageId];
if (!ov) return spec;
return {
languageId: spec.languageId,
extensions: ov.extensions ?? spec.extensions,
command: ov.command,
args: ov.args,
};
});
for (const [languageId, ov] of Object.entries(overrides)) {
if (merged.some((s) => s.languageId === languageId)) continue; // already a built-in we overrode
if (!ov.command || !ov.extensions || ov.extensions.length === 0) continue;
merged.push({ languageId, extensions: ov.extensions, command: ov.command, args: ov.args });
}
LANGUAGE_SPECS = merged;
}
/** Picks the LanguageSpec for a file path, or null if no extension matches. */
function specForFile(filePath: string): LanguageSpec | null {
const ext = path.extname(filePath).toLowerCase();
if (!ext) return null;
return LANGUAGE_SPECS.find((s) => s.extensions.includes(ext)) ?? null;
}
/** A per-workspace (cwd) registry of live LSP servers, keyed by language id. One server per
* language per cwd — a second project gets its own manager (locode is single-session-per-process
* today, but keying on cwd keeps it correct if that ever changes). */
const handles = new Map<string, LspHandle>();
/** Convert an absolute filesystem path to an LSP file:// URI string. */
function toUri(absPath: string): string {
return URI.file(absPath).toString();
}
interface LocResult {
path: string;
line: number;
column: number;
}
function toLocation(loc: Location): LocResult {
return {
path: URI.parse(loc.uri).fsPath,
line: loc.range.start.line + 1,
column: loc.range.start.character + 1,
};
}
/** Spawns the LSP server for `spec`, initializes it, and returns a live handle. Throws a clear,
* actionable error if the server binary isn't on the PATH (the most common failure) so the tool
* can surface "install typescript-language-server" instead of an opaque spawn ENOENT. */
async function startServer(spec: LanguageSpec, cwd: string): Promise<LspHandle> {
let child: ChildProcess;
try {
// npm installs global CLI packages on Windows as .cmd/.ps1 shims, not raw .exe files — spawn()
// can't resolve those without shell:true, so a genuinely-installed server would otherwise ENOENT.
child = spawn(spec.command, spec.args ?? [], {
cwd,
stdio: ["pipe", "pipe", "pipe"],
shell: process.platform === "win32",
});
} catch (err) {
throw new Error(
`Could not start the LSP server "${spec.command}" for ${spec.languageId}. Is it installed and on your PATH? (${(err as Error).message})`,
);
}
// spawn() itself rarely throws synchronously — a missing binary (ENOENT) instead fires an
// async 'error' event on the child process. With no listener, that event is unhandled and
// crashes the whole process, so wait for either a successful spawn or that error before
// proceeding.
await new Promise<void>((resolve, reject) => {
child.once("spawn", () => resolve());
child.once("error", (err) => {
reject(
new Error(
`Could not start the LSP server "${spec.command}" for ${spec.languageId}. Is it installed and on your PATH? (${(err as Error).message})`,
),
);
});
});
// After startup, a late 'error' (e.g. the process dying unexpectedly) must not go unhandled
// either — the 'exit' handler below already disposes the connection, so just swallow it here.
child.on("error", () => {});
if (!child.stdin || !child.stdout) {
child.kill();
throw new Error(`LSP server "${spec.command}" did not open stdio streams.`);
}
const reader = new StreamMessageReader(child.stdout);
const writer = new StreamMessageWriter(child.stdin);
const connection = createProtocolConnection(reader, writer);
// vscode-jsonrpc buffers all messages until listen() starts pumping them — without this,
// sendRequest hangs (nothing is ever written) or throws "Call listen() first."
connection.listen();
// Surface stderr so a crashing server isn't a silent void (matches locode's MCP stdio policy).
child.stderr?.on("data", () => {
// Discard by default; a future debug mode could surface this. Don't let it back up.
});
await connection.sendRequest("initialize", {
processId: process.pid,
rootUri: URI.file(cwd).toString(),
capabilities: {
// locode consumes definition/references/diagnostics; declare only those so a server doesn't
// waste effort enabling features we'll never query. Full text sync (change=1) is simplest and
// correct — we always resend the whole file, never a range edit.
textDocumentSync: { openClose: true, change: 1 },
definitionProvider: true,
referencesProvider: true,
},
workspaceFolders: [{ uri: URI.file(cwd).toString(), name: path.basename(cwd) || cwd }],
});
// Per LSP spec, the client must send `initialized` after the initialize response.
await connection.sendNotification("initialized", {});
// A server crash should reject any in-flight request rather than hanging forever — listen for
// exit and dispose the connection so the next call throws instead of awaiting a dead process.
child.on("exit", () => {
connection.dispose();
handles.delete(`${cwd}::${spec.languageId}`);
});
return { connection, child, languageId: spec.languageId, openDocs: new Set() };
}
/** Returns the (lazily-started) LSP handle for the language owning `filePath`, or throws if no
* server is configured/can't start. The first call for a language pays the initialize round-trip;
* every later call reuses the live server. */
async function handleForFile(filePath: string, cwd: string): Promise<LspHandle> {
const spec = specForFile(filePath);
if (!spec) {
throw new Error(`No LSP server configured for "${path.extname(filePath)}" (code intelligence supports: ${LANGUAGE_SPECS.map((s) => s.extensions[0]).join(", ")}).`);
}
const key = `${cwd}::${spec.languageId}`;
let handle = handles.get(key);
if (!handle) {
handle = await startServer(spec, cwd);
handles.set(key, handle);
}
return handle;
}
/** Ensures the LSP server knows the current on-disk contents of `filePath`. Sends didOpen the
* first time a file is touched, didChange on subsequent syncs (the file was edited on disk since).
* Reads the file fresh each time — locode's tools write to disk before this runs, so the disk is
* the source of truth, not any in-memory buffer. */
async function syncDocument(handle: LspHandle, absPath: string, cwd: string): Promise<void> {
const uri = toUri(absPath);
const content = await fsReadFile(absPath, "utf-8");
if (!handle.openDocs.has(uri)) {
await handle.connection.sendNotification(DidOpenTextDocumentNotification.type, {
textDocument: { uri, languageId: handle.languageId, version: 1, text: content },
});
handle.openDocs.add(uri);
} else {
await handle.connection.sendNotification(DidChangeTextDocumentNotification.type, {
textDocument: { uri, version: Date.now() },
contentChanges: [{ text: content }],
});
}
}
export interface DefinitionResult {
/** The file/line/column of the symbol's definition. Multiple entries if the symbol has more than
* one definition (interface implementations, overloads, partial classes). Empty if the server
* found none (undefined symbol, or the server couldn't resolve it). */
definitions: LocResult[];
}
/** Resolves where the symbol at `line`/`column` (1-indexed) in `filePath` is defined. Syncs the
* document first so the server's view matches disk. Returns an empty list (not an error) when the
* server has no definition to offer — that's a legitimate "not found", not a failure. */
export async function getDefinition(
filePath: string,
line: number,
column: number,
cwd: string,
): Promise<DefinitionResult> {
const absPath = path.resolve(cwd, filePath);
const handle = await handleForFile(filePath, cwd);
await syncDocument(handle, absPath, cwd);
const pos: Position = { line: line - 1, character: column - 1 };
const result = (await handle.connection.sendRequest(DefinitionRequest.type, {
textDocument: { uri: toUri(absPath) } as TextDocumentIdentifier,
position: pos,
})) as Location | Location[] | null;
const locs = Array.isArray(result) ? result : result ? [result] : [];
return { definitions: locs.map(toLocation) };
}
export interface ReferencesResult {
/** Every place the symbol at `line`/`column` is referenced (including its definition). */
references: LocResult[];
}
/** Finds every reference to the symbol at `line`/`column` in `filePath`. `includeDeclaration`
* defaults to true (matches most IDE "find all references" behavior). */
export async function getReferences(
filePath: string,
line: number,
column: number,
cwd: string,
includeDeclaration = true,
): Promise<ReferencesResult> {
const absPath = path.resolve(cwd, filePath);
const handle = await handleForFile(filePath, cwd);
await syncDocument(handle, absPath, cwd);
const pos: Position = { line: line - 1, character: column - 1 };
const result = (await handle.connection.sendRequest(ReferencesRequest.type, {
textDocument: { uri: toUri(absPath) } as TextDocumentIdentifier,
position: pos,
context: { includeDeclaration },
})) as Location[] | null;
return { references: (result ?? []).map(toLocation) };
}
/** Notifies the LSP server that `filePath` changed on disk, so a subsequent `diagnostics` call
* reflects the new content. Called from the FileChanged hook path after edit_file/write_file. If no
* server is running for this language (or the file isn't one we manage), this is a no-op — it must
* never throw from a hook context, since hooks fire on every mutating tool. */
export async function notifyFileChanged(filePath: string, cwd: string): Promise<void> {
try {
const spec = specForFile(filePath);
if (!spec) return;
const key = `${cwd}::${spec.languageId}`;
const handle = handles.get(key);
if (!handle) return; // No server started yet — diagnostics will sync on first query.
await syncDocument(handle, path.resolve(cwd, filePath), cwd);
} catch {
// Best-effort: a hook context can't propagate errors into the turn.
}
}
type Severity = "error" | "warning" | "information" | "hint";
export interface DiagnosticsResult {
diagnostics: { path: string; line: number; column: number; severity: Severity; message: string; source?: string }[];
}
/** The most recent diagnostics the server has published for `filePath`. LSP pushes diagnostics via
* `textDocument/publishDiagnostics` notifications; locode collects them per-URI as they arrive and
* returns the latest snapshot here. Forces a document sync first so the snapshot is current. */
const diagnosticsByUri = new Map<string, Diagnostic[]>();
// Per-URI resolvers waiting on the next publishDiagnostics notification. getDiagnostics arms one
// for the file it just synced, then races it against a timeout — so a slow server (tsserver on a
// large file) still gets a chance to publish the fresh snapshot rather than the caller reading a
// stale one after a single event-loop turn. Resolved and cleared by the publishDiagnostics handler.
const diagWaiters = new Map<string, () => void>();
/** Wait for the next publishDiagnostics for `uri`, or give up after `timeoutMs`. Resolves true if
* a publish arrived, false on timeout. The waiter is removed either way. */
function waitForDiagnostics(uri: string, timeoutMs: number): Promise<boolean> {
return new Promise((resolve) => {
const timer = setTimeout(() => {
diagWaiters.delete(uri);
resolve(false);
}, timeoutMs);
diagWaiters.set(uri, () => {
clearTimeout(timer);
diagWaiters.delete(uri);
resolve(true);
});
});
}
const SEVERITY_MAP: Record<number, Severity> = {
1: "error",
2: "warning",
3: "information",
4: "hint",
};
export async function getDiagnostics(filePath: string, cwd: string): Promise<DiagnosticsResult> {
const absPath = path.resolve(cwd, filePath);
const handle = await handleForFile(filePath, cwd);
const uri = toUri(absPath);
// Attach a per-connection diagnostic collector the first time we use this handle.
if (!(handle as unknown as { __diagWired?: boolean }).__diagWired) {
(handle as unknown as { __diagWired?: boolean }).__diagWired = true;
handle.connection.onNotification("textDocument/publishDiagnostics", (params: { uri: string; diagnostics: Diagnostic[] }) => {
diagnosticsByUri.set(params.uri, params.diagnostics);
// Wake a getDiagnostics call waiting on this URI, if any.
diagWaiters.get(params.uri)?.();
});
}
// Clear any stale snapshot for this URI before syncing so a timeout fallthrough can't return
// diagnostics from before the edit. The server publishes asynchronously after didChange; race
// its next publish against a short timeout so a slow server (tsserver on a large file) still
// gets a chance to compute fresh diagnostics rather than us reading a stale snapshot after one
// event-loop turn. Fall through to whatever's cached on timeout (possibly empty).
diagnosticsByUri.delete(uri);
await syncDocument(handle, absPath, cwd);
await waitForDiagnostics(uri, 1500);
const diags = diagnosticsByUri.get(uri) ?? [];
return {
diagnostics: diags.map((d) => ({
path: URI.parse(uri).fsPath,
line: (d.range?.start.line ?? 0) + 1,
column: (d.range?.start.character ?? 0) + 1,
severity: SEVERITY_MAP[d.severity ?? 1] ?? "information",
message: messageToString(d.message ?? ""),
source: d.source,
})),
};
}
/** Shuts down every live LSP server. Call on locode exit so spawned servers (tsserver, pyright,
* gopls, …) don't outlive the process as orphans. Awaits each shutdown so the signals land before
* teardown. Best-effort: a stuck server can't block exit forever (the child kill still fires). */
export async function shutdownAll(): Promise<void> {
const all = [...handles.values()];
handles.clear();
await Promise.allSettled(
all.map(async (h) => {
try {
await h.connection.sendRequest("shutdown", null);
h.connection.sendNotification("exit", {});
} catch {
// Already dead — fall through to kill.
}
h.child.kill();
}),
);
}
/** For tests only: clear the live-handle registry and diagnostic cache without spawning/killing. */
export function _resetForTests(): void {
handles.clear();
diagnosticsByUri.clear();
diagWaiters.clear();
}
/** For tests only: a snapshot of the currently configured language specs (after any
* configureLanguageSpecs merge), so tests can assert the merge without spawning a server. */
export function _specsForTests(): readonly LanguageSpec[] {
return LANGUAGE_SPECS;
}
// The immutable built-in spec list, kept so tests can restore LANGUAGE_SPECS to defaults after a
// configureLanguageSpecs call (the merge is forward-only by design — production applies it once).
const BUILTIN_LANGUAGE_SPECS: readonly LanguageSpec[] = [
{ languageId: "typescript", extensions: [".ts", ".tsx", ".mts", ".cts", ".js", ".jsx", ".mjs", ".cjs"], command: "typescript-language-server", args: ["--stdio"] },
{ languageId: "python", extensions: [".py", ".pyi"], command: "pyright-langserver", args: ["--stdio"] },
{ languageId: "go", extensions: [".go"], command: "gopls", args: ["serve"] },
{ languageId: "rust", extensions: [".rs"], command: "rust-analyzer" },
{ languageId: "c", extensions: [".c", ".h", ".cpp", ".cc", ".cxx", ".hpp", ".hh", ".hxx"], command: "clangd" },
];
/** For tests only: restore the built-in language specs (undo any configureLanguageSpecs merge). */
export function _resetSpecsForTests(): void {
LANGUAGE_SPECS = BUILTIN_LANGUAGE_SPECS.map((s) => ({ ...s }));
}
+52 -60
View File
@@ -3,8 +3,8 @@ import { existsSync, mkdtempSync, rmSync } from "node:fs";
import path from "node:path";
import os from "node:os";
import { _setConfigFilePathForTest, loadStoredConfig, saveStoredConfig } from "./store.js";
import { resolveAutoCompactThreshold, resolveContextWindowDefault, resolveMaxIterations, resolveMaxOutputTokens, resolveMaxRetries, resolveRequestTimeoutMs, resolveSubagentTimeoutMs } from "./config.js";
import { DEFAULT_AUTO_COMPACT_THRESHOLD, DEFAULT_CONTEXT_WINDOW_LOCAL, DEFAULT_CONTEXT_WINDOW_CLOUD, DEFAULT_MAX_ITERATIONS, DEFAULT_MAX_OUTPUT_TOKENS, DEFAULT_MAX_RETRIES, DEFAULT_REQUEST_TIMEOUT_MS, DEFAULT_SUBAGENT_TIMEOUT_MS } from "./defaults.js";
import { resolveAutoCompactThreshold, resolveContextWindowDefault, resolveMaxIterations, resolveRequestTimeoutMs, resolveSubagentMaxDepth, resolveSubagentMaxIterations, resolveSubagentTimeoutMs } from "./config.js";
import { DEFAULT_AUTO_COMPACT_THRESHOLD, DEFAULT_CONTEXT_WINDOW, DEFAULT_MAX_ITERATIONS, DEFAULT_REQUEST_TIMEOUT_MS, DEFAULT_SUBAGENT_MAX_DEPTH, DEFAULT_SUBAGENT_MAX_ITERATIONS, DEFAULT_SUBAGENT_TIMEOUT_MS } from "./defaults.js";
// Isolate the persisted config to a temp directory so the suite never reads or overwrites the
// user's real ~/.config/locode/config.json (the previous afterEach { saveStoredConfig({}) } wiped
@@ -25,10 +25,9 @@ describe("config resolution", () => {
delete process.env.LOCODE_AUTO_COMPACT_THRESHOLD;
delete process.env.LOCODE_CONTEXT_WINDOW;
delete process.env.LOCODE_MAX_ITERATIONS;
delete process.env.LOCODE_MAX_OUTPUT_TOKENS;
delete process.env.LOCODE_MAX_OUTPUT_TOKENS;
delete process.env.LOCODE_MAX_RETRIES;
delete process.env.LOCODE_REQUEST_TIMEOUT_MS;
delete process.env.LOCODE_SUBAGENT_MAX_ITERATIONS;
delete process.env.LOCODE_SUBAGENT_MAX_DEPTH;
delete process.env.LOCODE_SUBAGENT_TIMEOUT_MS;
});
@@ -53,46 +52,14 @@ describe("config resolution", () => {
expect(resolveAutoCompactThreshold()).toBe(DEFAULT_AUTO_COMPACT_THRESHOLD);
});
it("resolves context window default — small local model", () => {
expect(resolveContextWindowDefault("http://localhost:11434/v1", "llama3.2:latest")).toBe(DEFAULT_CONTEXT_WINDOW_LOCAL);
});
it("resolves context window default — remote cloud API", () => {
expect(resolveContextWindowDefault("https://api.openai.com/v1", "gpt-4")).toBe(DEFAULT_CONTEXT_WINDOW_CLOUD);
});
it("resolves context window default — Ollama's cloud-routed models on a local endpoint", () => {
// Regression: glm-5.2:cloud etc. run through the same localhost Ollama daemon as a genuinely
// local model, so the base URL alone can't distinguish them — see isSmallLocalModel().
expect(resolveContextWindowDefault("http://localhost:11434/v1", "glm-5.2:cloud")).toBe(DEFAULT_CONTEXT_WINDOW_CLOUD);
expect(resolveContextWindowDefault("http://localhost:11434/v1", "qwen3.5:397b-cloud")).toBe(DEFAULT_CONTEXT_WINDOW_CLOUD);
it("resolves context window default", () => {
expect(resolveContextWindowDefault()).toBe(DEFAULT_CONTEXT_WINDOW);
});
it("resolves max iterations default", () => {
expect(resolveMaxIterations()).toBe(DEFAULT_MAX_ITERATIONS);
});
it("resolves max output tokens default", () => {
expect(resolveMaxOutputTokens()).toBe(DEFAULT_MAX_OUTPUT_TOKENS);
});
it("reads max output tokens from env", () => {
process.env.LOCODE_MAX_OUTPUT_TOKENS = "16384";
expect(resolveMaxOutputTokens()).toBe(16_384);
});
it("reads max output tokens from stored config", () => {
saveStoredConfig({ maxOutputTokens: 4096 });
expect(resolveMaxOutputTokens()).toBe(4096);
});
it("rejects out-of-range max output tokens", () => {
saveStoredConfig({ maxOutputTokens: 100 });
expect(resolveMaxOutputTokens()).toBe(DEFAULT_MAX_OUTPUT_TOKENS);
process.env.LOCODE_MAX_OUTPUT_TOKENS = "5000000";
expect(resolveMaxOutputTokens()).toBe(DEFAULT_MAX_OUTPUT_TOKENS);
});
it("resolves request timeout default", () => {
expect(resolveRequestTimeoutMs()).toBe(DEFAULT_REQUEST_TIMEOUT_MS);
});
@@ -114,6 +81,52 @@ describe("config resolution", () => {
expect(resolveRequestTimeoutMs()).toBe(DEFAULT_REQUEST_TIMEOUT_MS);
});
it("resolves sub-agent max iterations default", () => {
expect(resolveSubagentMaxIterations()).toBe(DEFAULT_SUBAGENT_MAX_ITERATIONS);
});
it("reads sub-agent max iterations from env", () => {
process.env.LOCODE_SUBAGENT_MAX_ITERATIONS = "30";
expect(resolveSubagentMaxIterations()).toBe(30);
});
it("reads sub-agent max iterations from stored config", () => {
saveStoredConfig({ subagentMaxIterations: 25 });
expect(resolveSubagentMaxIterations()).toBe(25);
});
it("rejects out-of-range sub-agent max iterations", () => {
process.env.LOCODE_SUBAGENT_MAX_ITERATIONS = "0"; // below floor
expect(resolveSubagentMaxIterations()).toBe(DEFAULT_SUBAGENT_MAX_ITERATIONS);
process.env.LOCODE_SUBAGENT_MAX_ITERATIONS = "5000"; // above cap
expect(resolveSubagentMaxIterations()).toBe(DEFAULT_SUBAGENT_MAX_ITERATIONS);
saveStoredConfig({ subagentMaxIterations: 0 }); // stored below floor
expect(resolveSubagentMaxIterations()).toBe(DEFAULT_SUBAGENT_MAX_ITERATIONS);
});
it("resolves sub-agent max depth default", () => {
expect(resolveSubagentMaxDepth()).toBe(DEFAULT_SUBAGENT_MAX_DEPTH);
});
it("reads sub-agent max depth from env", () => {
process.env.LOCODE_SUBAGENT_MAX_DEPTH = "4";
expect(resolveSubagentMaxDepth()).toBe(4);
});
it("reads sub-agent max depth from stored config", () => {
saveStoredConfig({ subagentMaxDepth: 3 });
expect(resolveSubagentMaxDepth()).toBe(3);
});
it("rejects out-of-range sub-agent max depth", () => {
process.env.LOCODE_SUBAGENT_MAX_DEPTH = "-1"; // below floor
expect(resolveSubagentMaxDepth()).toBe(DEFAULT_SUBAGENT_MAX_DEPTH);
process.env.LOCODE_SUBAGENT_MAX_DEPTH = "99"; // above cap
expect(resolveSubagentMaxDepth()).toBe(DEFAULT_SUBAGENT_MAX_DEPTH);
saveStoredConfig({ subagentMaxDepth: -1 }); // stored below floor
expect(resolveSubagentMaxDepth()).toBe(DEFAULT_SUBAGENT_MAX_DEPTH);
});
it("resolves sub-agent timeout default", () => {
expect(resolveSubagentTimeoutMs()).toBe(DEFAULT_SUBAGENT_TIMEOUT_MS);
});
@@ -136,25 +149,4 @@ describe("config resolution", () => {
saveStoredConfig({ subagentTimeoutMs: 500 }); // stored below floor
expect(resolveSubagentTimeoutMs()).toBe(DEFAULT_SUBAGENT_TIMEOUT_MS);
});
it("resolves max retries default", () => {
expect(resolveMaxRetries()).toBe(DEFAULT_MAX_RETRIES);
});
it("reads max retries from env", () => {
process.env.LOCODE_MAX_RETRIES = "3";
expect(resolveMaxRetries()).toBe(3);
});
it("reads max retries from stored config", () => {
saveStoredConfig({ maxRetries: 5 });
expect(resolveMaxRetries()).toBe(5);
});
it("rejects out-of-range max retries", () => {
saveStoredConfig({ maxRetries: 11 });
expect(resolveMaxRetries()).toBe(DEFAULT_MAX_RETRIES);
process.env.LOCODE_MAX_RETRIES = "-1";
expect(resolveMaxRetries()).toBe(DEFAULT_MAX_RETRIES);
});
});
+32 -46
View File
@@ -1,15 +1,13 @@
import {
DEFAULT_AUTO_COMPACT_THRESHOLD,
DEFAULT_CONTEXT_WINDOW_LOCAL,
DEFAULT_CONTEXT_WINDOW_CLOUD,
DEFAULT_CONTEXT_WINDOW,
DEFAULT_MAX_ITERATIONS,
DEFAULT_MAX_OUTPUT_TOKENS,
DEFAULT_MAX_RETRIES,
DEFAULT_REQUEST_TIMEOUT_MS,
DEFAULT_SUBAGENT_MAX_DEPTH,
DEFAULT_SUBAGENT_MAX_ITERATIONS,
DEFAULT_SUBAGENT_TIMEOUT_MS,
KNOWN_BACKENDS,
type BackendName,
isSmallLocalModel,
} from "./defaults.js";
import { loadStoredConfig } from "./store.js";
@@ -51,28 +49,13 @@ export function resolveModel(cliModel?: string): string | undefined {
return cliModel ?? process.env.LOCODE_MODEL ?? stored.model;
}
/** The fallback context window size to use when it can't be auto-detected from the backend. Picks a
* small-local or cloud-scale default based on isSmallLocalModel() — the backend host alone isn't
* enough, since Ollama's cloud-routed models (e.g. "glm-5.2:cloud") share a local host with
* genuinely local ones. */
export function resolveContextWindowDefault(baseURL: string, model: string): number {
/** The fallback context window size to use when it can't be auto-detected from the backend. */
export function resolveContextWindowDefault(): number {
const stored = loadStoredConfig();
const envValue = Number(process.env.LOCODE_CONTEXT_WINDOW);
if (Number.isFinite(envValue) && envValue > 0) return envValue;
if (typeof stored.contextWindow === "number" && stored.contextWindow > 0) return stored.contextWindow;
return isSmallLocalModel(baseURL, model) ? DEFAULT_CONTEXT_WINDOW_LOCAL : DEFAULT_CONTEXT_WINDOW_CLOUD;
}
/** Ceiling on a single response's max_tokens (see DEFAULT_MAX_OUTPUT_TOKENS), independent of the
* context window. Bounded to 256–1,000,000 to reject pathological values. */
export function resolveMaxOutputTokens(): number {
const stored = loadStoredConfig();
const envValue = Number(process.env.LOCODE_MAX_OUTPUT_TOKENS);
if (Number.isFinite(envValue) && envValue >= 256 && envValue <= 1_000_000) return envValue;
if (typeof stored.maxOutputTokens === "number" && stored.maxOutputTokens >= 256 && stored.maxOutputTokens <= 1_000_000) {
return stored.maxOutputTokens;
}
return DEFAULT_MAX_OUTPUT_TOKENS;
return DEFAULT_CONTEXT_WINDOW;
}
/** Max tool calls allowed per turn before locode gives up. */
@@ -84,6 +67,31 @@ export function resolveMaxIterations(): number {
return DEFAULT_MAX_ITERATIONS;
}
/** Max tool calls in a single sub-agent turn (see DEFAULT_SUBAGENT_MAX_ITERATIONS). Bounded to
* 1–1000 to reject pathological values. */
export function resolveSubagentMaxIterations(): number {
const stored = loadStoredConfig();
const envValue = Number(process.env.LOCODE_SUBAGENT_MAX_ITERATIONS);
if (Number.isFinite(envValue) && envValue >= 1 && envValue <= 1000) return envValue;
if (typeof stored.subagentMaxIterations === "number" && stored.subagentMaxIterations >= 1 && stored.subagentMaxIterations <= 1000) {
return stored.subagentMaxIterations;
}
return DEFAULT_SUBAGENT_MAX_ITERATIONS;
}
/** Max nesting depth for sub-agents (see DEFAULT_SUBAGENT_MAX_DEPTH). Bounded to 0–10: 0 lets the
* main session still delegate (depth 1) but blocks that delegate from delegating further; values
* above ~10 serve no real purpose and just invite runaway recursion. */
export function resolveSubagentMaxDepth(): number {
const stored = loadStoredConfig();
const envValue = Number(process.env.LOCODE_SUBAGENT_MAX_DEPTH);
if (Number.isFinite(envValue) && envValue >= 0 && envValue <= 10) return envValue;
if (typeof stored.subagentMaxDepth === "number" && stored.subagentMaxDepth >= 0 && stored.subagentMaxDepth <= 10) {
return stored.subagentMaxDepth;
}
return DEFAULT_SUBAGENT_MAX_DEPTH;
}
/** Wall-clock budget for a single sub-agent turn (see DEFAULT_SUBAGENT_TIMEOUT_MS). Bounded to
* 1s–1h to reject pathological values. */
export function resolveSubagentTimeoutMs(): number {
@@ -107,22 +115,8 @@ export function resolveAutoCompactThreshold(): number {
return DEFAULT_AUTO_COMPACT_THRESHOLD;
}
/** Max retry attempts the OpenAI SDK makes on transient failures (connection errors, 429, 5xx)
* with exponential backoff. Bounded to 0–10 to reject pathological values. 0 = fail immediately,
* matching locode's old behavior of never retrying (a slow local backend usually means the model
* is genuinely stuck, not a transient blip — but some setups have occasional connection drops). */
export function resolveMaxRetries(): number {
const stored = loadStoredConfig();
const envValue = Number(process.env.LOCODE_MAX_RETRIES);
if (Number.isFinite(envValue) && envValue >= 0 && envValue <= 10) return envValue;
if (typeof stored.maxRetries === "number" && stored.maxRetries >= 0 && stored.maxRetries <= 10) {
return stored.maxRetries;
}
return DEFAULT_MAX_RETRIES;
}
/** Milliseconds to wait on a single chat completion request before giving up (see backend/client.ts
* for why locode defaults to no retries). Bounded to 10s–30min to reject pathological values. */
* for why locode doesn't retry on top of this). Bounded to 10s–30min to reject pathological values. */
export function resolveRequestTimeoutMs(): number {
const stored = loadStoredConfig();
const envValue = Number(process.env.LOCODE_REQUEST_TIMEOUT_MS);
@@ -132,11 +126,3 @@ export function resolveRequestTimeoutMs(): number {
}
return DEFAULT_REQUEST_TIMEOUT_MS;
}
/** User-configured LSP server overrides/additions (see StoredConfig.lspServers). An empty object
* means "use the built-in language→server mappings only". Validated loosely: entries without a
* command are dropped by configureLanguageSpecs, so we just pass them through. */
export function resolveLspServers(): Record<string, { command: string; args?: string[]; extensions?: string[] }> {
const stored = loadStoredConfig();
return stored.lspServers ?? {};
}
-61
View File
@@ -1,61 +0,0 @@
import { describe, expect, it } from "vitest";
import { DEFAULT_CONTEXT_WINDOW_CLOUD, DEFAULT_CONTEXT_WINDOW_LOCAL, isCloudRoutedModelName, isLocalBackendURL, isSmallLocalModel } from "./defaults.js";
describe("isLocalBackendURL", () => {
it("recognizes localhost, 127.0.0.1, and ::1", () => {
expect(isLocalBackendURL("http://localhost:11434/v1")).toBe(true);
expect(isLocalBackendURL("http://127.0.0.1:1234/v1")).toBe(true);
expect(isLocalBackendURL("http://[::1]:11434/v1")).toBe(true);
});
it("rejects a remote host", () => {
expect(isLocalBackendURL("https://api.openai.com/v1")).toBe(false);
});
it("returns false for an unparseable URL instead of throwing", () => {
expect(isLocalBackendURL("not a url")).toBe(false);
});
});
describe("isCloudRoutedModelName", () => {
it("recognizes Ollama's cloud tag conventions", () => {
expect(isCloudRoutedModelName("glm-5.2:cloud")).toBe(true);
expect(isCloudRoutedModelName("qwen3.5:397b-cloud")).toBe(true);
expect(isCloudRoutedModelName("gpt-oss:120b-cloud")).toBe(true);
});
it("does not match a genuinely local model", () => {
expect(isCloudRoutedModelName("llama3.2:latest")).toBe(false);
expect(isCloudRoutedModelName("gemma3:4b")).toBe(false);
expect(isCloudRoutedModelName("phi4:14b")).toBe(false);
});
it("does not false-positive on 'cloud' appearing outside the tag segment", () => {
expect(isCloudRoutedModelName("cloudmodel:latest")).toBe(false);
expect(isCloudRoutedModelName("cloud")).toBe(false); // no colon at all
});
});
describe("isSmallLocalModel", () => {
it("is true for a genuinely local model on a local endpoint", () => {
expect(isSmallLocalModel("http://localhost:11434/v1", "llama3.2:latest")).toBe(true);
});
it("is false for a cloud-routed model even on a local endpoint", () => {
expect(isSmallLocalModel("http://localhost:11434/v1", "glm-5.2:cloud")).toBe(false);
});
it("is false for a remote backend regardless of model name", () => {
expect(isSmallLocalModel("https://api.openai.com/v1", "gpt-4")).toBe(false);
});
});
describe("DEFAULT_CONTEXT_WINDOW_LOCAL / CLOUD", () => {
it("LOCAL is the small-local fallback (8 192)", () => {
expect(DEFAULT_CONTEXT_WINDOW_LOCAL).toBe(8192);
});
it("CLOUD is the cloud/large fallback (131 072)", () => {
expect(DEFAULT_CONTEXT_WINDOW_CLOUD).toBe(131072);
});
});
+29 -67
View File
@@ -8,82 +8,44 @@ export const KNOWN_BACKENDS = {
export type BackendName = keyof typeof KNOWN_BACKENDS;
/** Returns true if the given base URL looks like a local backend (Ollama or LM Studio on localhost).
* On its own this is NOT enough to tell a small local model from a large one — see
* isSmallLocalModel() below, which is what callers should actually use. */
export function isLocalBackendURL(baseURL: string): boolean {
try {
const url = new URL(baseURL);
// WHATWG URL keeps the brackets on a literal IPv6 host in .hostname (e.g. "[::1]", not "::1") —
// https://url.spec.whatwg.org/#concept-host-serializer.
return url.hostname === "localhost" || url.hostname === "127.0.0.1" || url.hostname === "[::1]";
} catch {
return false;
}
}
/** Used when the context window can't be auto-detected from the backend (see backend/contextWindow.ts)
* and the user hasn't configured one — a conservative size common among smaller local models. */
export const DEFAULT_CONTEXT_WINDOW = 8192;
/** Ollama's cloud-hosted models are proxied through the same local daemon as a genuinely local
* model — same base URL, same host — so isLocalBackendURL() alone can't tell them apart. Ollama
* names them with a "cloud" tag segment instead, e.g. "glm-5.2:cloud" or "qwen3.5:397b-cloud". */
export function isCloudRoutedModelName(model: string): boolean {
const tag = model.split(":")[1] ?? "";
return tag === "cloud" || tag.endsWith("-cloud");
}
/** Max tool calls per *round* within a turn. Local models often issue one tool call per round, so
* a multi-file edit + verify sequence can easily run past 50; 100 gives real tasks room to breathe
* while still bounding a genuinely stuck model per round. runTurn automatically chains up to
* MAX_CONTINUATION_ROUNDS rounds before pausing, so the effective per-turn budget is much larger. */
export const DEFAULT_MAX_ITERATIONS = 100;
/** Whether locode should treat this model as a small/less-capable local model — extra system-prompt
* guidance about unreliable tool calls and limited output, plus a conservative context-window
* fallback — rather than a large, reliable one. True only when the backend is local AND the model
* isn't one of Ollama's cloud-routed models under that same local endpoint; a backend hosted
* elsewhere (a real remote/cloud API) is never treated as "small local" regardless of model name. */
export function isSmallLocalModel(baseURL: string, model: string): boolean {
return isLocalBackendURL(baseURL) && !isCloudRoutedModelName(model);
}
/** Number of internal rounds runTurn will chain automatically when a single round exhausts its
* step budget without producing a final answer. This matches Claude Code's behavior of continuing
* a long tool chain rather than stopping after every N steps and asking the user to continue.
* Effective per-turn budget = DEFAULT_MAX_ITERATIONS * MAX_CONTINUATION_ROUNDS. */
export const MAX_CONTINUATION_ROUNDS = 5;
/** Fallback context window when auto-detection fails and the user hasn't configured one.
* Small local models (see isSmallLocalModel()) typically have 8k–32k context, so 8192 is a safe
* conservative default. Everything else (cloud APIs, and Ollama's own cloud-routed models) typically
* has 128k–1M+ context, so 131072 (128K) avoids severely underutilizing them. */
export const DEFAULT_CONTEXT_WINDOW_LOCAL = 8192;
export const DEFAULT_CONTEXT_WINDOW_CLOUD = 131072;
/** Max tool calls in a single sub-agent turn. Sub-agents run headless (one round, no
* auto-continue — see runSubAgentTurn) and are meant for one focused task, so this is intentionally
* smaller than the parent's per-round budget (DEFAULT_MAX_ITERATIONS): a runaway sub-agent fails
* fast and tells the parent to split the task rather than silently burning a large black-box
* budget. Configurable via `LOCODE_SUBAGENT_MAX_ITERATIONS` / `locode config set subagentMaxIterations`. */
export const DEFAULT_SUBAGENT_MAX_ITERATIONS = 50;
/** Ceiling on a single response's `max_tokens`, independent of the model's context window. Most
* backends cap how much a single completion can generate well below the total context window they
* advertise (e.g. Ollama's glm-5.1:cloud / glm-5.2:cloud report a 1,000,000-token context window
* but error on a request above their real 131,072-token output cap: "max_tokens (500000) exceeds
* model's maximum output tokens (131072)") — resolveMaxTokens (agent/loop.ts) used to request up to
* the whole remaining window, which such backends rejected outright as a context/length error even
* on the very first turn. 131072 (128K) matches that verified real ceiling exactly, so it's safe to
* use as the default without erroring on the very first turn — going straight to a backend's real
* cap only became reasonable once detectRepetitionLoop (agent/loop.ts) existed to abort a model
* that gets stuck generating instead of relying on this value alone as the safety valve. Raise it
* further via `locode config set maxOutputTokens` for backends known to allow more; lower it for
* ones with a smaller real ceiling. */
export const DEFAULT_MAX_OUTPUT_TOKENS = 131_072;
/** Max model requests per turn before locode pauses rather than looping forever. Each iteration
* is one model generation request (one tool-call round-trip), and local models commonly issue a
* single tool call per request — so a real multi-file task (read several files, edit each, grep
* to verify, re-read) easily needs 40–60 requests. 50 was too tight and caused frequent
* "Paused after 50 steps" soft-stops on legitimate work; 100 still caused frequent pauses on
* larger tasks. 300 gives real tasks ample room to finish while still bounding a genuinely stuck
* model. Hitting the cap is a soft pause, not a failure (the work so far is intact — send another
* message to resume). Configurable via `maxIterations`, e.g. `locode config set maxIterations 500`
* for very large batch jobs. */
export const DEFAULT_MAX_ITERATIONS = 300;
/** Max nesting depth for sub-agents. Depth 0 is the main session; a sub-agent it spawns is depth
* 1, a sub-agent that one spawns is depth 2, and so on. A sub-agent at the cap can't spawn further
* sub-agents (the toolset excludes `agent` for it, with an explicit depth-check backstop in
* agent/loop.ts). The default of 2 allows one level of delegation plus a focused sub-task under
* that, while keeping recursion shallow enough that a misbehaving model can't fan out
* uncontrollably. Configurable via `LOCODE_SUBAGENT_MAX_DEPTH` / `locode config set subagentMaxDepth`. */
export const DEFAULT_SUBAGENT_MAX_DEPTH = 2;
/** Fraction of the context window at which locode automatically summarizes the conversation.
* User-configurable via `locode config set autoCompactThreshold`. */
export const DEFAULT_AUTO_COMPACT_THRESHOLD = 0.85;
/** How long to wait on a single chat completion request before giving up. The OpenAI SDK retries
* transient failures (connection errors, 429, 5xx) up to `maxRetries` times with exponential
* backoff before surfacing the error; set to 0 to fail immediately like older locode versions.
* Raise this via `maxRetries` if your backend has occasional transient blips. See backend/client.ts. */
export const DEFAULT_MAX_RETRIES = 0;
/** How long to wait on a single chat completion request before giving up. Raise this via
* `requestTimeoutMs` if your backend queues requests behind a concurrency limit (e.g. Ollama's
* `OLLAMA_NUM_PARALLEL`) rather than serving them immediately. */
/** How long to wait on a single chat completion request before giving up (no retries — see
* backend/client.ts). Raise this via `requestTimeoutMs` if your backend queues requests behind a
* concurrency limit (e.g. Ollama's `OLLAMA_NUM_PARALLEL`) rather than serving them immediately. */
export const DEFAULT_REQUEST_TIMEOUT_MS = 180_000;
/** Wall-clock budget for a single sub-agent turn. Sub-agents make their own sequence of model
+19 -31
View File
@@ -1,6 +1,7 @@
import envPaths from "env-paths";
import { existsSync, mkdirSync, readFileSync, writeFileSync } from "node:fs";
import path from "node:path";
import type { PermissionRule } from "../permissions/types.js";
export interface StoredConfig {
backend?: string;
@@ -10,24 +11,20 @@ export interface StoredConfig {
contextWindow?: number;
/** Max tool calls allowed per turn before locode gives up rather than looping forever. */
maxIterations?: number;
/** Max tool calls in a single sub-agent turn (smaller than maxIterations — sub-agents are bounded
* to one focused task with no auto-continue). */
subagentMaxIterations?: number;
/** Max nesting depth for sub-agents (0 = main session, 1 = first delegation, …). */
subagentMaxDepth?: number;
/** Fraction of the context window (0.0–1.0) at which locode auto-compacts the conversation. */
autoCompactThreshold?: number;
/** Milliseconds to wait on a single chat completion request before giving up. */
requestTimeoutMs?: number;
/** Milliseconds of wall-clock budget for a single sub-agent turn. */
subagentTimeoutMs?: number;
/** Ceiling on a single response's max_tokens, independent of contextWindow. */
maxOutputTokens?: number;
/** Max retry attempts the OpenAI SDK makes on transient failures (connection errors, 429, 5xx)
* before surfacing the error. 0 = fail immediately (old behavior); the SDK uses exponential
* backoff between attempts. */
maxRetries?: number;
/** User-defined LSP server overrides/additions, keyed by language id (e.g. "java", "ruby",
* "typescript"). Each entry is { command, args?, extensions? }. An entry for a built-in id
* overrides its command/args; an entry with `extensions` also rewrites which file extensions
* route to that language. Entries for new ids add support for languages locode doesn't ship a
* server for. See `locode config set lspServers` (JSON value). */
lspServers?: Record<string, { command: string; args?: string[]; extensions?: string[] }>;
/** User-level permission rules (auto-approve/deny tool calls without prompting). Merged with
* project-level rules from .locode/settings.json; deny wins across layers. See PermissionRule. */
permissionRules?: PermissionRule[];
}
const paths = envPaths("locode", { suffix: "" });
@@ -41,39 +38,30 @@ export function configFilePath(): string {
return configFileOverride ?? configFile;
}
/** Directory holding the user-level config files (config.json, hooks.json, mcp.json, plugins.json,
* and the freeform memory.md). Derived from configFilePath() so the test override
* (_setConfigFilePathForTest) redirects this too. */
export function configDirPath(): string {
return path.dirname(configFilePath());
}
/** @internal For tests only — redirect config persistence to `path` (pass undefined to reset). */
export function _setConfigFilePathForTest(p: string | undefined): void {
configFileOverride = p;
invalidateConfigCache();
}
// In-memory cache so every resolve*() call in the same process doesn't re-read and re-parse
// the same small JSON file (8–10 calls during startup alone). Invalidated by saveStoredConfig
// and by tests that change the config path.
let cachedConfig: StoredConfig | undefined;
export function loadStoredConfig(): StoredConfig {
if (cachedConfig !== undefined) return cachedConfig;
const file = configFilePath();
if (!existsSync(file)) { cachedConfig = {}; return cachedConfig; }
if (!existsSync(file)) return {};
try {
cachedConfig = JSON.parse(readFileSync(file, "utf-8")) as StoredConfig;
return cachedConfig;
return JSON.parse(readFileSync(file, "utf-8")) as StoredConfig;
} catch {
cachedConfig = {};
return cachedConfig;
return {};
}
}
/** Invalidate the in-memory cache — called by saveStoredConfig and by tests that change
* the config file path. */
export function invalidateConfigCache(): void {
cachedConfig = undefined;
}
export function saveStoredConfig(cfg: StoredConfig): void {
const file = configFilePath();
mkdirSync(path.dirname(file), { recursive: true });
writeFileSync(file, JSON.stringify(cfg, null, 2));
cachedConfig = cfg;
}
+3 -14
View File
@@ -91,22 +91,11 @@ describe("runHooksForEvent", () => {
expect(result.additionalContext).toBeUndefined();
});
it("injects a prompt hook's message as additional context", async () => {
it("warns for unsupported prompt hooks", async () => {
loadMergedHooks.loadMergedHooks.mockReturnValueOnce({
SessionStart: [{ hooks: [{ type: "prompt", message: "Remember to check the changelog." }] }],
SessionStart: [{ hooks: [{ type: "prompt", message: "ok?" } as any] }],
});
const result = await runHooksForEvent("SessionStart", ctx, {});
expect(result.additionalContext).toBe("Remember to check the changelog.");
expect(result.warnings).toEqual([]);
});
it("combines a prompt hook's message with a command hook's stdout", async () => {
loadMergedHooks.loadMergedHooks.mockReturnValueOnce({
SessionStart: [{ hooks: [{ type: "prompt", message: "prompt message" }, { type: "command", command: "echo cmd" }] }],
});
execa.execa.mockResolvedValueOnce({ exitCode: 0, stdout: "cmd output", stderr: "", timedOut: false });
const result = await runHooksForEvent("SessionStart", ctx, {});
expect(result.additionalContext).toContain("prompt message");
expect(result.additionalContext).toContain("cmd output");
expect(result.warnings).toContain("1 prompt hook(s) skipped (not yet implemented).");
});
});
+11 -17
View File
@@ -2,7 +2,7 @@ import { execa } from "execa";
import path from "node:path";
import { loadMergedHooks } from "./config.js";
import { matcherMatches } from "./matcher.js";
import type { Hook, HookCommand, HookEventName, HookHttp, HookPrompt } from "./types.js";
import type { Hook, HookCommand, HookEventName, HookHttp } from "./types.js";
const DEFAULT_TIMEOUT_SECONDS = 30;
@@ -37,20 +37,6 @@ function isHttpHook(hook: Hook): hook is HookHttp {
return hook.type === "http";
}
function isPromptHook(hook: Hook): hook is HookPrompt {
return hook.type === "prompt";
}
/** A prompt hook has no process/response to run — it's just a static message that always
* "succeeds" and folds into additionalContext the same way a command/http hook's plain-text
* stdout does (see the outcome-handling loop in runHooksForEvent). Wrapped in a resolved promise
* so it can share the same Promise.all as the other hook kinds. */
async function runPromptHook(
hook: HookPrompt,
): Promise<{ exitCode: number; stdout: string; stderr: string; timedOut: boolean; json?: unknown; warning?: string }> {
return { exitCode: 0, stdout: hook.message, stderr: "", timedOut: false };
}
function parseOutput(output: string, schema?: "json"): { text?: string; json?: unknown; warning?: string } {
if (!schema) return { text: output };
if (schema !== "json") return { text: output };
@@ -209,18 +195,26 @@ export async function runHooksForEvent(
.flatMap((entry) => entry.hooks);
if (hooks.length === 0) return EMPTY_RESULT;
// Prompt hooks are declared in the type system but not implemented in this pass — they need a
// blocking UI flow the current runner doesn't have. Skip them with a warning so a config that
// includes them doesn't silently do nothing.
const skippedPrompts = hooks.filter((h) => h.type === "prompt").length;
const runnableHooks = hooks.filter((hook) => hook.type !== "prompt");
const stdinPayload = JSON.stringify({ hook_event_name: event, session_id: ctx.sessionId, cwd: ctx.cwd, ...payload });
const outcomes = await Promise.all(
hooks.map(async (hook) => {
runnableHooks.map(async (hook) => {
if (isCommandHook(hook)) return runCommandHook(hook, stdinPayload, ctx);
if (isHttpHook(hook)) return runHttpHook(hook, stdinPayload, ctx);
if (isPromptHook(hook)) return runPromptHook(hook);
return { exitCode: 1, stdout: "", stderr: `Unsupported hook type: ${(hook as Hook).type}`, timedOut: false };
}),
);
const result: HookResult = { blocked: false, warnings: [] };
if (skippedPrompts > 0) {
result.warnings.push(`${skippedPrompts} prompt hook(s) skipped (not yet implemented).`);
}
for (const outcome of outcomes) {
if (outcome.warning) {
result.warnings.push(outcome.warning);
+1 -3
View File
@@ -40,10 +40,8 @@ export interface HookHttp extends HookBase {
export interface HookPrompt extends HookBase {
type: "prompt";
/** Injected verbatim as additional context, the same way a command/http hook's stdout is —
* see runner.ts. Always "succeeds" (there's no process/response to fail); a prompt hook can't
* block an event the way a command hook's exit code 2 can. */
message: string;
/** Not yet implemented — prompt hooks require a UI blocking flow the current runner doesn't support. */
}
export type Hook = HookCommand | HookHttp | HookPrompt;
+23 -47
View File
@@ -3,7 +3,6 @@ import { StdioClientTransport } from "@modelcontextprotocol/sdk/client/stdio.js"
import { StreamableHTTPClientTransport } from "@modelcontextprotocol/sdk/client/streamableHttp.js";
import type { Transport } from "@modelcontextprotocol/sdk/shared/transport.js";
import { isHttpServerConfig, type McpServerConfig } from "./types.js";
import pkg from "../../package.json" with { type: "json" };
export interface McpTextContentBlock {
type: "text";
@@ -55,53 +54,30 @@ export interface ConnectedMcpServer {
tools: McpToolInfo[];
}
/** Max connection attempts for an MCP server that fails transiently (process slow to start,
* HTTP 503, etc.). A stdio server whose command genuinely doesn't exist fails immediately every
* time, so retries only help the transient case — kept small to avoid stalling startup. */
const MCP_CONNECT_MAX_ATTEMPTS = 3;
const MCP_CONNECT_BASE_DELAY_MS = 500;
async function sleep(ms: number, signal?: AbortSignal): Promise<void> {
return new Promise((resolve, reject) => {
const t = setTimeout(resolve, ms);
signal?.addEventListener("abort", () => { clearTimeout(t); reject(new Error("aborted")); }, { once: true });
});
}
export async function connectMcpServer(name: string, config: McpServerConfig): Promise<ConnectedMcpServer> {
let lastErr: unknown;
for (let attempt = 1; attempt <= MCP_CONNECT_MAX_ATTEMPTS; attempt++) {
const transport: Transport = isHttpServerConfig(config)
? new StreamableHTTPClientTransport(new URL(config.url), {
requestInit: config.headers ? { headers: config.headers } : undefined,
})
: new StdioClientTransport({
command: config.command,
args: config.args,
env: config.env,
// Default is "inherit", which would leak the child's stderr straight into the terminal
// and corrupt Ink's alternate-screen UI. Pipe it instead so it's just discarded.
stderr: "pipe",
});
const transport: Transport = isHttpServerConfig(config)
? new StreamableHTTPClientTransport(new URL(config.url), {
requestInit: config.headers ? { headers: config.headers } : undefined,
})
: new StdioClientTransport({
command: config.command,
args: config.args,
env: config.env,
// Default is "inherit", which would leak the child's stderr straight into the terminal
// and corrupt Ink's alternate-screen UI. Pipe it instead so it's just discarded.
stderr: "pipe",
});
const client = new Client({ name: "locode", version: pkg.version });
try {
await client.connect(transport);
const { tools } = await client.listTools();
return { name, client, transport, tools: tools as McpToolInfo[] };
} catch (err) {
// listTools failed or connect failed — close the transport so the stdio subprocess / HTTP
// connection isn't orphaned. manager.ts catches the rejection as an error status, but
// without this the child process keeps running.
await transport.close().catch(() => {});
lastErr = err;
// Retry with exponential backoff for transient failures; the last attempt's error is what
// the caller sees. A genuinely broken config (missing binary, bad URL) fails fast every time,
// so the retries just add a small delay — acceptable for the rare transient-startup case.
if (attempt < MCP_CONNECT_MAX_ATTEMPTS) {
await sleep(MCP_CONNECT_BASE_DELAY_MS * Math.pow(2, attempt - 1));
}
}
const client = new Client({ name: "locode", version: "0.3.1" });
await client.connect(transport);
try {
const { tools } = await client.listTools();
return { name, client, transport, tools: tools as McpToolInfo[] };
} catch (err) {
// listTools failed (server connected but never responded to the listing) — close the
// transport so the stdio subprocess / HTTP connection isn't orphaned. manager.ts catches
// the rejection as an error status, but without this the child process keeps running.
await transport.close().catch(() => {});
throw err;
}
throw lastErr;
}
-9
View File
@@ -76,12 +76,3 @@ export async function disconnectAllMcpServers(): Promise<void> {
await Promise.allSettled(connections.map((c) => c.transport.close()));
connections = [];
}
/** Re-connects to every configured MCP server (e.g. after the user edits `.mcp.json` or restarts a
* server process), swapping in the new connections and returning the fresh tool list so the caller
* can rebuild the session's toolset. Equivalent to a fresh `connectConfiguredMcpServers` call,
* exposed separately so callers can name the intent ("reconnect") without implying the first-run
* setup path. */
export async function reconnectMcpServers(cwd: string): Promise<ToolDef[]> {
return connectConfiguredMcpServers(cwd);
}
+130
View File
@@ -0,0 +1,130 @@
import { describe, expect, it } from "vitest";
import { PermissionManager } from "./permissionManager.js";
import type { PermissionRule } from "./types.js";
describe("PermissionManager rules", () => {
describe("checkRules", () => {
it("returns null when no rule matches the tool", () => {
const pm = new PermissionManager([{ tool: "bash", allow: true }]);
expect(pm.checkRules("edit_file", { command: "ls" })).toBeNull();
});
it("returns allow for a matching allow rule without argPattern", () => {
const pm = new PermissionManager([{ tool: "bash", allow: true }]);
expect(pm.checkRules("bash", { command: "rm -rf /" })).toBe("allow");
});
it("returns deny for a matching deny rule without argPattern", () => {
const pm = new PermissionManager([{ tool: "bash", allow: false }]);
expect(pm.checkRules("bash", { command: "ls" })).toBe("deny");
});
it("deny wins over allow regardless of order", () => {
const allowFirst: PermissionRule[] = [
{ tool: "bash", allow: true },
{ tool: "bash", allow: false },
];
const denyFirst: PermissionRule[] = [
{ tool: "bash", allow: false },
{ tool: "bash", allow: true },
];
expect(new PermissionManager(allowFirst).checkRules("bash", {})).toBe("deny");
expect(new PermissionManager(denyFirst).checkRules("bash", {})).toBe("deny");
});
it("deny with argPattern only denies matching args, falling through otherwise", () => {
const pm = new PermissionManager([{ tool: "bash", argPattern: "rm\\s+-rf", allow: false }]);
expect(pm.checkRules("bash", { command: "rm -rf /" })).toBe("deny");
// Non-matching args: no rule fires → null.
expect(pm.checkRules("bash", { command: "ls" })).toBeNull();
});
it("allow with argPattern auto-approves only matching args", () => {
// argPattern is a regex tested against JSON.stringify(args), so for {command:"npm test"} the
// serialized text is {"command":"npm test"} — match the substring (no ^ anchor, which would
// bind to the leading brace).
const pm = new PermissionManager([{ tool: "bash", argPattern: "npm (test|run)", allow: true }]);
expect(pm.checkRules("bash", { command: "npm test" })).toBe("allow");
expect(pm.checkRules("bash", { command: "npm install" })).toBeNull();
});
it("argPattern is tested against JSON.stringify(args), so nested fields match", () => {
const pm = new PermissionManager([{ tool: "edit_file", argPattern: "secret", allow: false }]);
expect(pm.checkRules("edit_file", { path: "/safe", old_string: "secret" })).toBe("deny");
expect(pm.checkRules("edit_file", { path: "/safe", old_string: "public" })).toBeNull();
});
it("an invalid regex in argPattern is skipped, not thrown", () => {
const pm = new PermissionManager([
{ tool: "bash", argPattern: "(", allow: false }, // invalid regex
{ tool: "bash", allow: true },
]);
// The invalid deny rule is skipped, so the allow rule fires.
expect(pm.checkRules("bash", { command: "ls" })).toBe("allow");
});
it("treats undefined args as an empty string for pattern matching", () => {
const pm = new PermissionManager([{ tool: "bash", argPattern: "^$", allow: true }]);
expect(pm.checkRules("bash", undefined)).toBe("allow");
});
});
describe("isAutoApproved", () => {
it("auto-approves when an allow rule matches", () => {
const pm = new PermissionManager([{ tool: "bash", argPattern: "npm test", allow: true }]);
expect(pm.isAutoApproved("bash", { command: "npm test" })).toBe(true);
expect(pm.isAutoApproved("bash", { command: "rm -rf" })).toBe(false);
});
it("does not auto-approve when a deny rule matches (deny handled in the gate, not here)", () => {
const pm = new PermissionManager([{ tool: "bash", allow: false }]);
expect(pm.isAutoApproved("bash", {})).toBe(false);
});
it("auto-edit mode auto-approves write_file/edit_file only", () => {
const pm = new PermissionManager([]);
pm.setMode("auto-edit");
expect(pm.isAutoApproved("write_file", { path: "x" })).toBe(true);
expect(pm.isAutoApproved("edit_file", { path: "x" })).toBe(true);
// bash is mutating but not a file-edit tool — still prompts even in auto-edit.
expect(pm.isAutoApproved("bash", { command: "ls" })).toBe(false);
});
it("auto-accept mode also only covers file-edit tools, not bash/git", () => {
const pm = new PermissionManager([]);
pm.setMode("auto-accept");
expect(pm.isAutoApproved("write_file", {})).toBe(true);
expect(pm.isAutoApproved("bash", {})).toBe(false);
});
it("default mode falls back to the session-allowed list", () => {
const pm = new PermissionManager([]);
pm.allowForSession("edit_file");
expect(pm.isAutoApproved("edit_file")).toBe(true);
expect(pm.isAutoApproved("write_file")).toBe(false);
});
it("an allow rule overrides even default mode (no session-allow needed)", () => {
const pm = new PermissionManager([{ tool: "bash", allow: true }]);
expect(pm.isAutoApproved("bash", {})).toBe(true);
});
});
describe("setRules / listRules", () => {
it("setRules replaces the active rule set", () => {
const pm = new PermissionManager([{ tool: "bash", allow: true }]);
pm.setRules([{ tool: "edit_file", allow: false }]);
expect(pm.checkRules("bash", {})).toBeNull();
expect(pm.checkRules("edit_file", {})).toBe("deny");
});
it("listRules returns a defensive copy", () => {
const pm = new PermissionManager([{ tool: "bash", allow: true }]);
const out = pm.listRules();
expect(out).toEqual([{ tool: "bash", allow: true }]);
out.push({ tool: "edit_file", allow: false });
// Mutating the returned array must not affect the manager.
expect(pm.listRules()).toEqual([{ tool: "bash", allow: true }]);
});
});
});
+47 -8
View File
@@ -1,9 +1,14 @@
import type { PermissionMode } from "./types.js";
import type { PermissionMode, PermissionRule } from "./types.js";
import { AUTO_EDIT_TOOLS } from "./types.js";
export class PermissionManager {
private allowedForSession = new Set<string>();
private mode: PermissionMode = "default";
private rules: PermissionRule[] = [];
constructor(rules: PermissionRule[] = []) {
this.rules = rules;
}
getMode(): PermissionMode {
return this.mode;
@@ -13,14 +18,48 @@ export class PermissionManager {
this.mode = mode;
}
/** Replace the active permission-rule set (used after loading/merging user + project rules). */
setRules(rules: PermissionRule[]): void {
this.rules = rules;
}
listRules(): PermissionRule[] {
return this.rules.map((r) => ({ ...r }));
}
/** Evaluate the persistent rules against a tool call. Returns `"deny"` if any deny rule matches
* (deny wins regardless of order or layer), `"allow"` if an allow rule matches and no deny does,
* or `null` when no rule matches (the caller falls through to mode/session logic and prompting).
* `argPattern` (when present) is a regex tested against JSON.stringify(args); an invalid regex is
* skipped rather than crashing the turn. */
checkRules(toolName: string, args: unknown): "deny" | "allow" | null {
const serialized = args === undefined ? "" : JSON.stringify(args);
let allowMatch = false;
for (const r of this.rules) {
if (r.tool !== toolName) continue;
if (r.argPattern !== undefined) {
try {
if (!new RegExp(r.argPattern).test(serialized)) continue;
} catch {
// Invalid regex in a rule — skip it rather than blocking every call to this tool.
continue;
}
}
if (!r.allow) return "deny";
allowMatch = true;
}
return allowMatch ? "allow" : null;
}
/** Check whether a mutating tool should be auto-approved (no confirmation needed). */
isAutoApproved(toolName: string): boolean {
// auto-accept approves EVERY mutating tool (bash, git_commit, write_file, edit_file, ...) —
// the "I trust everything, don't ask" mode. auto-edit is the narrower "approve file edits only"
// mode, auto-approving just the file-edit tools so a user can batch-edit without approving each
// one but still gate dangerous shell/git operations. default falls back to the session-allowed
// list (tools the user approved "for this session" in a prior prompt).
if (this.mode === "auto-accept") return true;
isAutoApproved(toolName: string, args?: unknown): boolean {
// An explicit allow rule auto-approves (a deny rule is handled separately, in the gate, before
// this is reached — so we only need to check for "allow" here).
if (this.checkRules(toolName, args) === "allow") return true;
// auto-accept only covers the same file-edit tools as auto-edit, not arbitrary mutating tools
// such as bash or git_commit. This prevents a user who intended "approve edits" from silently
// approving every dangerous operation.
if (this.mode === "auto-accept" && AUTO_EDIT_TOOLS.has(toolName)) return true;
if (this.mode === "auto-edit" && AUTO_EDIT_TOOLS.has(toolName)) return true;
// "default" — check session-allowed list
return this.allowedForSession.has(toolName);
+13
View File
@@ -5,6 +5,19 @@ export type PermissionMode = "default" | "auto-edit" | "auto-accept" | "plan";
/** Tools that are auto-accepted in "auto-edit" mode. */
export const AUTO_EDIT_TOOLS = new Set(["write_file", "edit_file"]);
/** A persistent permission rule (from user config or project .locode/settings.json) that either
* auto-approves or outright blocks a tool call, without prompting. When `argPattern` is omitted the
* rule matches any call to `tool`; when present it's a regex tested against JSON.stringify(args), so
* e.g. `{ tool: "bash", argPattern: "npm (test|run)", allow: true }` auto-approves test/run
* commands while still prompting for others (the regex matches a substring of the serialized args,
* so don't anchor with `^` — that would bind to the leading `{` of the JSON). Deny rules take
* precedence over allow rules. */
export interface PermissionRule {
tool: string;
argPattern?: string;
allow: boolean;
}
export type ConfirmFn = (opts: {
toolName: string;
args: unknown;
-87
View File
@@ -1,87 +0,0 @@
import { mkdtempSync, readFileSync, rmSync } from "node:fs";
import os from "node:os";
import path from "node:path";
import { describe, it, expect, afterEach, beforeEach } from "vitest";
import type { ChatCompletionMessageParam } from "openai/resources/chat/completions";
import { exportSession, sessionToJson, sessionToMarkdown, defaultExportFilename } from "./exportSession.js";
const meta = { model: "test-model", createdAt: "2026-01-01T00:00:00.000Z" };
const messages: ChatCompletionMessageParam[] = [
{ role: "user", content: "hello" },
{ role: "assistant", content: "let me check", tool_calls: [{ id: "call_1", type: "function", function: { name: "read_file", arguments: '{"path":"a.ts"}' } }] },
{ role: "tool", tool_call_id: "call_1", content: "file contents" },
{ role: "assistant", content: "done" },
];
describe("sessionToMarkdown (full transcript)", () => {
it("includes user text, assistant text, tool calls, and tool results", () => {
const md = sessionToMarkdown(messages, meta);
expect(md).toContain("### You");
expect(md).toContain("hello");
expect(md).toContain("let me check");
expect(md).toContain("#### Tool calls");
expect(md).toContain('"name": "read_file"');
expect(md).toContain("#### Tool result");
expect(md).toContain("file contents");
expect(md).toContain("done");
});
it("does not drop a tool-only assistant turn (no text)", () => {
const md = sessionToMarkdown(
[{ role: "assistant", content: null, tool_calls: [{ id: "c", type: "function", function: { name: "grep", arguments: "{}" } }] } as ChatCompletionMessageParam],
meta,
);
expect(md).toContain("#### Tool calls");
expect(md).toContain('"name": "grep"');
});
});
describe("sessionToJson", () => {
it("produces a JSON object with meta, exportedAt, and the verbatim messages", () => {
const json = sessionToJson(messages, meta);
const parsed = JSON.parse(json);
expect(parsed.model).toBe("test-model");
expect(parsed.exportedAt).toBeTruthy();
expect(parsed.messages).toHaveLength(4);
expect(parsed.messages[1].tool_calls[0].function.name).toBe("read_file");
});
});
describe("defaultExportFilename", () => {
it("defaults to a .md extension", () => {
expect(defaultExportFilename()).toMatch(/\.md$/);
});
it("uses .json for the json format", () => {
expect(defaultExportFilename("json")).toMatch(/\.json$/);
});
});
describe("exportSession", () => {
let cwd: string;
beforeEach(() => {
cwd = mkdtempSync(path.join(os.tmpdir(), "locode-export-"));
});
afterEach(() => {
rmSync(cwd, { recursive: true, force: true });
});
it("writes a markdown file by default", async () => {
const resolved = await exportSession(messages, meta, cwd, "out.md");
const content = readFileSync(resolved, "utf-8");
expect(content).toContain("# locode conversation");
expect(content).toContain("hello");
});
it("writes a JSON file when format is json", async () => {
const resolved = await exportSession(messages, meta, cwd, "out.json", "json");
const content = readFileSync(resolved, "utf-8");
const parsed = JSON.parse(content);
expect(parsed.messages).toHaveLength(4);
});
it("auto-generates a filename with the right extension when none given", async () => {
const resolved = await exportSession(messages, meta, cwd, undefined, "json");
expect(resolved).toMatch(/\.json$/);
});
});
+17 -58
View File
@@ -3,41 +3,15 @@ import path from "node:path";
import type { ChatCompletionMessageParam } from "openai/resources/chat/completions";
import { writeFileAtomic } from "../utils/writeFileAtomic.js";
/** Renders one message as a markdown section for a FULL transcript export — including tool
* calls and their results, which the old prose-only export dropped. A tool-call assistant turn
* lists each call as a fenced JSON block; a tool-result message is rendered as a fenced result.
* Multimodal user content (text + image parts) is reduced to its text parts plus an
* `[image attached]` placeholder. Returns null only for genuinely empty turns. */
/** Mirrors the filtering used when replaying a resumed session (see App.tsx initSessionFromRecord):
* only plain user/assistant text turns are human-readable — raw tool-call/tool-result payloads
* and fallback-mode `tool_result` blocks are internal bookkeeping, not conversation content. */
function messageSection(m: ChatCompletionMessageParam): string | null {
if (m.role === "user") {
if (typeof m.content === "string") {
if (m.content.startsWith("```tool_result")) {
// A fallback-mode tool result block — render it verbatim under a Tool result heading.
return `#### Tool result\n\n${m.content}`;
}
return `### You\n\n${m.content}`;
}
if (Array.isArray(m.content)) {
const parts = m.content.map((p) => (p.type === "text" ? p.text : "[image attached]")).join("\n");
return parts.trim() ? `### You\n\n${parts}` : null;
}
return null;
if (m.role === "user" && typeof m.content === "string" && !m.content.startsWith("```tool_result")) {
return `### You\n\n${m.content}`;
}
if (m.role === "assistant") {
const text = typeof m.content === "string" ? m.content : "";
const calls = (m as { tool_calls?: { id: string; function: { name: string; arguments: string } }[] }).tool_calls;
const parts: string[] = [];
if (text.trim()) parts.push(`### Assistant\n\n${text}`);
if (calls && calls.length) {
const block = calls.map((c) => `{"name": "${c.function.name}", "arguments": ${c.function.arguments}}`).join("\n");
parts.push(`#### Tool calls\n\n` + "```json\n" + block + "\n```");
}
return parts.length ? parts.join("\n\n") : null;
}
if (m.role === "tool") {
const tm = m as { content?: string; tool_call_id?: string };
const body = typeof tm.content === "string" ? tm.content : JSON.stringify(tm.content);
return `#### Tool result${tm.tool_call_id ? ` (${tm.tool_call_id})` : ""}\n\n` + "```\n" + body + "\n```";
if (m.role === "assistant" && typeof m.content === "string" && m.content) {
return `### Assistant\n\n${m.content}`;
}
return null;
}
@@ -47,8 +21,6 @@ export interface ExportMeta {
createdAt: string;
}
export type ExportFormat = "markdown" | "json";
export function sessionToMarkdown(messages: ChatCompletionMessageParam[], meta: ExportMeta): string {
const header = [
"# locode conversation",
@@ -61,35 +33,22 @@ export function sessionToMarkdown(messages: ChatCompletionMessageParam[], meta:
return [header, ...sections].join("\n\n");
}
/** A JSON export is the full record (messages verbatim + metadata), suitable for cross-machine
* replay/sharing or feeding into another tool. The markdown export is for humans. */
export function sessionToJson(messages: ChatCompletionMessageParam[], meta: ExportMeta): string {
return JSON.stringify({ ...meta, exportedAt: new Date().toISOString(), messages }, null, 2);
}
export function defaultExportFilename(format: ExportFormat = "markdown"): string {
export function defaultExportFilename(): string {
const stamp = new Date().toISOString().replace(/[:.]/g, "-");
return `locode-export-${stamp}.${format === "json" ? "json" : "md"}`;
return `locode-export-${stamp}.md`;
}
/** Writes the conversation to a file and returns the resolved absolute path. `target` may be a bare
* filename, a relative path, or an absolute path; a bare directory (or nothing at all) falls back
* to an auto-generated filename inside `cwd`. `format` selects a human markdown transcript
* (default, now including tool calls/results) or a machine-readable JSON dump. Written atomically
/** Writes the conversation to a markdown file and returns the resolved absolute path.
* `target` may be a bare filename, a relative path, or an absolute path; a bare directory
* (or nothing at all) falls back to an auto-generated filename inside `cwd`. Written atomically
* (temp file + rename), matching sessionStore's saves, so a crash mid-export can't leave a
* truncated file. */
export async function exportSession(
messages: ChatCompletionMessageParam[],
meta: ExportMeta,
cwd: string,
target?: string,
format: ExportFormat = "markdown",
): Promise<string> {
const filename = target?.trim() || defaultExportFilename(format);
export async function exportSession(messages: ChatCompletionMessageParam[], meta: ExportMeta, cwd: string, target?: string): Promise<string> {
const filename = target?.trim() || defaultExportFilename();
let resolved = path.isAbsolute(filename) ? filename : path.resolve(cwd, filename);
if (existsSync(resolved) && statSync(resolved).isDirectory()) {
resolved = path.join(resolved, defaultExportFilename(format));
resolved = path.join(resolved, defaultExportFilename());
}
await writeFileAtomic(resolved, format === "json" ? sessionToJson(messages, meta) : sessionToMarkdown(messages, meta));
await writeFileAtomic(resolved, sessionToMarkdown(messages, meta));
return resolved;
}
}
-81
View File
@@ -1,81 +0,0 @@
import { describe, expect, it } from "vitest";
import type { ChatCompletionMessageParam } from "openai/resources/chat/completions";
import { buildReplayHistory } from "./replayHistory.js";
describe("buildReplayHistory", () => {
it("replays plain user/assistant text turns", () => {
const messages: ChatCompletionMessageParam[] = [
{ role: "user", content: "hi" },
{ role: "assistant", content: "hello there" },
];
expect(buildReplayHistory(messages)).toMatchObject([
{ kind: "user", text: "hi" },
{ kind: "assistant", text: "hello there" },
]);
});
it("replays native-mode tool calls and results, correlating the result's name via tool_call_id", () => {
const messages: ChatCompletionMessageParam[] = [
{ role: "user", content: "read foo.txt" },
{
role: "assistant",
content: null,
tool_calls: [
{ id: "call_1", type: "function", function: { name: "read_file", arguments: JSON.stringify({ path: "foo.txt" }) } },
],
},
{ role: "tool", tool_call_id: "call_1", content: JSON.stringify({ totalLines: 3 }) },
{ role: "assistant", content: "It has 3 lines." },
];
expect(buildReplayHistory(messages)).toMatchObject([
{ kind: "user", text: "read foo.txt" },
{ kind: "tool_call", label: "Read(foo.txt)" },
{ kind: "tool_result", summary: "Read 3 lines", isError: false },
{ kind: "assistant", text: "It has 3 lines." },
]);
});
it("marks a native-mode tool error result", () => {
const messages: ChatCompletionMessageParam[] = [
{
role: "assistant",
content: null,
tool_calls: [{ id: "call_1", type: "function", function: { name: "bash", arguments: "{}" } }],
},
{ role: "tool", tool_call_id: "call_1", content: JSON.stringify({ error: "command not found" }) },
];
expect(buildReplayHistory(messages)).toMatchObject([
{ kind: "tool_call", label: "Bash()" },
{ kind: "tool_result", summary: "command not found", isError: true },
]);
});
it("replays fallback-mode tool_call blocks embedded in assistant text", () => {
const messages: ChatCompletionMessageParam[] = [
{
role: "assistant",
content: 'Let me check.\n```tool_call\n{"name":"read_file","arguments":{"path":"foo.txt"}}\n```',
},
{
role: "user",
content: '```tool_result\n{"name":"read_file","result":{"totalLines":3}}\n```',
},
];
expect(buildReplayHistory(messages)).toMatchObject([
{ kind: "assistant", text: "Let me check." },
{ kind: "tool_call", label: "Read(foo.txt)" },
{ kind: "tool_result", summary: "Read 3 lines", isError: false },
]);
});
it("drops the empty assistant text bubble when a native tool call has no accompanying prose", () => {
const messages: ChatCompletionMessageParam[] = [
{
role: "assistant",
content: null,
tool_calls: [{ id: "call_1", type: "function", function: { name: "grep", arguments: "{}" } }],
},
];
expect(buildReplayHistory(messages)).toMatchObject([{ kind: "tool_call", label: "Grep()" }]);
});
});
-108
View File
@@ -1,108 +0,0 @@
import type { ChatCompletionMessageParam } from "openai/resources/chat/completions";
import { parseFallbackToolCalls } from "../toolcalling/fallbackParser.js";
import { nextId, type HistoryItem, type NewHistoryItem } from "../ui/ink/types.js";
import { formatCallLabel, summarizeToolResult } from "../ui/toolSummary.js";
const TOOL_CALL_BLOCK_RE = /```tool_call\s*[\s\S]*?```/g;
const TOOL_RESULT_BLOCK_RE = /^```tool_result\n([\s\S]*?)\n```$/;
/** A message's `content` can be a plain string or an array of content parts (text/image_url) —
* see pushToolResultMessage in agent/loop.ts. Only the text part matters for replay display. */
function textContent(content: ChatCompletionMessageParam["content"]): string | null {
if (typeof content === "string") return content;
if (Array.isArray(content)) {
const part = content.find((p): p is { type: "text"; text: string } => (p as { type?: string }).type === "text");
return part?.text ?? null;
}
return null;
}
function isErrorResult(result: unknown): boolean {
return !!(result && typeof result === "object" && "error" in (result as object));
}
/** Rebuilds the tool_call/tool_result HistoryItems a resumed session's transcript is otherwise
* missing (see App.tsx initSessionFromRecord) — the persisted record (SessionRecord.messages) is
* the raw OpenAI-shape history, which carries everything needed (tool name, arguments, result)
* even though it was never saved as a pre-rendered display label. Handles both native mode
* (assistant `tool_calls` + matching `tool`-role messages, correlated by `tool_call_id`) and
* fallback mode (` ```tool_call``` ` blocks embedded in assistant text, ` ```tool_result``` `
* blocks embedded in user text — see fallbackParser.ts and pushToolResultMessage). */
export function buildReplayHistory(messages: ChatCompletionMessageParam[]): HistoryItem[] {
const items: HistoryItem[] = [];
// Native mode: `tool` messages only carry a tool_call_id, not the tool's name — remember the
// name from the assistant message that made the call so its later result can be labeled.
const pendingCallNames = new Map<string, string>();
const push = (item: NewHistoryItem) => items.push({ id: nextId(), ...item } as HistoryItem);
for (const m of messages) {
if (m.role === "user") {
const text = textContent(m.content);
if (text === null) continue;
const fallbackResult = TOOL_RESULT_BLOCK_RE.exec(text.trim());
if (fallbackResult) {
try {
const { name, result } = JSON.parse(fallbackResult[1]!) as { name: string; result: unknown };
push({ kind: "tool_result", summary: summarizeToolResult(name, result), isError: isErrorResult(result) });
} catch {
// Malformed persisted block (shouldn't happen since we wrote it) — drop rather than
// show the raw JSON to the user.
}
continue;
}
if (text) push({ kind: "user", text });
continue;
}
if (m.role === "assistant") {
const toolCalls = m.tool_calls;
if (toolCalls?.length) {
const text = textContent(m.content);
if (text) push({ kind: "assistant", text });
for (const call of toolCalls) {
if (call.type !== "function") continue;
let args: unknown = {};
try {
args = JSON.parse(call.function.arguments);
} catch {
// Leave args as {} — formatCallLabel degrades gracefully for a missing field.
}
pendingCallNames.set(call.id, call.function.name);
push({ kind: "tool_call", label: formatCallLabel(call.function.name, args) });
}
continue;
}
const text = textContent(m.content);
if (!text) continue;
const parsed = parseFallbackToolCalls(text);
if (parsed.calls.length) {
const stripped = text.replace(TOOL_CALL_BLOCK_RE, "").trim();
if (stripped) push({ kind: "assistant", text: stripped });
for (const call of parsed.calls) {
push({ kind: "tool_call", label: formatCallLabel(call.name, call.arguments) });
}
} else {
push({ kind: "assistant", text });
}
continue;
}
if (m.role === "tool") {
const name = pendingCallNames.get(m.tool_call_id) ?? "unknown";
const text = textContent(m.content);
let result: unknown = text;
if (text) {
try {
result = JSON.parse(text);
} catch {
result = text;
}
}
push({ kind: "tool_result", summary: summarizeToolResult(name, result), isError: isErrorResult(result) });
}
}
return items;
}
+2 -36
View File
@@ -1,14 +1,6 @@
import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
import { existsSync, mkdirSync, rmSync, unlinkSync, writeFileSync } from "node:fs";
import { afterEach, beforeEach, describe, expect, it } from "vitest";
import { existsSync, mkdirSync, rmSync, writeFileSync } from "node:fs";
import path from "node:path";
// unlinkSync is mocked (default: pass-through to the real implementation) only so the "can't
// actually delete the file" test below can make a single call fail — every other test's calls to
// unlinkSync still hit the real filesystem via this same mock.
vi.mock("node:fs", async (importOriginal) => {
const actual = await importOriginal<typeof import("node:fs")>();
return { ...actual, unlinkSync: vi.fn(actual.unlinkSync) };
});
import envPaths from "env-paths";
import {
deleteSession,
@@ -78,32 +70,6 @@ describe("sessionStore", () => {
expect(listSessions()).toHaveLength(0);
});
it("rebuilds the summary index from disk when no index file exists yet", () => {
writeFileSync(path.join(dir, "manual-session.json"), JSON.stringify(makeRecord("manual-session")));
expect(listSessions()).toHaveLength(1);
expect(listSessions()[0]?.id).toBe("manual-session");
});
it("reports failure (not success) when the file can't actually be deleted", async () => {
// Regression: a real unlink failure (e.g. Windows EPERM/EBUSY from a file lock) used to still
// return true and drop the entry from the index — reporting success while orphaning the file
// on disk with no way to reference it again.
await saveSession(makeRecord("locked-session"));
vi.mocked(unlinkSync).mockImplementationOnce(() => {
throw Object.assign(new Error("EBUSY: resource busy or locked"), { code: "EBUSY" });
});
expect(deleteSession("locked-session")).toBe(false);
expect(loadSession("locked-session")?.id).toBe("locked-session");
expect(listSessions()).toHaveLength(1);
});
it("self-heals when a session file is deleted outside of deleteSession()", async () => {
await saveSession(makeRecord("will-vanish"));
expect(listSessions()).toHaveLength(1);
unlinkSync(path.join(dir, "will-vanish.json"));
expect(listSessions()).toHaveLength(0);
});
it("deriveTitle extracts the first user message", () => {
const title = deriveTitle([
{ role: "system", content: "sys" },
+26 -118
View File
@@ -1,9 +1,9 @@
import envPaths from "env-paths";
import { existsSync, readdirSync, readFileSync, renameSync, unlinkSync, writeFileSync } from "node:fs";
import { existsSync, readdirSync, readFileSync, unlinkSync } from "node:fs";
import path from "node:path";
import type { ChatCompletionMessageParam } from "openai/resources/chat/completions";
import type { ToolCallMode } from "../backend/capabilityProbe.js";
import type { TaskStoreSnapshot } from "../tools/task.js";
import type { PermissionMode } from "../permissions/types.js";
import { writeFileAtomic } from "../utils/writeFileAtomic.js";
/** A saved conversation. `messages` excludes the system prompt — it's rebuilt fresh from the
@@ -16,11 +16,10 @@ export interface SessionRecord {
baseURL: string;
model: string;
mode: ToolCallMode;
/** Permission mode at save time, so plan/auto-edit/auto-accept survives a resume instead of
* always resetting to default. Omitted by older saved sessions — treated as "default". */
permissionMode?: PermissionMode;
messages: ChatCompletionMessageParam[];
/** Tool names the user approved "for this session" — preserved across resume. */
allowedTools?: string[];
/** Persisted task store snapshot so tasks survive session resume. */
tasks?: TaskStoreSnapshot;
}
export interface SessionSummary {
@@ -50,92 +49,6 @@ function filePath(id: string): string {
return path.join(dir, `${safeSessionId(id)}.json`);
}
const INDEX_FILENAME = "_index.json";
function indexFilePath(): string {
return path.join(dir, INDEX_FILENAME);
}
function summarize(record: SessionRecord): SessionSummary {
return {
id: record.id,
updatedAt: record.updatedAt,
title: deriveTitle(record.messages),
model: record.model,
baseURL: record.baseURL,
messageCount: record.messages.length,
};
}
/** Reads the on-disk summary index and reconciles it against the actual session files, so a stale or
* missing index self-heals instead of ever going wrong: entries whose file was deleted (by this or
* another locode process) are dropped, and files present on disk but missing from the index (a fresh
* install, a crash before the last index write, another process's save racing this one) are parsed
* individually. This keeps the common case to O(session count) stat calls instead of O(total
* transcript bytes) — listSessions() used to JSON.parse every saved session in full just to read 6
* summary fields off each one. */
function readIndex(): Map<string, SessionSummary> {
let index = new Map<string, SessionSummary>();
if (existsSync(indexFilePath())) {
try {
const entries = JSON.parse(readFileSync(indexFilePath(), "utf-8")) as SessionSummary[];
index = new Map(entries.map((e) => [e.id, e]));
} catch (err) {
// eslint-disable-next-line no-console
console.warn("[sessionStore] failed to parse index file, rebuilding:", err);
index = new Map();
}
}
for (const id of index.keys()) {
if (!existsSync(filePath(id))) index.delete(id);
}
const known = new Set([...index.keys()].map((id) => safeSessionId(id)));
if (existsSync(dir)) {
for (const entry of readdirSync(dir)) {
if (!entry.endsWith(".json") || entry === INDEX_FILENAME) continue;
const stem = entry.slice(0, -".json".length);
if (known.has(stem)) continue;
try {
const record = JSON.parse(readFileSync(path.join(dir, entry), "utf-8")) as SessionRecord;
index.set(record.id, summarize(record));
} catch (err) {
// Skip corrupt/partial session files — but log so disk issues aren't silent.
// eslint-disable-next-line no-console
console.warn(`[sessionStore] skipping corrupt session file ${entry}:`, err);
}
}
}
return index;
}
/** Best-effort atomic write of the index. A failed write just means the next readIndex() call
* re-parses whatever files it doesn't recognize yet — never incorrect data, only a missed
* optimization. */
function writeIndex(index: Map<string, SessionSummary>): void {
const tmp = `${indexFilePath()}.${process.pid}.${Date.now()}.tmp`;
try {
writeFileSync(tmp, JSON.stringify([...index.values()]));
renameSync(tmp, indexFilePath());
} catch (err) {
try {
unlinkSync(tmp);
} catch {
// tmp may not have been created if the write itself failed
}
// eslint-disable-next-line no-console
console.warn("[sessionStore] failed to write session index:", err);
}
}
/** Updates one entry in the persisted index. Two saves for different sessions racing this can lose
* one's index write, but never lose data: readIndex() picks up any on-disk session file it doesn't
* recognize, so the loser just costs the next listSessions() call one extra parse. */
function updateIndexEntry(record: SessionRecord): void {
const index = readIndex();
index.set(record.id, summarize(record));
writeIndex(index);
}
export function deriveTitle(messages: ChatCompletionMessageParam[]): string {
const first = messages.find((m) => m.role === "user");
const text = first && typeof first.content === "string" ? first.content.trim() : "";
@@ -154,8 +67,7 @@ const saveQueues = new Map<string, Promise<void>>();
* otherwise load as "no session found" (see loadSession, which treats an unparseable file as absent). */
export async function saveSession(record: SessionRecord): Promise<void> {
const file = filePath(record.id);
const write = () =>
writeFileAtomic(file, JSON.stringify(record, null, 2)).then(() => updateIndexEntry(record));
const write = () => writeFileAtomic(file, JSON.stringify(record, null, 2));
// Run whether or not the previous save rejected, so one failure can't stall the chain.
const prev = saveQueues.get(record.id);
const next = (prev ?? Promise.resolve()).then(write, write);
@@ -177,19 +89,31 @@ export function loadSession(id: string): SessionRecord | undefined {
if (!existsSync(file)) return undefined;
try {
return JSON.parse(readFileSync(file, "utf-8")) as SessionRecord;
} catch (err) {
// Corrupt or unreadable session file — treat as absent, but log so disk issues
// aren't completely silent. Callers can't distinguish "no file" from "corrupt file",
// but at least the log preserves the reason.
// eslint-disable-next-line no-console
console.warn(`[sessionStore] failed to load session ${id}, treating as absent:`, err);
} catch {
return undefined;
}
}
export function listSessions(): SessionSummary[] {
if (!existsSync(dir)) return [];
return [...readIndex().values()].sort((a, b) => b.updatedAt.localeCompare(a.updatedAt));
const summaries: SessionSummary[] = [];
for (const entry of readdirSync(dir)) {
if (!entry.endsWith(".json")) continue;
try {
const record = JSON.parse(readFileSync(path.join(dir, entry), "utf-8")) as SessionRecord;
summaries.push({
id: record.id,
updatedAt: record.updatedAt,
title: deriveTitle(record.messages),
model: record.model,
baseURL: record.baseURL,
messageCount: record.messages.length,
});
} catch {
// Skip corrupt/partial session files
}
}
return summaries.sort((a, b) => b.updatedAt.localeCompare(a.updatedAt));
}
export function mostRecentSessionId(): string | undefined {
@@ -199,22 +123,6 @@ export function mostRecentSessionId(): string | undefined {
export function deleteSession(id: string): boolean {
const file = filePath(id);
if (!existsSync(file)) return false;
try {
unlinkSync(file);
} catch (err) {
// ENOENT is harmless (race with another process) — the end state (file gone) is what we
// wanted anyway, so fall through and report success. Any other error (e.g. EPERM/EBUSY from
// a file lock, common on Windows) means the file is still on disk — report failure and leave
// the index entry alone, or listSessions()/resume would silently orphan a file no one could
// reference again (removed from the index, but never actually deleted).
if ((err as NodeJS.ErrnoException).code !== "ENOENT") {
// eslint-disable-next-line no-console
console.warn(`[sessionStore] failed to delete session file ${file}:`, err);
return false;
}
}
const index = readIndex();
index.delete(id);
writeIndex(index);
unlinkSync(file);
return true;
}
+1 -1
View File
@@ -36,7 +36,7 @@ export function buildPluginAgentTool(agent: PluginAgentDef): ToolDef<{ prompt: s
{ description: agent.name, prompt },
{ systemPrompt: agent.systemPrompt, toolNames: agent.tools },
);
return { agent: agent.name, result };
return { agent: agent.name, result: result.result, ...(result.resumable ? { agentId: result.agentId } : {}) };
},
};
}
-17
View File
@@ -46,21 +46,4 @@ describe("loadPlugin", () => {
expect(plugin.agents).toHaveLength(1);
expect(plugin.agents[0]).toMatchObject({ name: "review", description: "Code reviewer" });
});
it("parses a command's allowed-tools frontmatter, translating Claude Code tool names", () => {
mkdirSync(path.join(tempDir, ".claude-plugin"), { recursive: true });
writeFileSync(path.join(tempDir, ".claude-plugin", "plugin.json"), JSON.stringify({ name: "test-plugin" }));
mkdirSync(path.join(tempDir, "commands"), { recursive: true });
writeFileSync(
path.join(tempDir, "commands", "readonly.md"),
"---\ndescription: Look but don't touch\nallowed-tools: Read, Grep\n---\nInvestigate $ARGUMENTS",
);
writeFileSync(path.join(tempDir, "commands", "unrestricted.md"), "---\ndescription: No restriction\n---\nDo $ARGUMENTS");
const plugin = loadPlugin(tempDir);
const readonly = plugin.commands.find((c) => c.name === "readonly");
const unrestricted = plugin.commands.find((c) => c.name === "unrestricted");
expect(readonly?.allowedTools).toEqual(["read_file", "grep"]);
expect(unrestricted?.allowedTools).toBeUndefined();
});
});
+5 -13
View File
@@ -20,17 +20,6 @@ function listMarkdownFiles(dir: string): string[] {
return readdirSync(dir).filter((f) => f.endsWith(".md"));
}
/** Parses a comma-separated tool-name list from frontmatter (agents' `tools`, commands'
* `allowed-tools`) and translates each from Claude Code's built-in names to locode's. Returns
* undefined for an absent/empty field so callers can treat that as "no restriction". */
function parseToolList(field: string | undefined): string[] | undefined {
const tools = field
?.split(",")
.map((t) => resolveToolName(t.trim()))
.filter(Boolean);
return tools?.length ? tools : undefined;
}
function loadCommands(pluginRoot: string, pluginName: string): PluginCommand[] {
const dir = path.join(pluginRoot, "commands");
return listMarkdownFiles(dir).map((entry) => {
@@ -41,7 +30,6 @@ function loadCommands(pluginRoot: string, pluginName: string): PluginCommand[] {
description: frontmatter.description,
argumentHint: frontmatter["argument-hint"],
template: body,
allowedTools: parseToolList(frontmatter["allowed-tools"]),
};
});
}
@@ -51,11 +39,15 @@ function loadAgents(pluginRoot: string, pluginName: string): PluginAgentDef[] {
return listMarkdownFiles(dir).map((entry) => {
const { frontmatter, body } = parseFrontmatter(readFileSync(path.join(dir, entry), "utf-8"));
const name = frontmatter.name || entry.replace(/\.md$/, "");
const tools = frontmatter.tools
?.split(",")
.map((t) => resolveToolName(t.trim()))
.filter(Boolean);
return {
pluginName,
name,
description: frontmatter.description || name,
tools: parseToolList(frontmatter.tools),
tools: tools?.length ? tools : undefined,
systemPrompt: body,
};
});
+4 -1
View File
@@ -13,7 +13,10 @@ const CLAUDE_TOOL_NAME_MAP: Record<string, string> = {
webfetch: "web_fetch",
websearch: "web_search",
task: "agent",
todowrite: "todo_write",
taskcreate: "task_create",
tasklist: "task_list",
taskget: "task_get",
taskupdate: "task_update",
};
export function resolveToolName(name: string): string {
-4
View File
@@ -16,10 +16,6 @@ export interface PluginCommand {
/** The markdown body — expanded via $ARGUMENTS/$1../$9 (see expandTemplate.ts) and submitted
* as the turn's input when the command is invoked. */
template: string;
/** Tool names this command's turn is restricted to (already translated from Claude Code's
* built-in tool names — see toolNameMap.ts), from frontmatter `allowed-tools`; undefined means
* no restriction (the command runs with the session's full toolset). */
allowedTools?: string[];
}
export interface PluginAgentDef {
+281
View File
@@ -0,0 +1,281 @@
import { describe, expect, it, vi, beforeEach, afterEach } from "vitest";
import { mkdtempSync, rmSync, existsSync, readFileSync } from "node:fs";
import { tmpdir } from "node:os";
import path from "node:path";
import { CronStore, parseCron } from "./cron.js";
import { cronCreateTool, cronDeleteTool, cronListTool, scheduleWakeupTool } from "../tools/cron.js";
import type { ToolContext } from "../tools/types.js";
function ctxWith(store?: CronStore): ToolContext {
return { cwd: "/x", ...(store ? { cronStore: store } : {}) };
}
describe("parseCron", () => {
it("parses a 5-field expression into allowed values", () => {
const spec = parseCron("0 9 * * 1-5");
expect(spec.minute).toEqual([0]);
expect(spec.hour).toEqual([9]);
expect(spec.dom).toEqual(Array.from({ length: 31 }, (_, i) => i + 1));
expect(spec.month).toEqual([1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12]);
expect(spec.dow).toEqual([1, 2, 3, 4, 5]);
expect(spec.domStar).toBe(true);
expect(spec.dowStar).toBe(false);
});
it("supports */N step and comma-lists", () => {
const spec = parseCron("*/15 8-17 * * 0,6");
expect(spec.minute).toEqual([0, 15, 30, 45]);
expect(spec.hour).toEqual([8, 9, 10, 11, 12, 13, 14, 15, 16, 17]);
expect(spec.dow).toEqual([0, 6]);
});
it("throws on a wrong field count", () => {
expect(() => parseCron("0 9 * *")).toThrow(/5 fields/);
expect(() => parseCron("0 9 * * * *")).toThrow(/5 fields/);
});
it("throws on out-of-range values", () => {
expect(() => parseCron("60 9 * * *")).toThrow(/out of range/);
expect(() => parseCron("0 24 * * *")).toThrow(/out of range/);
expect(() => parseCron("0 9 32 * *")).toThrow(/out of range/);
});
it("matches a Date correctly (weekday cron with dom=* → dow governs)", () => {
// Every Monday: dom=* (star), dow=1. 2026-08-17 is a Monday; 2026-08-16 is a Sunday.
const s = parseCron("0 0 * * 1");
const mon = new Date(2026, 7, 17, 0, 0);
const sun = new Date(2026, 7, 16, 0, 0);
// Inline the Vixie-cron matcher logic (the store uses it internally).
const domMatch = (d: Date) => s.dom.includes(d.getDate());
const dowMatch = (d: Date) => s.dow.includes(d.getDay());
const ok = (d: Date) => (s.domStar ? dowMatch(d) : s.dowStar ? domMatch(d) : domMatch(d) || dowMatch(d));
expect(ok(mon)).toBe(true);
expect(ok(sun)).toBe(false);
});
it("fires on EITHER dom OR dow when both are restricted (Vixie semantics)", () => {
// dom=15, dow=0 (Sunday): fires on the 15th of any month OR any Sunday.
const s = parseCron("0 0 15 * 0");
const domMatch = (d: Date) => s.dom.includes(d.getDate());
const dowMatch = (d: Date) => s.dow.includes(d.getDay());
const ok = (d: Date) => (s.domStar ? dowMatch(d) : s.dowStar ? domMatch(d) : domMatch(d) || dowMatch(d));
// 2026-08-16 is Sunday the 16th (not the 15th) → matches via dow.
expect(ok(new Date(2026, 7, 16, 0, 0))).toBe(true);
// 2026-08-15 is Saturday the 15th → matches via dom.
expect(ok(new Date(2026, 7, 15, 0, 0))).toBe(true);
// 2026-08-14 is Friday the 14th → no match.
expect(ok(new Date(2026, 7, 14, 0, 0))).toBe(false);
});
});
describe("CronStore", () => {
it("create validates the cron expression up front", () => {
const s = new CronStore();
expect(() => s.create({ cron: "bad expr", prompt: "x" })).toThrow();
});
it("create/list/delete round-trips jobs with sequential ids", () => {
const s = new CronStore();
const a = s.create({ cron: "0 9 * * *", prompt: "morning" });
const b = s.create({ cron: "0 10 * * *", prompt: "late", recurring: false });
expect(a.id).toBe("c1");
expect(b.id).toBe("c2");
expect(s.list()).toHaveLength(2);
expect(s.delete("c1")).toBe(true);
expect(s.list()).toHaveLength(1);
expect(s.delete("nope")).toBe(false);
});
it("tick fires due recurring jobs (deduped within the minute) and enqueues the prompt", () => {
const s = new CronStore();
const enqueued: string[] = [];
s.start({ enqueue: (p) => enqueued.push(p), isIdle: () => true });
// A cron that matches every minute, so it's definitely due now.
s.create({ cron: "* * * * *", prompt: "tick" });
// Call the private tick via a cast (the real loop uses setInterval).
(s as unknown as { tick: () => void }).tick();
expect(enqueued).toEqual(["tick"]);
// A second tick in the same minute must NOT re-fire (deduped by lastFiredMinute).
(s as unknown as { tick: () => void }).tick();
expect(enqueued).toEqual(["tick"]);
s.stop();
});
it("tick does not fire while isIdle is false", () => {
const s = new CronStore();
const enqueued: string[] = [];
s.start({ enqueue: (p) => enqueued.push(p), isIdle: () => false });
s.create({ cron: "* * * * *", prompt: "tick" });
(s as unknown as { tick: () => void }).tick();
expect(enqueued).toEqual([]);
s.stop();
});
it("one-shot jobs (recurring:false) are deleted after firing once", () => {
const s = new CronStore();
const enqueued: string[] = [];
s.start({ enqueue: (p) => enqueued.push(p), isIdle: () => true });
s.create({ cron: "* * * * *", prompt: "once", recurring: false });
(s as unknown as { tick: () => void }).tick();
expect(enqueued).toEqual(["once"]);
expect(s.list()).toHaveLength(0); // removed after the one fire
s.stop();
});
it("scheduleWakeup clamps delay to [60,3600] and fires once then is removed", () => {
const s = new CronStore();
const enqueued: string[] = [];
s.start({ enqueue: (p) => enqueued.push(p), isIdle: () => true });
const res = s.scheduleWakeup({ delaySeconds: 5, prompt: "wake" }) as { id: string; fireAt: number };
expect(res.id).toMatch(/^w/);
// Clamped to 60s, so a tick right now (well before fireAt) must not enqueue.
(s as unknown as { tick: () => void }).tick();
expect(enqueued).toEqual([]);
// Force the wakeup into the past (mutate the STORE's internal entry, not the listWakeups copy)
// and tick again so the wakeup is now due.
const internal = (s as unknown as { wakeups: Map<string, { fireAt: number }> }).wakeups;
const id = [...internal.keys()][0]!;
internal.get(id)!.fireAt = Date.now() - 1000;
(s as unknown as { tick: () => void }).tick();
expect(enqueued).toEqual(["wake"]);
expect(s.listWakeups()).toHaveLength(0);
s.stop();
});
it("scheduleWakeup with stop:true clears all wakeups", () => {
const s = new CronStore();
s.scheduleWakeup({ delaySeconds: 60, prompt: "a" });
s.scheduleWakeup({ delaySeconds: 120, prompt: "b" });
expect(s.listWakeups()).toHaveLength(2);
const res = s.scheduleWakeup({ delaySeconds: 60, prompt: "", stop: true });
expect(res).toEqual({ stopped: true });
expect(s.listWakeups()).toHaveLength(0);
});
describe("durable persistence", () => {
let dir: string;
beforeEach(() => {
dir = mkdtempSync(path.join(tmpdir(), "locode-cron-test-"));
});
afterEach(() => {
try {
rmSync(dir, { recursive: true, force: true });
} catch {
/* leave for OS temp sweep */
}
});
it("persists durable jobs to scheduled_tasks.json and reloads them on construction", () => {
const s1 = new CronStore(dir);
s1.create({ cron: "0 9 * * *", prompt: "morning", durable: true });
s1.create({ cron: "0 10 * * *", prompt: "ephemeral" }); // not durable
expect(existsSync(path.join(dir, "scheduled_tasks.json"))).toBe(true);
const raw = JSON.parse(readFileSync(path.join(dir, "scheduled_tasks.json"), "utf-8")) as {
jobs: { prompt: string; durable: boolean }[];
};
expect(raw.jobs).toHaveLength(1);
expect(raw.jobs[0]!.prompt).toBe("morning");
// A fresh store pointed at the same dir reloads the durable job only.
const s2 = new CronStore(dir);
expect(s2.list()).toHaveLength(1);
expect(s2.list()[0]!.prompt).toBe("morning");
// The reloaded id shouldn't collide with a new one (seq was bumped past it).
const next = s2.create({ cron: "0 11 * * *", prompt: "next" });
expect(next.id).not.toBe("c1");
});
it("non-durable jobs are NOT persisted", () => {
const s1 = new CronStore(dir);
s1.create({ cron: "0 9 * * *", prompt: "ephemeral", durable: false });
const s2 = new CronStore(dir);
expect(s2.list()).toHaveLength(0);
});
});
});
describe("cron/schedule tools", () => {
it("cron_create returns the job id and snapshot, validating cron syntax", async () => {
const store = new CronStore();
const result = (await cronCreateTool.handler(
{ cron: "0 9 * * 1-5", prompt: "weekday standup", recurring: true },
ctxWith(store),
)) as { id: string; job: { cron: string; recurring: boolean } };
expect(result.id).toBe("c1");
expect(result.job.cron).toBe("0 9 * * 1-5");
});
it("cron_create returns an error for invalid cron syntax", async () => {
const store = new CronStore();
const result = (await cronCreateTool.handler(
{ cron: "not cron", prompt: "x" },
ctxWith(store),
)) as { error: string };
expect(result.error).toMatch(/field|5 fields|range/i);
});
it("cron_create returns an error when no store is available", async () => {
const result = (await cronCreateTool.handler({ cron: "0 9 * * *", prompt: "x" }, ctxWith(undefined))) as {
error: string;
};
expect(result.error).toMatch(/not available/i);
});
it("cron_list returns the jobs; empty (not error) when no store", async () => {
const store = new CronStore();
store.create({ cron: "0 9 * * *", prompt: "x" });
const result = (await cronListTool.handler({}, ctxWith(store))) as { jobs: { id: string }[] };
expect(result.jobs).toHaveLength(1);
const empty = (await cronListTool.handler({}, ctxWith(undefined))) as { jobs: unknown[] };
expect(empty.jobs).toEqual([]);
});
it("cron_delete removes a job and returns { deleted }", async () => {
const store = new CronStore();
store.create({ cron: "0 9 * * *", prompt: "x" });
const result = (await cronDeleteTool.handler({ id: "c1" }, ctxWith(store))) as { deleted: string };
expect(result.deleted).toBe("c1");
expect(store.list()).toHaveLength(0);
});
it("cron_delete returns an error for an unknown id", async () => {
const store = new CronStore();
const result = (await cronDeleteTool.handler({ id: "c99" }, ctxWith(store))) as { error: string };
expect(result.error).toMatch(/not found/i);
});
it("schedule_wakeup returns a wakeup id, or { stopped } when stop is true", async () => {
const store = new CronStore();
const result = (await scheduleWakeupTool.handler(
{ delaySeconds: 120, prompt: "check back" },
ctxWith(store),
)) as { id: string; fireAt: number };
expect(result.id).toMatch(/^w/);
const stop = (await scheduleWakeupTool.handler(
{ delaySeconds: 60, prompt: "", stop: true },
ctxWith(store),
)) as { stopped: boolean };
expect(stop.stopped).toBe(true);
});
it("all cron/schedule tools are non-mutating (no confirmation prompt)", () => {
expect(cronCreateTool.mutating).toBe(false);
expect(cronListTool.mutating).toBe(false);
expect(cronDeleteTool.mutating).toBe(false);
expect(scheduleWakeupTool.mutating).toBe(false);
});
});
describe("cron tool schema validation", () => {
it("requires a cron expression and prompt on cron_create", () => {
expect(() => cronCreateTool.schema.parse({ cron: "", prompt: "x" })).toThrow();
expect(() => cronCreateTool.schema.parse({ cron: "0 9 * * *", prompt: "" })).toThrow();
});
it("requires an id on cron_delete", () => {
expect(() => cronDeleteTool.schema.parse({ id: "" })).toThrow();
});
it("requires a positive integer delaySeconds on schedule_wakeup", () => {
expect(() => scheduleWakeupTool.schema.parse({ delaySeconds: 0, prompt: "x" })).toThrow();
expect(() => scheduleWakeupTool.schema.parse({ delaySeconds: 1.5, prompt: "x" })).toThrow();
});
});
+296
View File
@@ -0,0 +1,296 @@
import { readFileSync, writeFileSync, existsSync, mkdirSync } from "node:fs";
import path from "node:path";
/** A scheduled, recurring cron job (5-field cron expression in the user's LOCAL timezone, matching
* Claude Code). `recurring: false` is a one-shot that fires once then auto-deletes. `durable` jobs
* are persisted to disk so they survive a restart; session-only jobs die with the process. */
export interface CronJob {
id: string;
cron: string;
prompt: string;
recurring: boolean;
durable: boolean;
/** Epoch ms the job was created — used for the 7-day auto-expiry on recurring jobs. */
createdAt: number;
/** Epoch-minute of the most recent fire, so a job doesn't re-fire within the same minute. */
lastFiredMinute?: number;
/** Set true once the 7-day expiry has fired its final run, so the tick deletes it after enqueue. */
expired?: boolean;
}
/** A one-shot delayed prompt (the ScheduleWakeup primitive), used for self-paced loops. Fires once
* at fireAt (epoch ms) then is removed. Session-only — never persisted. */
export interface Wakeup {
id: string;
fireAt: number;
prompt: string;
}
/** Parsed 5-field cron spec. `domStar`/`dowStar` record whether the day-of-month / day-of-week fields
* were `*` — needed for Vixie-cron semantics (when both are restricted, fire on EITHER match). */
interface CronSpec {
minute: number[];
hour: number[];
dom: number[];
month: number[];
dow: number[];
domStar: boolean;
dowStar: boolean;
}
const FIELD_RANGES: Record<string, [number, number]> = {
minute: [0, 59],
hour: [0, 23],
dom: [1, 31],
month: [1, 12],
dow: [0, 6],
};
// Parses a single cron field into the sorted list of allowed values. Supports `*`, the `*/N` step
// form, `N`, `N-M`, `N-M/S`, and comma-lists of any of these. Throws on out-of-range or unparseable
// input. (Line comments, not JSDoc, because the `*/N` step syntax contains a `*/` that would close a
// block comment prematurely.)
function parseField(field: string, range: [number, number]): { values: number[]; isStar: boolean } {
const min = range[0];
const max = range[1];
const out = new Set<number>();
const isStar = field === "*" || field === "*/1";
for (const part of field.split(",")) {
const slashIdx = part.indexOf("/");
let rangePart = part;
let step = 1;
if (slashIdx !== -1) {
rangePart = part.slice(0, slashIdx);
step = Number(part.slice(slashIdx + 1));
}
let lo: number;
let hi: number;
if (rangePart === "*") {
lo = min;
hi = max;
} else if (rangePart.includes("-")) {
const [a, b] = rangePart.split("-");
lo = Number(a);
hi = Number(b);
} else {
lo = hi = Number(rangePart);
}
if (!Number.isFinite(lo) || !Number.isFinite(hi) || !Number.isFinite(step) || step < 1) {
throw new Error(`invalid cron field "${field}"`);
}
if (lo < min || hi > max || lo > hi) {
throw new Error(`cron field "${field}" out of range [${min}-${max}]`);
}
for (let v = lo; v <= hi; v += step) out.add(v);
}
return { values: [...out].sort((a, b) => a - b), isStar };
}
/** Parses a 5-field cron expression into a matcher spec. Throws on malformed input. */
export function parseCron(expr: string): CronSpec {
const fields = expr.trim().split(/\s+/);
if (fields.length !== 5) throw new Error(`cron expression must have 5 fields, got ${fields.length}`);
const minute = fields[0]!;
const hour = fields[1]!;
const dom = fields[2]!;
const month = fields[3]!;
const dow = fields[4]!;
const m = parseField(minute, FIELD_RANGES.minute!);
const h = parseField(hour, FIELD_RANGES.hour!);
const dm = parseField(dom, FIELD_RANGES.dom!);
const mo = parseField(month, FIELD_RANGES.month!);
const dw = parseField(dow, FIELD_RANGES.dow!);
return { minute: m.values, hour: h.values, dom: dm.values, month: mo.values, dow: dw.values, domStar: dm.isStar, dowStar: dw.isStar };
}
/** Whether a cron spec matches a given local Date. */
function cronMatches(spec: CronSpec, d: Date): boolean {
if (!spec.minute.includes(d.getMinutes())) return false;
if (!spec.hour.includes(d.getHours())) return false;
if (!spec.month.includes(d.getMonth() + 1)) return false;
const domMatch = spec.dom.includes(d.getDate());
const dowMatch = spec.dow.includes(d.getDay());
// Vixie-cron: when both day fields are restricted, fire on EITHER; when one is *, the other governs.
if (spec.domStar && spec.dowStar) return true;
if (spec.domStar) return dowMatch;
if (spec.dowStar) return domMatch;
return domMatch || dowMatch;
}
const SEVEN_DAYS_MS = 7 * 24 * 60 * 60 * 1000;
const TICK_INTERVAL_MS = 15_000;
/** In-memory scheduler for cron jobs and one-shot wakeups. Held by the Session; the App starts it
* with an `enqueue` callback (submit a turn) and an `isIdle` predicate (true when no turn is
* running) so jobs only fire while the REPL is idle, matching Claude Code. Durable jobs persist
* to `<configDir>/scheduled_tasks.json`; session-only jobs don't. */
export class CronStore {
private jobs = new Map<string, CronJob>();
private wakeups = new Map<string, Wakeup>();
private seq = 0;
private wakeSeq = 0;
private timer: ReturnType<typeof setInterval> | undefined;
private enqueue: ((prompt: string) => void) | undefined;
private isIdle: (() => boolean) | undefined;
private readonly file: string | undefined;
constructor(configDir?: string) {
if (configDir) {
this.file = path.join(configDir, "scheduled_tasks.json");
this.loadDurable();
}
}
private nextId(): string {
this.seq += 1;
return `c${this.seq}`;
}
private nextWakeupId(): string {
this.wakeSeq += 1;
return `w${this.wakeSeq}`;
}
/** Begins the tick loop. Called once by the App after the session is wired. */
start(opts: { enqueue: (prompt: string) => void; isIdle: () => boolean }): void {
this.enqueue = opts.enqueue;
this.isIdle = opts.isIdle;
if (this.timer) return;
this.timer = setInterval(() => this.tick(), TICK_INTERVAL_MS);
// setInterval keeps the event loop alive; unref so the process can still exit naturally when the
// UI closes (the App stops the store on unmount anyway, but this is a backstop).
this.timer.unref?.();
}
/** Stops the tick loop (e.g. on session/App teardown). */
stop(): void {
if (this.timer) {
clearInterval(this.timer);
this.timer = undefined;
}
}
private tick(): void {
if (!this.enqueue || !this.isIdle) return;
if (!this.isIdle()) return; // only fire while the REPL is idle
const now = Date.now();
const nowMinute = Math.floor(now / 60_000);
const nowDate = new Date(now);
const toDelete: string[] = [];
for (const job of this.jobs.values()) {
const age = now - job.createdAt;
// 7-day auto-expiry: recurring jobs fire one final time then are deleted.
if (job.recurring && age >= SEVEN_DAYS_MS) {
this.enqueue(job.prompt);
toDelete.push(job.id);
continue;
}
if (job.lastFiredMinute === nowMinute) continue;
const spec = parseCron(job.cron);
if (cronMatches(spec, nowDate)) {
this.enqueue(job.prompt);
job.lastFiredMinute = nowMinute;
if (!job.recurring) toDelete.push(job.id); // one-shot fires once then is removed
}
}
for (const id of toDelete) this.delete(id);
// Wakeups: fire when their time has come.
const wakeDelete: string[] = [];
for (const w of this.wakeups.values()) {
if (now >= w.fireAt) {
this.enqueue(w.prompt);
wakeDelete.push(w.id);
}
}
for (const id of wakeDelete) this.wakeups.delete(id);
}
create(input: { cron: string; prompt: string; recurring?: boolean; durable?: boolean }): CronJob {
parseCron(input.cron); // validate syntax up front
const recurring = input.recurring ?? true;
const durable = input.durable ?? false;
const job: CronJob = {
id: this.nextId(),
cron: input.cron,
prompt: input.prompt,
recurring,
durable,
createdAt: Date.now(),
};
this.jobs.set(job.id, job);
if (durable) this.persist();
return job;
}
list(): CronJob[] {
return [...this.jobs.values()].map((j) => ({ ...j }));
}
get(id: string): CronJob | undefined {
const j = this.jobs.get(id);
return j ? { ...j } : undefined;
}
delete(id: string): boolean {
const existed = this.jobs.delete(id);
if (existed) this.persist();
return existed;
}
/** Schedules a one-shot wakeup `delaySeconds` from now. If `stop` is true, cancels ALL wakeups
* instead (used to end a self-paced loop). Returns the wakeup id, or `{ stopped: true }`. */
scheduleWakeup(input: { delaySeconds: number; prompt: string; stop?: boolean }):
| { id: string; fireAt: number }
| { stopped: true } {
if (input.stop) {
this.wakeups.clear();
return { stopped: true };
}
const delay = Math.max(60, Math.min(3600, input.delaySeconds));
const id = this.nextWakeupId();
const fireAt = Date.now() + delay * 1000;
this.wakeups.set(id, { id, fireAt, prompt: input.prompt });
return { id, fireAt };
}
listWakeups(): Wakeup[] {
return [...this.wakeups.values()].map((w) => ({ ...w }));
}
private loadDurable(): void {
if (!this.file || !existsSync(this.file)) return;
try {
const data = JSON.parse(readFileSync(this.file, "utf-8")) as { jobs?: CronJob[] };
for (const j of data.jobs ?? []) {
// Only durable jobs are persisted; skip any that slipped in without the flag.
if (!j.durable) continue;
this.jobs.set(j.id, { ...j });
// Bump the seq past any restored ids so new ids don't collide.
const n = Number(j.id.replace(/^c/, ""));
if (Number.isFinite(n) && n > this.seq) this.seq = n;
}
} catch {
// A corrupt persistence file shouldn't block startup — just start with no durable jobs.
}
}
private persist(): void {
if (!this.file) return;
try {
const dir = path.dirname(this.file);
if (!existsSync(dir)) mkdirSync(dir, { recursive: true });
const durable = [...this.jobs.values()].filter((j) => j.durable).map((j) => ({
id: j.id,
cron: j.cron,
prompt: j.prompt,
recurring: j.recurring,
durable: j.durable,
createdAt: j.createdAt,
}));
writeFileSync(this.file, JSON.stringify({ jobs: durable }, null, 2), "utf-8");
} catch {
// Persistence is best-effort; a write failure must not crash the scheduler.
}
}
}
+22 -69
View File
@@ -1,78 +1,31 @@
import { describe, it, expect } from "vitest";
import { describe, expect, it } from "vitest";
import { parseFallbackToolCalls } from "./fallbackParser.js";
describe("parseFallbackToolCalls", () => {
it("parses the instructed ```tool_call fenced format", () => {
const out = parseFallbackToolCalls(
'```tool_call\n{"name":"read_file","arguments":{"path":"src/a.ts"}}\n```',
it("parses a plain tool_call block", () => {
const result = parseFallbackToolCalls("```tool_call\n{\"name\": \"read_file\", \"arguments\": {\"path\": \"src/x.ts\"}}\n```");
expect(result.malformed).toBe(false);
expect(result.calls).toHaveLength(1);
expect(result.calls[0]).toEqual({ name: "read_file", arguments: { path: "src/x.ts" } });
});
it("strips inner ```json fences", () => {
const result = parseFallbackToolCalls(
"```tool_call\n```json\n{\"name\": \"read_file\", \"arguments\": {\"path\": \"src/x.ts\"}}\n```\n```",
);
expect(out.calls).toEqual([{ name: "read_file", arguments: { path: "src/a.ts" } }]);
expect(out.malformed).toBe(false);
expect(result.calls).toHaveLength(1);
expect(result.calls[0]).toEqual({ name: "read_file", arguments: { path: "src/x.ts" } });
});
it("returns no calls and not-malformed for plain prose with no tool attempt", () => {
const out = parseFallbackToolCalls("Here is the answer: use foo().");
expect(out.calls).toEqual([]);
expect(out.malformed).toBe(false);
it("marks top-level arguments (not nested in 'arguments') as malformed", () => {
const result = parseFallbackToolCalls("```tool_call\n{\"name\": \"read_file\", \"path\": \"src/x.ts\"}\n```");
expect(result.malformed).toBe(true);
expect(result.calls).toHaveLength(0);
});
it("flags a tool_call block whose JSON is unparseable as malformed", () => {
const out = parseFallbackToolCalls("```tool_call\n{not valid json\n```");
expect(out.calls).toEqual([]);
expect(out.malformed).toBe(true);
it("marks invalid JSON as malformed", () => {
const result = parseFallbackToolCalls("```tool_call\nnot json\n```");
expect(result.malformed).toBe(true);
expect(result.calls).toHaveLength(0);
});
it("accepts a ```json fenced block when no tool_call block is present", () => {
const out = parseFallbackToolCalls(
'```json\n{"name":"grep","arguments":{"pattern":"foo"}}\n```',
);
expect(out.calls).toEqual([{ name: "grep", arguments: { pattern: "foo" } }]);
expect(out.malformed).toBe(false);
});
it("prefers a ```tool_call block over a ```json block when both appear", () => {
const out = parseFallbackToolCalls(
'```tool_call\n{"name":"read_file","arguments":{"path":"a"}}\n```\n' +
'```json\n{"name":"grep","arguments":{"pattern":"x"}}\n```',
);
expect(out.calls).toHaveLength(1);
expect(out.calls[0]!.name).toBe("read_file");
});
it("does not treat a ```json fence inside a tool_call block as the terminator", () => {
const out = parseFallbackToolCalls(
"```tool_call\n```json\n{\"name\":\"read_file\",\"arguments\":{\"path\":\"a\"}}\n```\n```",
);
expect(out.calls).toEqual([{ name: "read_file", arguments: { path: "a" } }]);
});
it("extracts a bare (unfenced) tool-call object from surrounding prose", () => {
const out = parseFallbackToolCalls(
'Let me read that file.\n{"name":"read_file","arguments":{"path":"src/loop.ts"}}\nThat should help.',
);
expect(out.calls).toEqual([{ name: "read_file", arguments: { path: "src/loop.ts" } }]);
});
it("ignores bare braces in prose that don't look like a tool call", () => {
const out = parseFallbackToolCalls("The config is { key: value } and that's it.");
expect(out.calls).toEqual([]);
expect(out.malformed).toBe(false);
});
it("repairs a truncated JSON tool_call block via partial-JSON repair", () => {
// max_tokens clipped the closing brace and quote
const out = parseFallbackToolCalls('```tool_call\n{"name":"read_file","arguments":{"path":"src/lo');
expect(out.calls).toEqual([{ name: "read_file", arguments: { path: "src/lo" } }]);
});
it("accepts a tool call with omitted arguments as empty arguments", () => {
const out = parseFallbackToolCalls('```tool_call\n{"name":"git_status"}\n```');
expect(out.calls).toEqual([{ name: "git_status", arguments: {} }]);
});
it("flags a tool call whose arguments is not an object", () => {
const out = parseFallbackToolCalls('```tool_call\n{"name":"x","arguments":"foo"}\n```');
expect(out.calls).toEqual([]);
expect(out.malformed).toBe(true);
});
});
});
+18 -145
View File
@@ -1,5 +1,3 @@
import { repairPartialJson } from "./partialJson.js";
export interface FallbackToolCall {
name: string;
arguments: Record<string, unknown>;
@@ -10,155 +8,30 @@ export interface FallbackParseResult {
malformed: boolean;
}
// Local models in fallback mode (no native function calling) are asked to emit tool calls as a
// fenced ```tool_call block. In practice they frequently deviate, so the parser is lenient about
// FORMAT but strict about CONTENT: anything we extract must still parse to { name, arguments }.
// Accepted shapes, in priority order:
// 1. A ```tool_call fenced block (the instructed format). The opening fence is ```tool_call on
// its own line; the closing fence is ``` on its own line — anchored so a ```json block INSIDE
// isn't mistaken for the terminator. Some models wrap the JSON in an inner ```json fence; we
// strip that inner fence before parsing.
// 2. A ```json fenced block whose content is a tool-call object (name + arguments). Models that
// ignore the custom "tool_call" fence name but reach for the familiar "json" one.
// 3. A bare tool-call object appearing in the response with no fence at all. We scan for the
// first balanced {...} that contains a string "name" and an object "arguments". To avoid
// matching arbitrary prose-embedded JSON, we require the recognizable keys.
//
// Every extracted candidate goes through `coerce`, which parses (with partial-JSON repair as a
// last resort for truncated streaming output) and validates the name/arguments shape. A candidate
// that doesn't yield a valid call sets `malformed` — the loop nudges the model to retry rather than
// silently ending the task with no tool executed.
// Opening fence ```tool_call on its own line; closing ``` on its own line. Multiline-anchored so
// an inner ```json fence can't be read as the terminator.
const TOOL_CALL_FENCE_RE = /^```tool_call\s*\n([\s\S]*?)\n```(?:\n|$)/gm;
// A ```json block — only used if no ```tool_call block matched, since a json fence may carry prose.
const JSON_FENCE_RE = /^```json\s*\n([\s\S]*?)\n```(?:\n|$)/gm;
/** Pull the textual content out of a fenced block, stripping any inner ```json fence a model may
* have nested inside it. Returns the cleaned, trimmed body. */
function cleanFencedBody(raw: string): string {
return raw.replace(/^```(?:json)?\s*|\s*```$/g, "").trim();
}
/** Parse a candidate string into a tool call, or null if it isn't one. Tries a direct parse, then a
* partial-JSON repair (truncation / trailing comma / unbalanced braces) for clipped streaming. */
function coerce(candidate: string): FallbackToolCall | null {
const trimmed = candidate.trim();
if (!trimmed) return null;
// Direct parse first.
let parsed: unknown = null;
try {
parsed = JSON.parse(trimmed);
} catch {
parsed = null;
}
// Repair pass for truncated / sloppy JSON from local streaming.
if (parsed === null) parsed = repairPartialJson(trimmed);
if (!parsed || typeof parsed !== "object") return null;
const obj = parsed as Record<string, unknown>;
if (typeof obj.name !== "string") return null;
if (obj.arguments === undefined || obj.arguments === null) {
// Some models omit arguments entirely when the tool takes none — treat as empty.
return { name: obj.name, arguments: {} };
}
if (typeof obj.arguments !== "object" || Array.isArray(obj.arguments)) return null;
return { name: obj.name, arguments: obj.arguments as Record<string, unknown> };
}
/** Scan `content` for the first balanced {...} object containing a `"name"` string and an
* `"arguments"` object, with no fence at all. Tracks string literals so braces inside strings
* don't affect nesting, and cuts to the first complete top-level object. */
function findBareToolCall(content: string): string | null {
const start = content.indexOf("{");
if (start < 0) return null;
let depth = 0;
let inStr = false;
for (let i = start; i < content.length; i++) {
const c = content[i]!;
if (inStr) {
if (c === "\\") {
i++;
continue;
}
if (c === '"') inStr = false;
continue;
}
if (c === '"') {
inStr = true;
continue;
}
if (c === "{") depth++;
else if (c === "}") {
depth--;
if (depth === 0) {
const candidate = content.slice(start, i + 1);
// Only accept it if it actually looks like a tool call — otherwise keep scanning.
if (/"name"\s*:/.test(candidate) && /"arguments"\s*:/.test(candidate)) {
return candidate;
}
// Reset to the next brace past this point to keep looking.
const next = content.indexOf("{", i + 1);
if (next < 0) return null;
i = next - 1;
depth = 0;
}
}
}
// Ran off the end with an unclosed object — a truncated tool call (max_tokens clipped the
// closing brace, and possibly the closing fence too). Hand the unbalanced substring back;
// `coerce` will run it through partial-JSON repair to close what's open.
if (depth > 0) {
const candidate = content.slice(start);
if (/"name"\s*:/.test(candidate) && /"arguments"\s*:/.test(candidate)) {
return candidate;
}
}
return null;
}
// Opening fence is ```tool_call on its own line; closing fence is ``` on its own line.
// This prevents ```json inside the block from being mistaken for the terminator.
const BLOCK_RE = /^```tool_call\s*\n([\s\S]*?)\n```(?:\n|$)/gm;
export function parseFallbackToolCalls(content: string): FallbackParseResult {
const calls: FallbackToolCall[] = [];
let malformed = false;
let sawAnyCandidate = false;
// (1) ```tool_call fenced blocks (the instructed format) — there may be several.
for (const match of content.matchAll(TOOL_CALL_FENCE_RE)) {
sawAnyCandidate = true;
const body = cleanFencedBody(match[1] ?? "");
const call = coerce(body);
if (call) calls.push(call);
else malformed = true;
}
if (calls.length > 0) return { calls, malformed };
// (2) ```json fenced blocks — only if no tool_call block matched. Take the first that coerces.
for (const match of content.matchAll(JSON_FENCE_RE)) {
const call = coerce(cleanFencedBody(match[1] ?? ""));
if (call) {
calls.push(call);
return { calls, malformed: false };
for (const match of content.matchAll(BLOCK_RE)) {
const raw = match[1]?.trim() ?? "";
// Some local models emit markdown fences inside the tool_call block (e.g. ```json ... ```).
// Strip them so the inner JSON can be parsed.
const cleaned = raw.replace(/^```(?:json)?\s*|\s*```$/g, "").trim();
try {
const parsed = JSON.parse(cleaned || raw);
if (parsed && typeof parsed.name === "string" && typeof parsed.arguments === "object" && parsed.arguments !== null) {
calls.push({ name: parsed.name, arguments: parsed.arguments });
} else {
malformed = true;
}
} catch {
malformed = true;
}
sawAnyCandidate = true;
malformed = true;
}
if (calls.length > 0) return { calls, malformed };
// (3) Bare (unfenced) tool-call object — last resort. At most one: the prompt says one call per
// response, and extracting multiple bare objects from prose is too error-prone.
const bare = findBareToolCall(content);
if (bare !== null) {
sawAnyCandidate = true;
const call = coerce(bare);
if (call) {
calls.push(call);
return { calls, malformed: false };
}
malformed = true;
}
// If we never saw anything that even looked like a tool-call attempt, that's not malformed —
// the model simply answered in prose (no tool needed). `malformed` stays false.
void sawAnyCandidate;
return { calls, malformed };
}
}
+17 -7
View File
@@ -1,14 +1,24 @@
export const FALLBACK_TOOL_INSTRUCTIONS = `This model does not support native function calling. To call a tool, write a fenced code block:
export const FALLBACK_TOOL_INSTRUCTIONS = `This model does not support native function calling. To call a tool, write exactly one fenced code block of the form:
\`\`\`tool_call
{"name": "read_file", "arguments": {"path": "src/index.ts"}}
{"name": "TOOL_NAME", "arguments": {"arg1": "value1", "arg2": "value2"}}
\`\`\`
Rules:
- One tool call per response. Wait for the result before calling another.
- Must contain valid JSON with "name" and "arguments" keys.
- If no tool is needed, answer normally without a fenced block.
- One tool call per response. Wait for the \`\`\`tool_result\`\`\` before calling another.
- The fenced block must contain a single JSON object with exactly two keys: "name" and "arguments".
- "arguments" must be an object matching the tool's schema. Do not put the arguments at the top level.
- If no tool is needed, answer normally without any \`\`\`tool_call\`\`\` block.
The result is returned in a \`\`\`tool_result\`\`\` block. Then answer normally or call another tool.`;
Example:
\`\`\`tool_call
{"name": "read_file", "arguments": {"path": "src/index.ts", "limit": 50}}
\`\`\`
export const FALLBACK_RETRY_NUDGE = `Your last \`tool_call\` block wasn't valid JSON with "name" and "arguments" keys. Try again using the correct format, or answer without a tool call.`;
When a tool result shows an error or empty output, do not repeat the exact same call. Adjust your arguments or ask the user.`;
export const FALLBACK_RETRY_NUDGE = `Your last \`tool_call\` block was invalid. Check:
- It must be a single JSON object inside the fence, not plain text or multiple objects.
- It must have "name" (string) and "arguments" (object) keys.
- Argument values must match the tool's expected types.
Try again with the correct format, or answer without a tool call.`;
-45
View File
@@ -1,45 +0,0 @@
import { describe, expect, it } from "vitest";
import { z } from "zod";
import { resolveToolCall } from "./nativeAdapter.js";
import type { ToolDef } from "../tools/types.js";
function makeRegistry(): Map<string, ToolDef> {
const tool: ToolDef = {
name: "write_file",
description: "d",
schema: z.object({ path: z.string(), content: z.string() }),
mutating: true,
handler: async () => ({}),
};
return new Map([["write_file", tool]]);
}
function call(args: string) {
return {
id: "1",
type: "function" as const,
function: { name: "write_file", arguments: args },
};
}
describe("resolveToolCall — argument repair", () => {
it("resolves normally when arguments are already valid JSON", () => {
const r = resolveToolCall(call('{"path":"a.ts","content":"hi"}'), makeRegistry());
expect("tool" in r).toBe(true);
});
it("repairs arguments containing a raw (unescaped) newline instead of failing outright", () => {
// A local model echoing multi-line file content into a non-streaming completion often pastes
// a real newline byte into the JSON string rather than escaping it — this path (unlike the
// streaming path in loop.ts) previously had no repair attempt at all.
const args = '{"path":"a.ts","content":"line1\nline2"}';
const r = resolveToolCall(call(args), makeRegistry());
expect("tool" in r).toBe(true);
if ("tool" in r) expect((r.args as { content: string }).content).toBe("line1\nline2");
});
it("still errors when the arguments are unsalvageable", () => {
const r = resolveToolCall(call("not json at all"), makeRegistry());
expect(r).toEqual({ error: "arguments were not valid JSON" });
});
});
+1 -8
View File
@@ -1,7 +1,6 @@
import type { ChatCompletionMessageToolCall, ChatCompletionTool } from "openai/resources/chat/completions";
import { z } from "zod";
import type { ToolDef } from "../tools/types.js";
import { repairPartialJson } from "./partialJson.js";
import { resolveToolInvocation, type ResolvedToolCall } from "./resolve.js";
export function toOpenAITools(tools: ToolDef[]): ChatCompletionTool[] {
@@ -26,13 +25,7 @@ export function resolveToolCall(
try {
parsedArgs = JSON.parse(call.function.arguments || "{}");
} catch {
// Non-streaming completions land here directly (unlike the streaming path in loop.ts, which
// already repairs before this point) — without a repair attempt here too, a call whose
// arguments contain e.g. an unescaped literal newline (a local model echoing multi-line file
// content raw) fails outright instead of being salvaged.
const repaired = repairPartialJson(call.function.arguments || "");
if (repaired === null) return { error: "arguments were not valid JSON" };
parsedArgs = repaired;
return { error: "arguments were not valid JSON" };
}
return resolveToolInvocation(call.function.name, parsedArgs, registry);
}
-83
View File
@@ -1,83 +0,0 @@
import { describe, it, expect } from "vitest";
import { repairPartialJson } from "./partialJson.js";
describe("repairPartialJson", () => {
it("parses already-valid JSON unchanged", () => {
expect(repairPartialJson('{"path":"src/a.ts"}')).toEqual({ path: "src/a.ts" });
expect(repairPartialJson("[]")).toEqual([]);
expect(repairPartialJson(" 42 ")).toBe(42);
expect(repairPartialJson('{"a":1}\n')).toEqual({ a: 1 });
});
it("returns null for empty / whitespace input", () => {
expect(repairPartialJson("")).toBeNull();
expect(repairPartialJson(" ")).toBeNull();
});
it("strips a trailing comma before a closing bracket", () => {
expect(repairPartialJson('{"a":1,}')).toEqual({ a: 1 });
expect(repairPartialJson('[1,2,]')).toEqual([1, 2]);
expect(repairPartialJson('{"a":{"b":2,},}')).toEqual({ a: { b: 2 } });
});
it("strips stray trailing content after a complete value", () => {
expect(repairPartialJson('{"a":1}\n```')).toEqual({ a: 1 });
expect(repairPartialJson('{"a":1}garbage')).toEqual({ a: 1 });
expect(repairPartialJson('[1,2] }')).toEqual([1, 2]);
});
it("closes a truncated string value", () => {
// max_tokens clipped mid-value: {"path":"src/lo → needs closing quote + brace
expect(repairPartialJson('{"path":"src/lo')).toEqual({ path: "src/lo" });
expect(repairPartialJson('{"a":"hello wor')).toEqual({ a: "hello wor" });
});
it("balances unclosed braces and brackets from truncation", () => {
expect(repairPartialJson('{"a":1')).toEqual({ a: 1 });
expect(repairPartialJson('{"a":{"b":2')).toEqual({ a: { b: 2 } });
expect(repairPartialJson("[1,2")).toEqual([1, 2]);
expect(repairPartialJson('{"items":[1,2')).toEqual({ items: [1, 2] });
});
it("does not count braces inside string literals", () => {
// The braces/brackets inside the string are content, not nesting.
expect(repairPartialJson('{"code":"func() { return [1] "')).toEqual({
code: "func() { return [1] ",
});
expect(repairPartialJson('{"s":"\\\"escaped\\\""}')).toEqual({ s: '"escaped"' });
});
it("handles escaped quotes inside strings during truncation repair", () => {
// Unterminated string with an escaped quote inside: {"s":"a\"b
expect(repairPartialJson('{"s":"a\\"b')).toEqual({ s: 'a"b' });
});
it("combined: trailing comma exposed after balancing", () => {
// {"a":1,"b":2, (truncated with trailing comma) → close brace, then strip comma
expect(repairPartialJson('{"a":1,"b":2,')).toEqual({ a: 1, b: 2 });
});
it("returns null when input is not salvageable as object/array/scalar", () => {
expect(repairPartialJson("just prose with no json")).toBeNull();
expect(repairPartialJson("{:}")).toBeNull();
});
it("escapes a raw (unescaped) literal newline inside a string value", () => {
// A local model echoing multi-line file content often pastes real \n bytes into the JSON
// string instead of writing the two-char `\n` escape — JSON.parse rejects that outright.
expect(repairPartialJson('{"content":"line1\nline2"}')).toEqual({ content: "line1\nline2" });
});
it("escapes a raw carriage return inside a string value (CRLF source content)", () => {
expect(repairPartialJson('{"content":"line1\r\nline2"}')).toEqual({ content: "line1\r\nline2" });
});
it("leaves an already-escaped \\n sequence untouched", () => {
expect(repairPartialJson('{"content":"line1\\nline2"}')).toEqual({ content: "line1\nline2" });
});
it("combines raw-newline escaping with truncation repair", () => {
// Truncated mid-value AND containing a raw newline earlier in the string.
expect(repairPartialJson('{"content":"line1\nline2')).toEqual({ content: "line1\nline2" });
});
});
-252
View File
@@ -1,252 +0,0 @@
// Partial/truncated JSON repair for native streaming tool-call arguments.
//
// Local-model backends (Ollama, LM Studio) streaming tool calls accumulate the `arguments` string
// across deltas. Two common pathologies produce a string that JSON.parse rejects but that contains
// all the semantic content the model intended:
//
// 1. TRUNCATION — max_tokens clipped the JSON mid-value. The string ends inside a string value,
// an array, or an object: `{"path": "src/lo`, `{"items": [1, 2`, `{"a": {"b": 1`.
// 2. LOCAL-MODEL SLOPPINESS — a trailing comma, an unbalanced brace/bracket, or a trailing
// garbage token after the closing brace: `{"path": "x.ts",}`, `{"a": 1 `.
//
// This module attempts a cheap, conservative repair BEFORE the caller falls back to a full
// non-streaming regeneration (which is expensive on a local backend and often fails identically
// when the cause was max_tokens). It only closes what's open and trims what's stray — it never
// invents keys or values, so a genuinely malformed call still fails downstream at schema validation.
//
// The repair is best-effort: if it can't produce parseable JSON, it returns null and the caller
// keeps its existing retry path. It is deliberately string-based (no AST) so it's trivially fast and
// has no dependencies, and so it handles truncated input that a strict parser can't even build an
// AST from.
/** Parse `s` as JSON; on success return the value, on failure return null (never throws). */
function tryParse(s: string): unknown {
try {
return JSON.parse(s);
} catch {
return null;
}
}
/** JSON disallows raw control characters (0x00-0x1F — notably literal newline, CR, tab) inside
* string literals; they must be written as `\n`/`\r`/`\t`/`\u00XX`. Local models echoing
* multi-line file content (very common for write_file/edit_file on this project's CRLF-heavy
* source) routinely paste it in unescaped, which makes an otherwise complete, semantically
* correct tool call fail JSON.parse with "Bad control character in string literal". Escaping
* only touches raw bytes found *inside* a string (tracked the same way `skipString` does, so an
* existing `\\n` escape sequence is left alone) — it never changes where a string starts/ends or
* where a brace/bracket falls outside one, so it's safe to run before the other repair steps. */
function escapeRawControlCharsInStrings(s: string): string {
let out = "";
let inStr = false;
let changed = false;
for (let i = 0; i < s.length; i++) {
const c = s[i]!;
if (!inStr) {
if (c === '"') inStr = true;
out += c;
continue;
}
if (c === "\\") {
// Preserve an existing escape sequence verbatim — don't touch the char after the backslash.
out += c + (s[i + 1] ?? "");
i++;
continue;
}
if (c === '"') {
inStr = false;
out += c;
continue;
}
const code = c.charCodeAt(0);
if (code < 0x20) {
changed = true;
if (c === "\n") out += "\\n";
else if (c === "\r") out += "\\r";
else if (c === "\t") out += "\\t";
else out += "\\u" + code.toString(16).padStart(4, "0");
continue;
}
out += c;
}
return changed ? out : s;
}
/** Skip past the next JSON string literal starting at `i` (the opening quote). Returns the index
* just past the closing quote. Strings are the only place braces/brackets can appear without
* affecting nesting, so we must not count them while inside one. Handles `\"` and other escapes. */
function skipString(s: string, i: number): number {
let j = i + 1; // past opening quote
for (; j < s.length; j++) {
const c = s[j]!;
if (c === "\\") {
j++; // skip the escaped char (covers \", \\, etc.)
continue;
}
if (c === '"') return j + 1; // past closing quote
}
return j; // ran off the end — unterminated string
}
/** Attempts to repair `raw` into parseable JSON. Returns the parsed value on success, or null if no
* repair produced valid JSON. Steps, applied in order of how cheap and safe they are:
*
* 1. Maybe it already parses (trailing whitespace/newlines are fine for JSON.parse) — return as-is.
* 2. Escape raw control characters (literal newline/CR/tab) found inside string literals — a
* model echoing multi-line file content unescaped, common on this project's CRLF sources.
* 3. Strip a trailing comma before an expected-but-absent `}` or `]` (common local-model slip).
* 4. Strip stray non-JSON tokens after the first complete top-level value (`{"a":1}\n` → `{"a":1}`,
* and `{"a":1}garbage` → `{"a":1}` — JSON.parse rejects trailing content, so trim to the first
* complete value).
* 5. Close unterminated strings, then balance still-open braces/brackets (truncation repair).
*
* Each step re-attempts a parse, so the cheapest fix that works wins. */
export function repairPartialJson(raw: string): unknown | null {
if (!raw) return null;
const trimmed = raw.trim();
if (!trimmed) return null;
// (1) Already valid?
const direct = tryParse(trimmed);
if (direct !== null) return direct;
// (2) Raw control characters (literal newlines/CR/tab) inside string literals — see
// escapeRawControlCharsInStrings for why this is common. Escaping never moves a quote or
// brace, so every remaining step below runs against this version instead of the original.
const working = escapeRawControlCharsInStrings(trimmed);
if (working !== trimmed) {
const v = tryParse(working);
if (v !== null) return v;
}
// (3) Trailing comma before end-of-object/array: `{"a":1,}` or `[1,2,]`. Repeat until none
// left so a nested shape like `{"a":{"b":2,},}` clears both commas (innermost-first).
let noTrailing = working;
let prev: string;
do {
prev = noTrailing;
noTrailing = noTrailing.replace(/,\s*([\]}]+\s*$)/, "$1");
} while (noTrailing !== prev);
if (noTrailing !== working) {
const v = tryParse(noTrailing);
if (v !== null) return v;
}
// (4) Stray trailing content after the first complete value. JSON.parse refuses trailing tokens,
// but a model often emits a closing brace then a stray newline, a repeated token, or prose.
// Find the end of the first balanced top-level value and cut there.
const cut = cutToFirstCompleteValue(working);
if (cut !== null && cut !== working) {
const v = tryParse(cut);
if (v !== null) return v;
}
// (5) Truncation repair: close an unterminated string, then balance open braces/brackets.
const balanced = balanceAndClose(working);
if (balanced !== null && balanced !== working) {
// Re-run the earlier cheap fixes on the balanced result (a trailing comma may now be exposed).
const v = tryParse(balanced);
if (v !== null) return v;
const v2 = tryParse(balanced.replace(/,\s*([\]}]\s*$)/, "$1"));
if (v2 !== null) return v2;
}
return null;
}
/** If `s` starts with a complete top-level JSON value followed by stray content, return just that
* value (as a substring). Returns null if we can't find a clean boundary (e.g. the value is itself
* truncated). Walks the string tracking string literals and nesting depth. */
function cutToFirstCompleteValue(s: string): string | null {
let i = 0;
// Skip leading whitespace.
while (i < s.length && /\s/.test(s[i]!)) i++;
if (i >= s.length) return null;
const start = i;
const stack: string[] = [];
let inStr = false;
while (i < s.length) {
const c = s[i]!;
if (inStr) {
if (c === "\\") {
i += 2;
continue;
}
if (c === '"') inStr = false;
i++;
continue;
}
if (c === '"') {
inStr = true;
i++;
continue;
}
if (c === "{" || c === "[") {
stack.push(c);
i++;
continue;
}
if (c === "}" || c === "]") {
stack.pop();
i++;
// If the stack is empty, this was the end of the top-level value — cut here.
if (stack.length === 0) return s.slice(start, i);
continue;
}
// A bare scalar (number/true/false/null) ends at the next delimiter/comma/whitespace.
if (stack.length === 0 && (c === "," || c === "}" || c === "]" || /\s/.test(c))) {
return s.slice(start, i);
}
i++;
}
// Ran off the end without closing the top-level value → it's truncated, not "complete + stray".
if (stack.length > 0) return null;
return null;
}
/** Closes an unterminated trailing string and balances any open braces/brackets. Returns the
* repaired string, or null if nothing needed closing (caller can compare to skip a no-op parse). */
function balanceAndClose(s: string): string | null {
let out = s;
const stack: string[] = [];
let inStr = false;
let i = 0;
for (; i < out.length; i++) {
const c = out[i]!;
if (inStr) {
if (c === "\\") {
i++;
continue;
}
if (c === '"') inStr = false;
continue;
}
if (c === '"') {
inStr = true;
continue;
}
if (c === "{") stack.push("}");
else if (c === "[") stack.push("]");
else if (c === "}" || c === "]") stack.pop();
}
// If we ended inside a string, close it. A truncated value like `{"path":"src/lo` needs a closing
// quote before we can balance the outer braces.
if (inStr) {
out += '"';
}
// Close anything still open, innermost-first. Truncation mid-array/object → append the closers.
if (stack.length === 0 && !inStr) {
// Nothing to close — but a trailing comma may have been the only issue; let the caller handle it.
return inStr ? out : null;
}
while (stack.length) {
out += stack.pop();
}
return out;
}
+64 -48
View File
@@ -1,5 +1,6 @@
import { z } from "zod";
import type { SubAgentResult } from "./types.js";
import { AGENT_TYPE_NAMES, getAgentType } from "./agentTypes.js";
import type { SubAgentOverrides } from "./types.js";
import type { ToolDef } from "./types.js";
const schema = z.object({
@@ -8,64 +9,79 @@ const schema = z.object({
.string()
.describe(
"Full, self-contained task description. The sub-agent has no conversation memory and cannot ask follow-ups.",
)
.optional(),
tasks: z
.array(
z.object({
description: z.string().describe("Short label for this sub-task."),
prompt: z
.string()
.describe(
"Full, self-contained task description for this sub-task. The sub-agent has no conversation memory.",
),
}),
)
.min(2)
),
agentType: z
.enum(AGENT_TYPE_NAMES)
.optional()
.describe(
"Run several sub-agents IN PARALLEL (read-heavy research/audit tasks). Each gets its own " +
"isolated context and returns independently. Use this for many files, e.g. one sub-agent " +
"per directory or per concern (security, performance, tests). Prefer `prompt` for a single task.",
)
.optional(),
}).refine((v) => v.prompt || v.tasks, {
message: "Provide either `prompt` (single sub-agent) or `tasks` (parallel batch).",
"Specialist sub-agent type. 'general-purpose' (default) has full tool access and may edit files. " +
"'explore' is read-only research (locate code, map structure). 'code-reviewer' is read-only review " +
"(find bugs, verify claims, report findings). 'planner' is read-only planning (design an implementation " +
"plan with steps and tradeoffs). 'debugger' reproduces and fixes a bug with a minimal, verified fix. " +
"'test-writer' writes focused tests, runs them, and iterates until they pass. Read-only types " +
"(explore, code-reviewer, planner) cannot modify files.",
),
name: z
.string()
.optional()
.describe(
"Optional name for a shared-cwd (single, non-parallel) delegation. Naming it creates a named TEAMMATE: " +
"you can then continue it with send_message by name and see it via list_teammates, without tracking its " +
"agentId. Use this when you'll send a teammate several messages (e.g. a 'researcher' you'll re-query). " +
"Names must be unique within your session — re-using an existing name returns an error instead of " +
"clobbering the teammate. Parallel (worktree-isolated) delegations are fire-and-forget and ignore the name.",
),
});
export const agentTool: ToolDef<z.infer<typeof schema>> = {
name: "agent",
description:
"Delegate a task to a sub-agent with its own tool loop (no nested agents). Only the final answer is returned. " +
"For many files, split into multiple sub-agents — pass a `tasks` array to run several in PARALLEL " +
"(each returns independently; one failure doesn't discard the others). Sub-agents have a smaller " +
"step budget; if one runs out, narrow the task rather than retrying. Mutating tool calls from any " +
"sub-agent still go through the same permission prompts (serialized, so parallel sub-agents that " +
"edit don't race).",
"For many files, split into multiple sub-agents. Sub-agents have a smaller step budget; if one runs out, " +
"narrow the task rather than retrying. Pick an agentType: general-purpose (full access, may edit), explore " +
"(read-only research), code-reviewer (read-only review), planner (read-only design of an implementation plan), " +
"debugger (reproduce and fix a bug with a minimal verified fix), or test-writer (write and run focused tests). " +
"A single delegation runs in your working directory " +
"and persists its edits; when you emit several 'agent'/'agent__*' calls in one response they run in parallel, " +
"each in an isolated throwaway git worktree whose file changes are discarded — use parallel batches for " +
"research/review/analysis (the returned answer is the deliverable), and a single call for implementation. " +
"A single (non-parallel) delegation is resumable: it returns an agentId you can pass to send_message to " +
"continue it. Pass `name` to give it a stable teammate name you can address by name instead of the agentId.",
schema,
mutating: false,
handler: async (args, ctx) => {
const runSubAgent = ctx.runSubAgent;
if (!runSubAgent) {
if (!ctx.runSubAgent) {
throw new Error("Sub-agents are not available in this context.");
}
// Parallel batch: run each task as its own sub-agent, concurrently. A single failure surfaces
// as that task's `error` rather than rejecting the whole batch — the model gets every sibling's
// result and can retry just the one that failed instead of paying for all of them again.
if (args.tasks && args.tasks.length > 0) {
const results = await Promise.all(
args.tasks.map(async (t): Promise<SubAgentResult> => {
try {
const result = await runSubAgent({ description: t.description, prompt: t.prompt });
return { description: t.description, result };
} catch (err) {
return { description: t.description, error: (err as Error).message ?? String(err) };
}
}),
);
return { description: args.description, results };
// Pre-flight a name collision so we never clobber an existing teammate (and never waste a
// delegation whose name we'd refuse to register). If resolveTeammate is absent (a non-session
// context), skip the check — there's no roster to clobber and nothing to register later either.
if (args.name && ctx.resolveTeammate?.(args.name)) {
return {
error: `A teammate named "${args.name}" already exists. Use send_message with name "${args.name}" to continue it, or pick a different name for this new delegation.`,
};
}
// Single sub-agent (the original path).
const result = await runSubAgent({ description: args.description, prompt: args.prompt ?? "" });
return { description: args.description, result };
const spec = getAgentType(args.agentType);
// general-purpose inherits the full toolset and uses the generic prompt (no overrides). Specialist
// types restrict tools and add their identity as a prompt addendum on top of the generic prompt.
const overrides: SubAgentOverrides | undefined =
spec.name === "general-purpose" ? undefined : { toolNames: spec.toolNames, systemPromptAddendum: spec.systemPromptAddendum };
const result = await ctx.runSubAgent({ description: args.description, prompt: args.prompt }, overrides);
// Register the name on the session's roster only when the sub-agent is resumable (shared-cwd, not
// worktree-isolated) AND a roster is wired. Isolated parallel agents are fire-and-forget, so a
// name wouldn't be addressable — the result simply omits agentId and name, signalling
// non-resumability. We only echo `name` when it was actually registered, so the model never sees a
// name that send_message can't resolve (e.g. a non-session context with no roster).
let registered = false;
if (args.name && result.resumable && ctx.registerTeammate) {
ctx.registerTeammate(args.name, result.agentId);
registered = true;
}
return {
description: args.description,
result: result.result,
...(result.resumable ? { agentId: result.agentId } : {}),
...(registered ? { name: args.name } : {}),
};
},
};
};
+57
View File
@@ -0,0 +1,57 @@
import { describe, expect, it } from "vitest";
import { AGENT_TYPES, AGENT_TYPE_NAMES, getAgentType, isReadOnlyAgentType } from "./agentTypes.js";
describe("agentTypes registry", () => {
it("exposes the six built-in types", () => {
expect(AGENT_TYPE_NAMES).toEqual(["general-purpose", "explore", "code-reviewer", "planner", "debugger", "test-writer"]);
expect(AGENT_TYPES).toHaveLength(6);
});
it("every type has a non-empty addendum", () => {
for (const t of AGENT_TYPES) expect(t.systemPromptAddendum.length).toBeGreaterThan(0);
});
it("getAgentType resolves known names and falls back to general-purpose", () => {
expect(getAgentType("explore").name).toBe("explore");
expect(getAgentType("code-reviewer").name).toBe("code-reviewer");
expect(getAgentType("planner").name).toBe("planner");
expect(getAgentType("debugger").name).toBe("debugger");
expect(getAgentType("test-writer").name).toBe("test-writer");
expect(getAgentType("general-purpose").name).toBe("general-purpose");
expect(getAgentType(undefined).name).toBe("general-purpose");
expect(getAgentType("bogus").name).toBe("general-purpose");
});
it("read-only types (explore, code-reviewer, planner) vs mutating types (general-purpose, debugger, test-writer)", () => {
expect(isReadOnlyAgentType("explore")).toBe(true);
expect(isReadOnlyAgentType("code-reviewer")).toBe(true);
expect(isReadOnlyAgentType("planner")).toBe(true);
expect(isReadOnlyAgentType("general-purpose")).toBe(false);
expect(isReadOnlyAgentType("debugger")).toBe(false);
expect(isReadOnlyAgentType("test-writer")).toBe(false);
expect(isReadOnlyAgentType(undefined)).toBe(false);
});
it("read-only toolsets contain only known non-mutating tools", () => {
const explore = getAgentType("explore");
const reviewer = getAgentType("code-reviewer");
const planner = getAgentType("planner");
expect(explore.toolNames).not.toContain("write_file");
expect(explore.toolNames).not.toContain("bash");
expect(explore.toolNames).toContain("read_file");
expect(reviewer.toolNames).toContain("git_status");
expect(reviewer.toolNames).not.toContain("edit_file");
expect(planner.toolNames).toContain("git_status");
expect(planner.toolNames).not.toContain("edit_file");
expect(planner.toolNames).not.toContain("bash");
});
it("debugger and test-writer can mutate (reproduce/fix and write/run tests)", () => {
const debugger_ = getAgentType("debugger");
const testWriter = getAgentType("test-writer");
expect(debugger_.toolNames).toContain("bash");
expect(debugger_.toolNames).toContain("edit_file");
expect(testWriter.toolNames).toContain("write_file");
expect(testWriter.toolNames).toContain("bash");
});
});
+110
View File
@@ -0,0 +1,110 @@
import type { ToolDef } from "./types.js";
/** A built-in specialist sub-agent type the model can request via the `agent` tool's `agentType`
* field. Each restricts the sub-agent's toolset (read-only types can't mutate) and prepends a
* specialist addendum to the generic sub-agent system prompt. Plugin agents (agent__*) are a
* separate, full-replacement mechanism; these built-in types layer on top of the generic prompt so
* the standard tool discipline still applies. */
export interface AgentTypeSpec {
name: string;
/** One-line summary surfaced to the model via the schema enum description. */
description: string;
/** Tools the sub-agent may use (by ToolDef name). Omit to inherit the parent's full toolset minus
* further-nesting agent tools (same as a generic sub-agent). */
toolNames?: string[];
/** Prepended (as an addendum) to the generic sub-agent system prompt — NOT a full replacement, so
* the standard tool-use discipline survives. */
systemPromptAddendum: string;
}
export const AGENT_TYPES: AgentTypeSpec[] = [
{
name: "general-purpose",
description: "Full tool access (default). Use for implementation and any task that may edit files.",
// toolNames omitted → inherit the parent's full toolset.
systemPromptAddendum:
"You are a general-purpose sub-agent. You may read, search, and edit files to complete the delegated task.",
},
{
name: "explore",
description: "Read-only research: locate code and map structure across many files. Cannot modify anything.",
toolNames: ["read_file", "list_files", "grep", "web_search", "web_fetch"],
systemPromptAddendum:
"You are an Explore agent — a read-only research specialist. Your job is to locate code, map structure, and gather " +
"facts across the codebase to answer a specific question. You have only read/search tools and must not modify anything. " +
"Read excerpts rather than whole files; report conclusions with the file:line references that back them, not file dumps. " +
"If the answer isn't findable, say so plainly.",
},
{
name: "code-reviewer",
description: "Read-only code review: find bugs, verify claims against code, report ranked findings. Cannot modify.",
toolNames: ["read_file", "list_files", "grep", "git_status", "web_search", "web_fetch"],
systemPromptAddendum:
"You are a code-review specialist. Review the relevant code for correctness, edge cases, and likely bugs. You have only " +
"read/search tools. Verify every claim against the actual code rather than assuming. Report concrete findings with " +
"file:line anchors, ranked most-severe first; if you find nothing wrong, say so rather than inventing issues. Do not " +
"modify code — report only.",
},
{
name: "planner",
description: "Read-only planning: design an implementation plan with files to change, steps, and tradeoffs. Cannot modify.",
toolNames: ["read_file", "list_files", "grep", "git_status", "web_search", "web_fetch"],
systemPromptAddendum:
"You are a planning specialist. Investigate the codebase enough to design a concrete implementation plan — which files to " +
"change, in what order, and how, with the key code anchors (file:line) that justify each step. Surface tradeoffs and " +
"risks between approaches, and call out anything you'd need to verify before implementing. You have only read/search " +
"tools and must not modify anything. Return a step-by-step plan, not code dumps.",
},
{
name: "debugger",
description: "Reproduce and fix a bug: form hypotheses, read code, run commands to reproduce, apply a minimal fix, verify.",
toolNames: ["read_file", "list_files", "grep", "bash", "bash_output", "edit_file", "multi_edit"],
systemPromptAddendum:
"You are a debugging specialist. Investigate a reported bug by forming a hypothesis, reading the relevant code, and " +
"reproducing it with shell commands before touching anything. Apply the minimal fix that addresses the root cause (not " +
"the symptom), then verify the fix actually resolves the reproduction. Prefer a small, targeted edit over a rewrite. " +
"If you can't reproduce the bug, say so and report what you found instead of guessing at a fix.",
},
{
name: "test-writer",
description: "Write focused tests for a feature or bug fix, run them, and iterate until they pass.",
toolNames: ["read_file", "list_files", "grep", "write_file", "edit_file", "bash", "bash_output"],
systemPromptAddendum:
"You are a test-writing specialist. Write focused, meaningful tests (not trivial smoke tests) for the delegated feature " +
"or fix, following the project's existing test conventions and runner. Run the tests with shell commands and iterate " +
"until they pass — a test that's never run is unfinished. Cover the important edge cases, but don't over-test. If the " +
"code under test is wrong, fix it minimally rather than writing a test around the bug.",
},
];
export const AGENT_TYPE_NAMES = AGENT_TYPES.map((t) => t.name) as [string, ...string[]];
/** Looks up a built-in agent type by name. Falls back to general-purpose for an unknown/missing
* name so a model that omits the field or typo's it still gets a working sub-agent. */
export function getAgentType(name?: string): AgentTypeSpec {
if (name) {
const found = AGENT_TYPES.find((t) => t.name === name);
if (found) return found;
}
return AGENT_TYPES[0]!;
}
/** Whether a given agent type is read-only (no mutating tools), used to decide worktree isolation:
* read-only parallel agents don't need isolation and should see the current working state. */
export function isReadOnlyAgentType(name?: string): boolean {
const spec = getAgentType(name);
return spec.toolNames !== undefined && !spec.toolNames.some((n) => MUTATING_TOOL_NAMES.has(n));
}
/** Mutating tool names, for isReadOnlyAgentType. Kept here (not imported from the tool defs) so this
* stays a static decision without instantiating tools. */
const MUTATING_TOOL_NAMES = new Set([
"write_file",
"edit_file",
"multi_edit",
"notebook_edit",
"bash",
"bash_kill",
"git_commit",
"memory_write",
]);
+87
View File
@@ -0,0 +1,87 @@
import { describe, expect, it, vi } from "vitest";
import { askQuestionTool } from "./askQuestion.js";
import type { AskQuestionAnswer, AskQuestionSpec, ToolContext } from "./types.js";
const validArgs = {
questions: [
{
question: "Which auth method?",
header: "Auth method",
options: [
{ label: "OAuth", description: "delegate to provider" },
{ label: "API key", description: "simple header token" },
],
},
],
};
function ctxWith(askQuestion?: ToolContext["askQuestion"]): ToolContext {
return { cwd: "/x", ...(askQuestion ? { askQuestion } : {}) };
}
describe("ask_user_question tool", () => {
it("parses a well-formed single question", () => {
expect(() => askQuestionTool.schema.parse(validArgs)).not.toThrow();
});
it("rejects fewer than 2 options per question", () => {
expect(() =>
askQuestionTool.schema.parse({ questions: [{ question: "q?", header: "h", options: [{ label: "only" }] }] }),
).toThrow();
});
it("rejects more than 4 options per question", () => {
const opts = Array.from({ length: 5 }, (_, i) => ({ label: `o${i}` }));
expect(() => askQuestionTool.schema.parse({ questions: [{ question: "q?", header: "h", options: opts }] })).toThrow();
});
it("rejects zero questions", () => {
expect(() => askQuestionTool.schema.parse({ questions: [] })).toThrow();
});
it("rejects more than 4 questions", () => {
const qs = Array.from({ length: 5 }, (_, i) => ({
question: `q${i}?`,
header: `h${i}`,
options: [{ label: "a" }, { label: "b" }],
}));
expect(() => askQuestionTool.schema.parse({ questions: qs })).toThrow();
});
it("rejects a header longer than 12 chars", () => {
expect(() =>
askQuestionTool.schema.parse({
questions: [{ question: "q?", header: "this is too long", options: [{ label: "a" }, { label: "b" }] }],
}),
).toThrow();
});
it("returns the user's answers when askQuestion is wired", async () => {
const answers: AskQuestionAnswer[] = [{ question: "Which auth method?", selected: ["OAuth"] }];
const askQuestion = vi.fn(async (_questions: AskQuestionSpec[]) => answers);
const result = await askQuestionTool.handler(validArgs as any, ctxWith(askQuestion));
expect(askQuestion).toHaveBeenCalledTimes(1);
// The spec passed to the callback preserves question/header/options and omits undefined fields.
const first = askQuestion.mock.calls[0]![0][0]!;
expect(first.question).toBe("Which auth method?");
expect(first.header).toBe("Auth method");
expect(first.options[0]).toEqual({ label: "OAuth", description: "delegate to provider" });
expect("multiSelect" in first).toBe(false);
expect(result).toEqual({ answers });
});
it("preserves multiSelect when set", async () => {
const askQuestion = vi.fn(async (_questions: AskQuestionSpec[]) => [{ question: "q?", selected: ["a", "b"] }]);
await askQuestionTool.handler(
{ questions: [{ question: "q?", header: "h", options: [{ label: "a" }, { label: "b" }], multiSelect: true }] } as any,
ctxWith(askQuestion),
);
expect(askQuestion.mock.calls[0]![0][0]!.multiSelect).toBe(true);
});
it("returns a clear error (not a hang) when no interactive UI is available", async () => {
const result = await askQuestionTool.handler(validArgs as any, ctxWith(undefined));
expect("error" in (result as object)).toBe(true);
expect((result as { error: string }).error).toMatch(/no interactive UI|Can't ask/i);
});
});
+68
View File
@@ -0,0 +1,68 @@
import { z } from "zod";
import type { AskQuestionAnswer, AskQuestionSpec, ToolDef } from "./types.js";
const optionSchema = z.object({
label: z.string().describe("A concise (1-5 word) label for the option."),
description: z
.string()
.optional()
.describe("Explanation of what this option means or the trade-off it implies, shown dimmed under the label."),
});
const questionSchema = z.object({
question: z.string().describe("The complete question to ask, ending with a question mark."),
header: z
.string()
.max(12)
.describe("A very short label (max ~12 chars) shown as a chip beside the question, e.g. \"Auth method\"."),
options: z
.array(optionSchema)
.min(2)
.max(4)
.describe("Two to four mutually exclusive options (unless multiSelect). The user can also type a custom \"Other\" answer."),
multiSelect: z
.boolean()
.optional()
.describe("Set true to allow several options to be selected instead of just one."),
});
const schema = z.object({
questions: z
.array(questionSchema)
.min(1)
.max(4)
.describe("One to four questions to ask. The UI asks them one at a time and returns all answers together."),
});
/** Lets the model ask the user a structured multiple-choice question when it is blocked on a decision
* that is genuinely the user's to make — one it can't resolve from the code, request, or sensible
* defaults. Non-mutating (it changes nothing on the filesystem), so it's allowed in every permission
* mode including plan mode. The UI shows each question with its options (plus an implicit "Other"
* path for a freeform answer) and returns the selected label(s); in a headless context with no UI
* (sub-agents) the callback is absent and the tool fails with a clear "can't ask" error instead of
* hanging. Reserve this for real decision points — don't ask questions you could answer yourself by
* reading the code or following an obvious default. */
export const askQuestionTool: ToolDef<z.infer<typeof schema>> = {
name: "ask_user_question",
description:
"Ask the user a structured multiple-choice question when blocked on a decision only they can make. " +
"Pass 1-4 questions, each with 2-4 options and a short header chip. The user can pick an option or type a " +
"custom \"Other\" answer. Use this instead of a prose question when a discrete choice would clarify the path. " +
"Only ask when the request is genuinely ambiguous after you've explored — don't offload decisions you could " +
"make yourself.",
schema,
mutating: false,
handler: async (args, ctx) => {
if (!ctx.askQuestion) {
return { error: "Can't ask the user a question in this context (no interactive UI). Make a sensible default choice and proceed, or explain the trade-off in prose." };
}
const specs: AskQuestionSpec[] = args.questions.map((q) => ({
question: q.question,
header: q.header,
options: q.options.map((o) => ({ label: o.label, ...(o.description ? { description: o.description } : {}) })),
...(q.multiSelect ? { multiSelect: true } : {}),
}));
const answers: AskQuestionAnswer[] = await ctx.askQuestion(specs);
return { answers };
},
};
-33
View File
@@ -1,5 +1,4 @@
import { beforeEach, describe, expect, it, vi } from "vitest";
import { execa } from "execa";
import type { ResultPromise } from "execa";
import { bashTool } from "./bash.js";
import { killProcessTree } from "../utils/processTree.js";
@@ -81,36 +80,4 @@ describe("bash tool — sub-agent abort", () => {
expect(killProcessTree).not.toHaveBeenCalled();
});
});
describe("bash tool — safety guards", () => {
beforeEach(() => {
vi.mocked(execa).mockClear();
});
it("refuses a command that wipes the filesystem root before ever spawning it", async () => {
const ctx: ToolContext = { cwd: process.cwd() };
await expect(bashTool.handler({ command: "rm -rf /" }, ctx)).rejects.toThrow(/Refusing to run/);
expect(execa).not.toHaveBeenCalled();
});
it("does not flag an ordinary rm -rf on a project subdirectory", async () => {
// preview (not handler) so this doesn't spawn the fake child, which only ever resolves when
// killed/backgrounded — nothing here would do either, so awaiting handler() would hang.
const ctx: ToolContext = { cwd: process.cwd() };
const preview = await bashTool.preview!({ command: "rm -rf node_modules" }, ctx);
expect(preview).not.toMatch(/^Blocked:/);
});
it("refuses a cwd override that escapes the working directory", async () => {
const ctx: ToolContext = { cwd: process.cwd() };
await expect(bashTool.handler({ command: "ls", cwd: "../../" }, ctx)).rejects.toThrow(/Refusing to write outside/);
expect(execa).not.toHaveBeenCalled();
});
it("preview surfaces the block reason instead of running the command", async () => {
const ctx: ToolContext = { cwd: process.cwd() };
const preview = await bashTool.preview!({ command: "mkfs.ext4 /dev/sda1" }, ctx);
expect(preview).toMatch(/^Blocked:/);
});
});
+24 -33
View File
@@ -1,11 +1,11 @@
import path from "node:path";
import { execa } from "execa";
import { z } from "zod";
import { registerBackgroundJob } from "./backgroundJobs.js";
import { riskyBashCommandReason } from "./bashGuard.js";
import { resolveWithinCwd } from "./pathGuard.js";
import { killProcessTree } from "../utils/processTree.js";
import { truncate } from "../utils/truncate.js";
import { resolveShell } from "../utils/shell.js";
import { assertWithinWorkspace } from "../utils/path.js";
import type { ToolDef } from "./types.js";
const schema = z.object({
@@ -22,39 +22,28 @@ function delay(ms: number): Promise<"pending"> {
export const bashTool: ToolDef<z.infer<typeof schema>> = {
name: "bash",
description: "Run a shell command and return its stdout, stderr, and exit code. Use for building, running tests, git operations, or inspecting the environment. Output is capped (head+tail preserved); long-running commands can be backgrounded with Ctrl+B and checked with bash_output.",
description:
"Run a shell command and return its stdout, stderr, and exit code. Be careful with destructive " +
"operations (rm, git push, etc.). For long-running commands, increase timeout_ms or use Ctrl+B " +
"to background the command while it's running.",
schema,
mutating: true,
preview: async ({ command, cwd }, ctx) => {
const riskyReason = riskyBashCommandReason(command);
if (riskyReason) return `Blocked: this command ${riskyReason}.`;
if (cwd) {
try {
resolveWithinCwd(ctx.cwd, cwd);
} catch (err) {
return (err as Error).message;
}
}
return `Run shell command: ${command}${cwd ? ` (cwd: ${cwd})` : ""}`;
},
preview: async ({ command, cwd }) => `Run shell command: ${command}${cwd ? ` (cwd: ${cwd})` : ""}`,
handler: async ({ command, cwd, timeout_ms }, ctx) => {
const riskyReason = riskyBashCommandReason(command);
if (riskyReason) {
throw new Error(`Refusing to run: this command ${riskyReason}.`);
}
const workDir = cwd ? resolveWithinCwd(ctx.cwd, cwd) : ctx.cwd;
const workDir = cwd ? path.resolve(ctx.cwd, cwd) : ctx.cwd;
if (cwd) assertWithinWorkspace(workDir, ctx.cwd, cwd);
// Timeout is enforced by our own timer rather than execa's built-in `timeout` option, so that
// backgrounding via Ctrl+B can cancel it below — execa's own timeout kills the process on a
// fixed schedule regardless of what happens to it afterward, which would silently kill a
// long-running command right after the user chose to keep it running in the background.
const child = execa(command, { shell: resolveShell(), cwd: workDir, reject: false });
const stdoutChunks: string[] = [];
const stderrChunks: string[] = [];
let stdout = "";
let stderr = "";
const onStdout = (d: Buffer) => {
stdoutChunks.push(d.toString());
stdout += d.toString();
};
const onStderr = (d: Buffer) => {
stderrChunks.push(d.toString());
stderr += d.toString();
};
child.stdout?.on("data", onStdout);
child.stderr?.on("data", onStderr);
@@ -88,13 +77,10 @@ export const bashTool: ToolDef<z.infer<typeof schema>> = {
for (;;) {
if (ctx.backgroundControl?.requested) {
clearTimeout(foregroundTimer);
// Detach our own capture listeners before handing the streams to the background registry —
// otherwise both this closure's listeners and registerBackgroundJob's keep appending to
// separate buffers forever, doubling the work and growing memory without bound for a
// long-running backgrounded job. The buffers captured so far seed the job.
child.stdout?.off("data", onStdout);
child.stderr?.off("data", onStderr);
const job = registerBackgroundJob(command, workDir, child, stdoutChunks.join(""), stderrChunks.join(""));
// Hand the streams off to the background registry; the finally block below will detach
// our own capture listeners so both closures don't keep appending to separate buffers
// forever, doubling the work and growing memory without bound for a long-running job.
const job = registerBackgroundJob(command, workDir, child, stdout, stderr);
return {
backgrounded: true,
jobId: job.id,
@@ -106,13 +92,18 @@ export const bashTool: ToolDef<z.infer<typeof schema>> = {
clearTimeout(foregroundTimer);
return {
exitCode: settled.exitCode,
stdout: truncate(stdoutChunks.join("")),
stderr: truncate(stderrChunks.join("")),
stdout: truncate(stdout),
stderr: truncate(stderr),
timedOut,
};
}
}
} finally {
// Detach our capture listeners on every exit path so a long-running command doesn't keep
// orphaned handlers alive after the tool returns. On backgrounding this also stops the
// foreground closure from competing with the background registry for stream data.
child.stdout?.off("data", onStdout);
child.stderr?.off("data", onStderr);
// Remove the abort listener on every exit path. On backgrounding this is what stops a
// sub-agent timeout from killing a job the user explicitly chose to keep running; on normal
// completion it's just cleanup. (The listener is `{ once: true }`, but it may never fire.)
-45
View File
@@ -1,45 +0,0 @@
import { describe, expect, it } from "vitest";
import { riskyBashCommandReason } from "./bashGuard.js";
describe("riskyBashCommandReason", () => {
describe("blocks", () => {
const cases: [name: string, command: string][] = [
["rm -rf /", "rm -rf /"],
["rm -fr / (flag order swapped)", "rm -fr /"],
["rm -rf / with a trailing slash-star", "rm -rf /*"],
["rm -Rf ~ (home dir)", "rm -Rf ~"],
["rm --recursive --force /", "rm --recursive --force /"],
["sudo rm -rf /", "sudo rm -rf /"],
["classic fork bomb", ":(){ :|:& };:"],
["fork bomb with extra whitespace", ": ( ) { : | : & } ; :"],
["mkfs.ext4", "mkfs.ext4 /dev/sda1"],
["dd to a raw device", "dd if=/dev/zero of=/dev/sda bs=1M"],
["redirect onto a raw device", "echo oops > /dev/sda"],
["Windows format", "format C:"],
["Windows rd /s /q on a drive root", "rd /s /q C:\\"],
["PowerShell Remove-Item -Recurse -Force on a drive", "Remove-Item -Recurse -Force C:\\"],
];
for (const [name, command] of cases) {
it(name, () => {
expect(riskyBashCommandReason(command)).not.toBeNull();
});
}
});
describe("does not block", () => {
const cases: [name: string, command: string][] = [
["rm -rf on a project subdirectory", "rm -rf node_modules"],
["rm -rf on a relative build dir", "rm -rf ./dist"],
["rm without force/recursive on root-looking arg", "rm /tmp/foo.txt"],
["dd between two regular files", "dd if=file.img of=out.img"],
["a command that merely contains the word format", "echo 'please format your PR title'"],
["a normal git command", "git status"],
["listing a directory named format", "ls format"],
];
for (const [name, command] of cases) {
it(name, () => {
expect(riskyBashCommandReason(command)).toBeNull();
});
}
});
});
-87
View File
@@ -1,87 +0,0 @@
/** Blocks a small set of unambiguously catastrophic shell commands — wiping the whole filesystem
* or a whole drive, formatting a device, a fork bomb — before they ever reach the confirmation
* prompt (or, under auto-accept, before they'd run with no prompt at all). This is not a general
* command sandbox: it doesn't stop a model from `rm -rf`-ing some *other* directory it shouldn't,
* running a slow fork loop that isn't the canonical bomb syntax, or anything merely inadvisable —
* only the handful of patterns whose only realistic purpose is destroying the whole machine, where
* a false negative is far more likely than a false positive. Deliberately narrow so it doesn't
* reject legitimate commands like `rm -rf node_modules` or `dd if=file.img of=out.img`. */
interface RiskyPattern {
test: (command: string) => boolean;
reason: string;
}
/** Splits on whitespace for a crude token scan — good enough for a blocklist (not a security
* boundary; execa still runs the raw string through a real shell either way) and avoids a brittle
* do-everything regex that has to encode flag ordering itself. */
function tokenize(command: string): string[] {
return command.trim().split(/\s+/);
}
const ROOT_TARGETS = new Set(["/", "/*", "~", "~/", "~/*", "$home", "${home}"]);
/** `rm -rf /`, `rm -fr ~`, `sudo rm -Rf --no-preserve-root /`, etc. — recursive+forced deletion
* whose target is the filesystem root or the whole home directory, in any flag order/spelling. */
function isRmWipingRootOrHome(command: string): boolean {
const tokens = tokenize(command).map((t) => t.toLowerCase());
const rmIdx = tokens.findIndex((t) => t === "rm" || t.endsWith("/rm"));
if (rmIdx === -1) return false;
const rest = tokens.slice(rmIdx + 1);
const isFlag = (t: string) => t.startsWith("-");
const hasForce = rest.some((t) => (isFlag(t) && !t.startsWith("--") && t.includes("f")) || t === "--force");
const hasRecursive = rest.some((t) => (isFlag(t) && !t.startsWith("--") && (t.includes("r") || t.includes("R"))) || t === "--recursive");
const targets = rest.filter((t) => !isFlag(t));
return hasForce && hasRecursive && targets.some((t) => ROOT_TARGETS.has(t));
}
const RISKY_PATTERNS: RiskyPattern[] = [
{
test: isRmWipingRootOrHome,
reason: "recursively force-deletes the filesystem root or home directory",
},
{
// Classic bash fork bomb: ":(){ :|:& };:" (whitespace-tolerant).
test: (cmd) => /:\s*\(\s*\)\s*\{\s*:\s*\|\s*:\s*&?\s*;?\s*\}\s*;\s*:/.test(cmd),
reason: "is a fork bomb (unbounded process spawning)",
},
{
test: (cmd) => /\bmkfs(\.\w+)?\b/i.test(cmd),
reason: "formats a filesystem (mkfs)",
},
{
test: (cmd) => /\bdd\b[^\n]*\bof=\/dev\/(sd|hd|nvme|disk|xvd|rdisk)\w*/i.test(cmd),
reason: "writes raw data directly to a block device (dd of=/dev/...)",
},
{
test: (cmd) => />\s*\/dev\/(sd|hd|nvme|disk|xvd|rdisk)\w*\b/i.test(cmd),
reason: "redirects output directly onto a block device",
},
{
// `format C:`, `format /Y D:` — Windows drive format.
test: (cmd) => /\bformat\b[^\n]*\b[a-zA-Z]:/i.test(cmd),
reason: "formats a Windows drive (format)",
},
{
// `rd /s /q C:\`, `rmdir /s /q D:\` — recursive quiet delete of a bare drive root.
test: (cmd) => /\b(rd|rmdir)\b[^\n]*\/s\b[^\n]*\b[a-zA-Z]:\\?\s*(\/q\b[^\n]*)?$/im.test(cmd),
reason: "recursively deletes an entire Windows drive",
},
{
// PowerShell `Remove-Item -Recurse -Force C:\` (or -Path C:\, or $env:SystemDrive), flag order-tolerant.
test: (cmd) =>
/remove-item\b/i.test(cmd) &&
/-recurse\b/i.test(cmd) &&
/-force\b/i.test(cmd) &&
(/\b[a-zA-Z]:\\?\s*($|['")\s;])/.test(cmd) || /\$env:systemdrive\b/i.test(cmd)),
reason: "recursively force-deletes an entire Windows drive (Remove-Item)",
},
];
/** Returns a human-readable reason if `command` matches a known catastrophic pattern, else null. */
export function riskyBashCommandReason(command: string): string | null {
for (const pattern of RISKY_PATTERNS) {
if (pattern.test(command)) return pattern.reason;
}
return null;
}
-68
View File
@@ -1,68 +0,0 @@
import { describe, it, expect, vi, beforeEach } from "vitest";
// The LSP tools (definition/references/diagnostics) are thin dispatchers over lspManager. We mock
// the manager functions so the tests run without spawning a language server, and assert each tool
// forwards the right arguments (path resolved against cwd, 1-indexed→handled by the manager,
// includeDeclaration default) and returns the manager's result verbatim.
vi.mock("../codeintel/lspManager.js", () => ({
getDefinition: vi.fn(async () => ({ definitions: [{ path: "/abs/a.ts", line: 3, column: 5 }] })),
getReferences: vi.fn(async () => ({ references: [{ path: "/abs/a.ts", line: 3, column: 5 }] })),
getDiagnostics: vi.fn(async () => ({
diagnostics: [{ path: "/abs/a.ts", line: 1, column: 1, severity: "error", message: "oops" }],
})),
}));
import { definitionTool, referencesTool, diagnosticsTool } from "./codeIntel.js";
import { getDefinition, getReferences, getDiagnostics } from "../codeintel/lspManager.js";
const ctx = { cwd: "/proj" };
describe("definition tool", () => {
beforeEach(() => vi.clearAllMocks());
it("forwards path/line/column/cwd to getDefinition and returns its result", async () => {
const out = await definitionTool.handler({ path: "src/a.ts", line: 3, column: 5 }, ctx);
expect(getDefinition).toHaveBeenCalledWith("src/a.ts", 3, 5, "/proj");
expect(out).toEqual({ definitions: [{ path: "/abs/a.ts", line: 3, column: 5 }] });
});
it("is read-only (no permission prompt)", () => {
expect(definitionTool.mutating).toBe(false);
});
});
describe("references tool", () => {
beforeEach(() => vi.clearAllMocks());
it("defaults includeDeclaration to true when omitted", async () => {
await referencesTool.handler({ path: "src/a.ts", line: 3, column: 5 }, ctx);
expect(getReferences).toHaveBeenCalledWith("src/a.ts", 3, 5, "/proj", true);
});
it("passes an explicit includeDeclaration through", async () => {
await referencesTool.handler({ path: "src/a.ts", line: 3, column: 5, include_declaration: false }, ctx);
expect(getReferences).toHaveBeenCalledWith("src/a.ts", 3, 5, "/proj", false);
});
it("returns the manager's references result", async () => {
const out = await referencesTool.handler({ path: "src/a.ts", line: 3, column: 5 }, ctx);
expect(out).toEqual({ references: [{ path: "/abs/a.ts", line: 3, column: 5 }] });
});
});
describe("diagnostics tool", () => {
beforeEach(() => vi.clearAllMocks());
it("forwards path/cwd to getDiagnostics and returns its result", async () => {
const out = await diagnosticsTool.handler({ path: "src/a.ts" }, ctx);
expect(getDiagnostics).toHaveBeenCalledWith("src/a.ts", "/proj");
expect(out).toEqual({
diagnostics: [{ path: "/abs/a.ts", line: 1, column: 1, severity: "error", message: "oops" }],
});
});
it("is read-only", () => {
expect(diagnosticsTool.mutating).toBe(false);
});
});
-61
View File
@@ -1,61 +0,0 @@
import { z } from "zod";
import type { ToolDef } from "./types.js";
import { getDefinition, getReferences, getDiagnostics } from "../codeintel/lspManager.js";
// All three tools are read-only LSP queries. They share a common shape: point them at a file
// (relative to cwd) and a 1-indexed line/column, and they ask the language server for the answer.
// The server is lazily started on first use per language (tsserver, pyright, gopls, clangd,
// rust-analyzer) and reused across the whole session — see lspManager.ts for the lifecycle.
//
// The error messages from lspManager are written to be actionable (e.g. "install
// typescript-language-server"), so we let them surface verbatim rather than wrapping them — a
// generic "LSP unavailable" would hide the one piece of info the model needs to recover.
const positionSchema = z.object({
path: z
.string()
.describe("File path, relative to the working directory. Must match an extension with a configured LSP server (.ts/.tsx/.js/.jsx/.py/.go/.rs/.c/.cpp/…)."),
line: z.number().int().min(1).describe("1-indexed line number of the symbol to query."),
column: z.number().int().min(1).describe("1-indexed column number of the symbol to query."),
});
export const definitionTool: ToolDef<z.infer<typeof positionSchema>> = {
name: "definition",
description:
"Resolve where a symbol is DEFINED using the language server (LSP). Use when grep finds a call site but you need the actual declaration — e.g. a function/variable/type name at a line:column. Returns one or more file:line:column locations (empty list if the server couldn't resolve it, which is a legitimate 'not found', not an error). Requires the relevant language server on PATH (typescript-language-server, pyright-langserver, gopls, clangd, or rust-analyzer).",
schema: positionSchema,
mutating: false,
handler: async (args, ctx) => getDefinition(args.path, args.line, args.column, ctx.cwd),
};
const referencesSchema = positionSchema.extend({
include_declaration: z
.boolean()
.optional()
.describe("Whether to include the symbol's own declaration among the references. Defaults to true (matches most IDE 'find all references' behavior)."),
});
export const referencesTool: ToolDef<z.infer<typeof referencesSchema>> = {
name: "references",
description:
"Find every reference to a symbol using the language server (LSP) — the same as an IDE's 'find all references'. Use to enumerate all call/usage sites of a function/variable/type at a line:column before a rename or to gauge impact. Returns a list of file:line:column locations. Requires the relevant language server on PATH.",
schema: referencesSchema,
mutating: false,
handler: async (args, ctx) =>
getReferences(args.path, args.line, args.column, ctx.cwd, args.include_declaration ?? true),
};
const diagnosticsSchema = z.object({
path: z
.string()
.describe("File path, relative to the working directory, to check for type/syntax errors."),
});
export const diagnosticsTool: ToolDef<z.infer<typeof diagnosticsSchema>> = {
name: "diagnostics",
description:
"Get the latest type/syntax diagnostics (errors and warnings) the language server has published for a file — equivalent to an editor's Problems panel. Use right after an edit_file/write_file to verify the change didn't introduce a type error, or when `tsc --noEmit`/`pyright` would be the alternative. Forces a document sync first so the snapshot is current. Returns severity (error/warning/information/hint), line, column, and message for each diagnostic. Requires the relevant language server on PATH.",
schema: diagnosticsSchema,
mutating: false,
handler: async (args, ctx) => getDiagnostics(args.path, ctx.cwd),
};
+95
View File
@@ -0,0 +1,95 @@
import { z } from "zod";
import type { ToolDef } from "./types.js";
export const cronCreateTool: ToolDef<z.infer<typeof cronCreateSchema>> = {
name: "cron_create",
description:
"Schedule a prompt to run on a recurring cron schedule (5-field cron in the user's LOCAL timezone: minute hour " +
"day-of-month month day-of-week, e.g. '0 9 * * 1-5' = weekdays at 9am). Use for recurring checks, reminders, or " +
"self-paced loops. The prompt fires only while the REPL is idle. Recurring jobs auto-expire after 7 days. Set " +
"recurring: false for a one-shot that fires once then deletes itself. Set durable: true to persist across restarts. " +
"Returns the new job id.",
schema: z.object({
cron: z
.string()
.min(1)
.describe("5-field cron expression (minute hour day-of-month month day-of-week) in local time."),
prompt: z.string().min(1).describe("The prompt to enqueue when the job fires."),
recurring: z.boolean().optional().describe("True (default) to fire on every match; false to fire once then delete."),
durable: z
.boolean()
.optional()
.describe("True to persist the job to disk so it survives a restart (default false = session-only)."),
}),
// Scheduling is reversible (cron_delete) and not a destructive filesystem op — no confirmation prompt.
mutating: false,
handler: async (args, ctx) => {
if (!ctx.cronStore) return { error: "Scheduling is not available in this context." };
try {
const job = ctx.cronStore.create(args);
return { id: job.id, job };
} catch (err) {
return { error: (err as Error).message };
}
},
};
const cronCreateSchema = z.object({
cron: z.string().min(1),
prompt: z.string().min(1),
recurring: z.boolean().optional(),
durable: z.boolean().optional(),
});
export const cronListTool: ToolDef<z.infer<typeof cronListSchema>> = {
name: "cron_list",
description: "List all scheduled cron jobs with their id, schedule, prompt, and whether they're recurring/durable.",
schema: z.object({}),
mutating: false,
handler: async (_args, ctx) => {
return { jobs: ctx.cronStore?.list() ?? [] };
},
};
const cronListSchema = z.object({});
export const cronDeleteTool: ToolDef<z.infer<typeof cronDeleteSchema>> = {
name: "cron_delete",
description: "Cancel a scheduled cron job by id (from cron_list or cron_create's return). Returns { deleted: id } on success.",
schema: z.object({ id: z.string().min(1) }),
mutating: false,
handler: async (args, ctx) => {
const store = ctx.cronStore;
if (!store) return { error: "Scheduling is not available in this context." };
return store.delete(args.id) ? { deleted: args.id } : { error: `Job ${args.id} not found.` };
},
};
const cronDeleteSchema = z.object({ id: z.string().min(1) });
export const scheduleWakeupTool: ToolDef<z.infer<typeof scheduleWakeupSchema>> = {
name: "schedule_wakeup",
description:
"Schedule a one-shot prompt to fire after delaySeconds (60-3600), for self-paced loops that check back on external " +
"state. Pass stop: true to cancel ALL pending wakeups and end the loop. The prompt fires once then is removed. " +
"Only fires while the REPL is idle.",
schema: z.object({
delaySeconds: z.number().int().min(1),
prompt: z.string().min(1),
stop: z.boolean().optional(),
reason: z.string().optional(),
}),
mutating: false,
handler: async (args, ctx) => {
if (!ctx.cronStore) return { error: "Scheduling is not available in this context." };
const result = ctx.cronStore.scheduleWakeup(args);
return result;
},
};
const scheduleWakeupSchema = z.object({
delaySeconds: z.number().int().min(1),
prompt: z.string().min(1),
stop: z.boolean().optional(),
reason: z.string().optional(),
});
-121
View File
@@ -1,121 +0,0 @@
import { mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs";
import os from "node:os";
import path from "node:path";
import { afterEach, beforeEach, describe, expect, it } from "vitest";
import { detectEol, editFileTool, fromLF, toLF } from "./editFile.js";
import type { ToolContext } from "./types.js";
describe("editFile tool — path containment", () => {
let cwd: string;
let ctx: ToolContext;
beforeEach(() => {
cwd = mkdtempSync(path.join(os.tmpdir(), "locode-editfile-"));
ctx = { cwd };
});
afterEach(() => {
rmSync(cwd, { recursive: true, force: true });
});
it("edits a file inside the working directory", async () => {
writeFileSync(path.join(cwd, "note.txt"), "hello world");
const result = (await editFileTool.handler({ path: "note.txt", old_string: "world", new_string: "there" }, ctx)) as { replacements: number };
expect(result.replacements).toBe(1);
});
it("refuses to edit a file outside the working directory via ../ traversal", async () => {
await expect(
editFileTool.handler({ path: "../escape.txt", old_string: "a", new_string: "b" }, ctx),
).rejects.toThrow(/outside the working directory/);
});
it("preview reports the block instead of reading the target file", async () => {
const preview = await editFileTool.preview!({ path: "../escape.txt", old_string: "a", new_string: "b" }, ctx);
expect(preview).toMatch(/outside the working directory/);
});
it("handler suggests the closest match when old_string is not found", async () => {
writeFileSync(
path.join(cwd, "code.ts"),
"function greet(name: string): string {\n return `Hello, ${name}!`;\n}\n",
);
// Close but wrong: single quotes instead of backticks, "Hi" instead of "Hello".
await expect(
editFileTool.handler(
{ path: "code.ts", old_string: "return 'Hi, ${name}!';", new_string: "return `Hi, ${name}!`;" },
ctx,
),
).rejects.toThrow(/closest match/);
});
it("preview warns and shows the closest match when old_string is not found", async () => {
writeFileSync(path.join(cwd, "note.txt"), "the quick brown fox jumps over the lazy dog");
const preview = await editFileTool.preview!(
{ path: "note.txt", old_string: "the quick red fox jumps over the lazy cat", new_string: "x" },
ctx,
);
expect(preview).toMatch(/not found/);
expect(preview).toMatch(/closest match/);
expect(preview).toContain("quick brown fox");
});
it("does not suggest a match when nothing is remotely similar", async () => {
writeFileSync(path.join(cwd, "note.txt"), "aaaaaaaaaaaaaaaaaaaaaaaa");
await expect(
editFileTool.handler(
{ path: "note.txt", old_string: "completely different text xyz", new_string: "b" },
ctx,
),
).rejects.toThrow(/not found/);
// No "closest match" suffix when similarity is below the threshold.
await expect(
editFileTool.handler(
{ path: "note.txt", old_string: "completely different text xyz", new_string: "b" },
ctx,
),
).rejects.not.toThrow(/closest match/);
});
// read_file shows the model LF-normalized content regardless of the file's real line endings
// (see readFile.ts), so old_string/new_string from a model are always LF — matching must happen
// in that same space or every CRLF file in a project like this one fails with "not found".
it("matches an LF old_string against a CRLF file (mirrors what read_file shows the model)", async () => {
writeFileSync(path.join(cwd, "code.ts"), "function greet() {\r\n return 1;\r\n}\r\n");
const result = (await editFileTool.handler(
{ path: "code.ts", old_string: " return 1;", new_string: " return 2;" },
ctx,
)) as { replacements: number };
expect(result.replacements).toBe(1);
const onDisk = readFileSync(path.join(cwd, "code.ts"), "utf-8");
expect(onDisk).toBe("function greet() {\r\n return 2;\r\n}\r\n");
});
it("preserves CRLF line endings on disk after an edit spanning multiple lines", async () => {
writeFileSync(path.join(cwd, "code.ts"), "a\r\nb\r\nc\r\n");
await editFileTool.handler({ path: "code.ts", old_string: "a\nb", new_string: "a\nx\nb" }, ctx);
const onDisk = readFileSync(path.join(cwd, "code.ts"), "utf-8");
expect(onDisk).toBe("a\r\nx\r\nb\r\nc\r\n");
});
it("leaves a pure-LF file untouched by EOL conversion", async () => {
writeFileSync(path.join(cwd, "code.ts"), "a\nb\nc\n");
await editFileTool.handler({ path: "code.ts", old_string: "b", new_string: "x" }, ctx);
const onDisk = readFileSync(path.join(cwd, "code.ts"), "utf-8");
expect(onDisk).toBe("a\nx\nc\n");
});
});
describe("EOL helpers", () => {
it("detectEol finds CRLF, defaults to LF otherwise", () => {
expect(detectEol("a\r\nb")).toBe("\r\n");
expect(detectEol("a\nb")).toBe("\n");
expect(detectEol("a")).toBe("\n");
});
it("toLF/fromLF round-trip", () => {
expect(toLF("a\r\nb\r\nc")).toBe("a\nb\nc");
expect(fromLF("a\nb\nc", "\r\n")).toBe("a\r\nb\r\nc");
expect(fromLF("a\nb\nc", "\n")).toBe("a\nb\nc");
});
});
+23 -159
View File
@@ -1,8 +1,9 @@
import { createPatch } from "diff";
import { randomBytes } from "node:crypto";
import { readFile as fsReadFile, rename as fsRename, unlink as fsUnlink, writeFile as fsWriteFile } from "node:fs/promises";
import path from "node:path";
import { z } from "zod";
import { resolveWithinCwd } from "./pathGuard.js";
import { assertWithinWorkspace } from "../utils/path.js";
import type { ToolDef } from "./types.js";
const schema = z.object({
@@ -12,26 +13,6 @@ const schema = z.object({
replace_all: z.boolean().optional().describe("Replace every occurrence instead of requiring a unique match."),
});
/** `read_file` shows the model LF-normalized content (`content.split(/\r?\n/).join(...)` — see
* readFile.ts), regardless of the file's actual line endings on disk. A model's `old_string`/
* `new_string` are built from what it read, so they're always LF. Matching that against this
* tool's raw (real `\r\n`-preserving) file read would fail on every CRLF file in the project —
* which is most of them (see the repo's CRLF/LF notes). Detect the file's line ending once, do
* all matching/editing in LF space (so `old_string` from the model lines up), then convert the
* result back before writing so the file's on-disk convention is preserved rather than silently
* flipped to LF. */
export function detectEol(raw: string): "\r\n" | "\n" {
return raw.includes("\r\n") ? "\r\n" : "\n";
}
export function toLF(s: string): string {
return s.replace(/\r\n/g, "\n");
}
export function fromLF(s: string, eol: "\r\n" | "\n"): string {
return eol === "\n" ? s : s.replace(/\n/g, eol);
}
export function countOccurrences(haystack: string, needle: string): number {
return needle === "" ? 0 : haystack.split(needle).length - 1;
}
@@ -44,161 +25,41 @@ export function applyEdit(original: string, oldString: string, newString: string
return replaceAll ? original.split(oldString).join(newString) : original.replace(oldString, () => newString);
}
/**
* Normalise a candidate snippet for fuzzy comparison: collapse runs of whitespace to single spaces
* and trim. This makes the similarity score tolerant to indentation/line-ending differences, which
* are the most common reasons a local model's old_string almost-matches but not quite.
*/
function normaliseForCompare(s: string): string {
return s.replace(/\s+/g, " ").trim();
}
/**
* Compute a Levenshtein distance limited to `maxDist` — early-exits once the distance exceeds it,
* making it O(n*m) worst case but far cheaper in practice when we only care about "close enough".
*/
function boundedLevenshtein(a: string, b: string, maxDist: number): number {
const al = a.length;
const bl = b.length;
if (Math.abs(al - bl) > maxDist) return maxDist + 1;
if (al === 0) return bl;
if (bl === 0) return al;
let prev: number[] = new Array<number>(bl + 1);
let curr: number[] = new Array<number>(bl + 1);
for (let j = 0; j <= bl; j++) prev[j] = j;
for (let i = 1; i <= al; i++) {
curr[0] = i;
let rowMin = i;
const ai = a.charCodeAt(i - 1);
for (let j = 1; j <= bl; j++) {
const cost = ai === b.charCodeAt(j - 1) ? 0 : 1;
const del = (prev[j] ?? 0) + 1;
const ins = (curr[j - 1] ?? 0) + 1;
const sub = (prev[j - 1] ?? 0) + cost;
curr[j] = Math.min(del, ins, sub);
const cell = curr[j] ?? 0;
if (cell < rowMin) rowMin = cell;
}
// If every cell in this row already exceeds maxDist, the final answer can only be worse.
if (rowMin > maxDist) return maxDist + 1;
[prev, curr] = [curr, prev];
}
return prev[bl] ?? maxDist + 1;
}
interface SimilarMatch {
/** 0..1 similarity ratio (1 = identical, 0 = unrelated). */
score: number;
/** The exact text from the file at the best matching window. */
snippet: string;
/** 1-indexed line number where the snippet starts. */
line: number;
}
/**
* Find the region of `content` most similar to `needle`. Slides a window of the needle's length
* (±50%) across the file in word steps, scoring normalised text with bounded Levenshtein. Returns
* the best candidate when its similarity is at least 0.5 — clearly worth suggesting to the model.
* Returns null when nothing is close enough, in which case the caller falls back to the plain
* "not found" message.
*/
function findSimilarMatch(content: string, needle: string): SimilarMatch | null {
const needleNorm = normaliseForCompare(needle);
if (needleNorm.length < 3) return null;
const words = needleNorm.split(" ");
const minLen = Math.floor(needleNorm.length * 0.5);
const maxLen = Math.ceil(needleNorm.length * 1.5);
let best: SimilarMatch | null = null;
let bestDist = Infinity;
// Walk the file by character, treating every position as a potential window start is O(n*len)
// and too slow for big files. Instead, step at every Nth character (≈ word boundaries) to keep
// it cheap while still landing near real matches.
const step = Math.max(1, Math.floor(needleNorm.length / 8));
const contentLen = content.length;
for (let start = 0; start < contentLen; start += step) {
for (let len = minLen; len <= maxLen; len += step) {
const end = Math.min(start + len, contentLen);
const candidate = content.slice(start, end);
const candNorm = normaliseForCompare(candidate);
if (candNorm.length < minLen) continue;
// Only spend Levenshtein effort if the lengths are plausibly close.
const maxDist = Math.floor(needleNorm.length * 0.5);
const dist = boundedLevenshtein(needleNorm, candNorm, maxDist);
if (dist >= bestDist) continue;
bestDist = dist;
const score = 1 - dist / Math.max(needleNorm.length, candNorm.length);
// 1-indexed line: count newlines before `start`.
let line = 1;
for (let k = 0; k < start; k++) if (content.charCodeAt(k) === 10) line++;
best = { score, snippet: candidate.trim(), line };
}
}
if (best && best.score >= 0.5) return best;
return null;
}
/** Build a "did you mean" suffix for error/preview messages. Returns "" if nothing useful. */
function similarHint(original: string, oldString: string): string {
const m = findSimilarMatch(original, oldString);
if (!m) return "";
// Truncate long snippets so the message stays readable.
const snippet =
m.snippet.length > 300 ? `${m.snippet.slice(0, 300)}…` : m.snippet;
return `\n\nThe closest match in the file (line ${m.line}, ~${Math.round(m.score * 100)}% similar):\n"""\n${snippet}\n"""\nUse this exact text (or a unique subset of it) as old_string.`;
}
export const editFileTool: ToolDef<z.infer<typeof schema>> = {
name: "edit_file",
description:
"Replace exact text in a file. old_string must match exactly. Unless replace_all is set, it must be unique — include enough context. " +
"Use for small, targeted changes to an existing file. On a mismatch, the closest similar text is suggested to help retry.",
"Replace exact text in a file. old_string must match exactly. Read the file first, then use enough " +
"surrounding context in old_string to make it unique. Unless replace_all is set, duplicate matches are " +
"rejected. Use this for small, targeted changes; prefer write_file for new files or full rewrites.",
schema,
mutating: true,
preview: async ({ path: filePath, old_string, new_string, replace_all }, ctx) => {
let resolved: string;
const resolved = path.resolve(ctx.cwd, filePath);
assertWithinWorkspace(resolved, ctx.cwd, filePath);
let original: string;
try {
resolved = resolveWithinCwd(ctx.cwd, filePath);
} catch (err) {
return (err as Error).message;
}
let raw: string;
try {
raw = await fsReadFile(resolved, "utf-8");
original = await fsReadFile(resolved, "utf-8");
} catch {
return `File ${resolved} does not exist.`;
}
const eol = detectEol(raw);
const original = toLF(raw);
const oldLF = toLF(old_string);
const newLF = toLF(new_string);
const occurrences = countOccurrences(original, oldLF);
const occurrences = countOccurrences(original, old_string);
if (occurrences === 0) {
return `Warning: old_string not found in ${resolved} — this edit will fail.${similarHint(original, oldLF)}`;
return `Warning: old_string not found in ${resolved} — this edit will fail.`;
}
if (occurrences > 1 && !replace_all) {
return `Warning: old_string appears ${occurrences} times in ${resolved} — this edit will fail unless replace_all is set.`;
}
const updated = fromLF(applyEdit(original, oldLF, newLF, replace_all), eol);
return createPatch(resolved, raw, updated, "", "");
const updated = applyEdit(original, old_string, new_string, replace_all);
return createPatch(resolved, original, updated, "", "");
},
handler: async ({ path: filePath, old_string, new_string, replace_all }, ctx) => {
const resolved = resolveWithinCwd(ctx.cwd, filePath);
const raw = await fsReadFile(resolved, "utf-8");
const eol = detectEol(raw);
const original = toLF(raw);
const oldLF = toLF(old_string);
const newLF = toLF(new_string);
const occurrences = countOccurrences(original, oldLF);
const resolved = path.resolve(ctx.cwd, filePath);
assertWithinWorkspace(resolved, ctx.cwd, filePath);
const original = await fsReadFile(resolved, "utf-8");
const occurrences = countOccurrences(original, old_string);
if (occurrences === 0) {
throw new Error(
`old_string not found in ${filePath}. Make sure it matches the file exactly, including whitespace.${similarHint(original, oldLF)}`,
`old_string not found in ${filePath}. Make sure it matches the file exactly, including whitespace.`,
);
}
if (occurrences > 1 && !replace_all) {
@@ -206,7 +67,7 @@ export const editFileTool: ToolDef<z.infer<typeof schema>> = {
`old_string appears ${occurrences} times in ${filePath}. Provide more surrounding context to make it unique, or set replace_all: true.`,
);
}
const updated = fromLF(applyEdit(original, oldLF, newLF, replace_all), eol);
const updated = applyEdit(original, old_string, new_string, replace_all);
// Write to a temp file in the same directory, then rename — rename is atomic within a single
// directory, so a crash mid-write can't leave the user's source file half-overwritten (the live
// file stays intact until the rename swaps in the full new content). Clean up the temp file if
@@ -219,6 +80,9 @@ export const editFileTool: ToolDef<z.infer<typeof schema>> = {
await fsUnlink(tmp).catch(() => {});
throw err;
}
if (ctx.setLastEdit) {
ctx.setLastEdit({ path: filePath, previousContent: original });
}
return { path: resolved, replacements: replace_all ? occurrences : 1 };
},
};
};
+44
View File
@@ -0,0 +1,44 @@
import { z } from "zod";
import type { ToolDef } from "./types.js";
const schema = z.object({
plan: z
.string()
.describe(
"The full implementation plan in prose: the files you would change, the approach for each, and the key edits. " +
"Be concrete and actionable so the user can review it at a glance.",
),
});
/** The structured way to exit plan mode: the model calls this once it has finished researching and
* has a concrete plan, instead of presenting the plan as a final prose message. The handler shows
* the plan to the user via the same Approve/Reject prompt the prose path uses (see maybePresentPlan
* in agent/loop.ts); on approval plan mode ends and the model proceeds to implement in the same
* turn, on rejection it stays in plan mode and can refine. The prose path remains as a fallback for
* models that present a plan without calling this tool. Non-mutating: it changes permission mode, not
* the filesystem, so it passes the plan-mode mutating-tool gate. Only callable in plan mode — the
* ctx callback is withheld otherwise, so a call at the wrong time returns a clear error. */
export const exitPlanModeTool: ToolDef<z.infer<typeof schema>> = {
name: "exit_plan_mode",
description:
"Exit plan mode by presenting your implementation plan for the user's approval. Call this once you've " +
"finished researching and have a concrete plan (which files you'd change and how). On approval, plan mode " +
"ends and you implement the plan in this same turn. On rejection, stay in plan mode, refine the plan (explore " +
"more if needed), and call exit_plan_mode again. Only available in plan mode — don't call it otherwise.",
schema,
mutating: false,
handler: async (args, ctx) => {
if (!ctx.exitPlanMode) {
return { error: "exit_plan_mode is only available while plan mode is active." };
}
const { approved } = await ctx.exitPlanMode(args.plan);
if (approved) {
return { result: "Plan approved. Plan mode is now off — proceed to implement the plan now." };
}
return {
result:
"The user rejected the plan. Stay in plan mode, refine it (explore more if needed), and call " +
"exit_plan_mode again when ready. Do not call any mutating tool yet.",
};
},
};
+4 -2
View File
@@ -29,7 +29,8 @@ const statusSchema = z.object({
export const gitStatusTool: ToolDef<z.infer<typeof statusSchema>> = {
name: "git_status",
description:
"Inspect the git repo: `status`, `diff`, `log`, `show`, or `branches`. Read-only — no confirmation needed.",
"Inspect the git repo: `status`, `diff`, `log`, `show`, or `branches`. Read-only — no confirmation needed. " +
"Always check status/diff before mutating git_commit operations.",
schema: statusSchema,
mutating: false,
handler: async ({ operation, paths, staged, ref, maxCount }, ctx) => {
@@ -109,7 +110,8 @@ export const gitCommitTool: ToolDef<z.infer<typeof commitSchema>> = {
name: "git_commit",
description:
"Git operations: `add`, `commit`, `create_branch`, `checkout`, `push`, `reset`, `stash`, `merge`, `rebase`, " +
"`delete_branch`. Mutating operations require user confirmation with a preview.",
"`delete_branch`. Mutating operations require user confirmation with a preview. For commit, run git_status " +
"or git_commit add first, then provide a clear, concise message.",
schema: commitSchema,
mutating: true,
preview: async (args, ctx) => buildPreview(args, ctx.cwd),
+4 -1
View File
@@ -14,7 +14,10 @@ const schema = z.object({
export const grepTool: ToolDef<z.infer<typeof schema>> = {
name: "grep",
description: "Search file contents for a regular expression pattern using ripgrep. Use to find where a symbol/function/word is used across the codebase, or to locate files containing specific text. Faster than reading files one by one.",
description:
"Search file contents for a regular expression pattern using ripgrep. This is the best first step " +
"when exploring a codebase: use it to find where a symbol, function, or pattern is used, then " +
"read only the relevant files. If results are too broad, refine with `path` or `glob`.",
schema,
mutating: false,
handler: async ({ pattern, path: searchPath, glob, case_insensitive, max_results }, ctx) => {
+23 -8
View File
@@ -1,19 +1,25 @@
import { agentTool } from "./agentTool.js";
import { definitionTool, referencesTool, diagnosticsTool } from "./codeIntel.js";
import { askQuestionTool } from "./askQuestion.js";
import { bashTool } from "./bash.js";
import { bashKillTool } from "./bashKill.js";
import { bashOutputTool } from "./bashOutput.js";
import { cronCreateTool, cronDeleteTool, cronListTool, scheduleWakeupTool } from "./cron.js";
import { editFileTool } from "./editFile.js";
import { multiEditTool } from "./multiEdit.js";
import { exitPlanModeTool } from "./exitPlanMode.js";
import { gitCommitTool, gitStatusTool } from "./git.js";
import { grepTool } from "./grep.js";
import { listFilesTool } from "./listFiles.js";
import { memoryTool, memoryWriteTool } from "./memory.js";
import { multiEditTool } from "./multiEdit.js";
import { notebookEditTool } from "./notebookEdit.js";
import { readFileTool } from "./readFile.js";
import { todoWriteTool } from "./todoWrite.js";
import { taskCreateTool, taskListTool, taskGetTool, taskUpdateTool } from "./task.js";
import { sendMessageTool } from "./sendMessage.js";
import { listTeammatesTool } from "./teammates.js";
import { taskCreateTool, taskGetTool, taskListTool, taskUpdateTool } from "./task.js";
import { webFetchTool } from "./webFetch.js";
import { workflowTool } from "./workflow.js";
import { webSearchTool } from "./webSearch.js";
import { enterWorktreeTool, exitWorktreeTool } from "./worktreeSession.js";
import { writeFileTool } from "./writeFile.js";
import type { ToolDef } from "./types.js";
@@ -21,9 +27,6 @@ export const TOOLS: ToolDef[] = [
readFileTool,
listFilesTool,
grepTool,
definitionTool,
referencesTool,
diagnosticsTool,
webSearchTool,
webFetchTool,
gitStatusTool,
@@ -35,12 +38,24 @@ export const TOOLS: ToolDef[] = [
bashOutputTool,
bashKillTool,
gitCommitTool,
todoWriteTool,
taskCreateTool,
taskListTool,
taskGetTool,
taskUpdateTool,
memoryTool,
memoryWriteTool,
agentTool,
sendMessageTool,
listTeammatesTool,
workflowTool,
exitPlanModeTool,
askQuestionTool,
cronCreateTool,
cronListTool,
cronDeleteTool,
scheduleWakeupTool,
enterWorktreeTool,
exitWorktreeTool,
];
export const TOOL_REGISTRY: Map<string, ToolDef> = new Map(TOOLS.map((t) => [t.name, t]));
+4 -1
View File
@@ -17,7 +17,10 @@ const MAX_MATCHES = 500;
export const listFilesTool: ToolDef<z.infer<typeof schema>> = {
name: "list_files",
description: "List files matching a glob pattern (e.g. `src/**/*.ts`). Use to explore the project structure or find files by name/extension before reading them.",
description:
"List files matching a glob pattern. Use this to understand directory structure or find files " +
"by name. For searching file contents, use grep instead. Large result sets are truncated; narrow " +
"the pattern if you get too many matches.",
schema,
mutating: false,
handler: async ({ pattern, cwd }, ctx) => {
+123
View File
@@ -0,0 +1,123 @@
import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
import { existsSync, mkdtempSync, readFileSync, rmSync } from "node:fs";
import path from "node:path";
import os from "node:os";
import { _setConfigFilePathForTest } from "../config/store.js";
import { userMemoryDir, userMemoryIndexPath } from "../utils/userMemory.js";
import { memoryTool, memoryWriteTool } from "./memory.js";
import type { ToolContext } from "./types.js";
let tempDir: string;
const noCtx = {} as ToolContext;
beforeEach(() => {
tempDir = mkdtempSync(path.join(os.tmpdir(), "locode-memory-tool-"));
_setConfigFilePathForTest(path.join(tempDir, "config.json"));
});
afterEach(() => {
_setConfigFilePathForTest(undefined);
if (existsSync(tempDir)) rmSync(tempDir, { recursive: true, force: true });
});
describe("memory tool (read)", () => {
it("reports empty memory when nothing is saved", async () => {
const result = await memoryTool.handler({}, noCtx);
expect(result).toEqual({ content: "(memory is empty)", count: 0 });
});
it("reads a specific fact by name", async () => {
await memoryWriteTool.handler(
{ action: "write", name: "prefers-concise", description: "short answers", type: "user", content: "Be brief." },
noCtx,
);
const result = (await memoryTool.handler({ name: "prefers-concise" }, noCtx)) as { content: string };
expect(result.content).toContain("Be brief.");
expect(result.content).toContain("name: prefers-concise");
});
it("returns a not-found message for an unknown name", async () => {
const result = (await memoryTool.handler({ name: "nope" }, noCtx)) as { content: string };
expect(result.content).toContain("No memory fact named 'nope'");
});
it("lists all facts with no argument", async () => {
await memoryWriteTool.handler({ action: "write", name: "a-fact", description: "a", type: "user", content: "aa" }, noCtx);
await memoryWriteTool.handler({ action: "write", name: "b-fact", description: "b", type: "project", content: "bb" }, noCtx);
const result = (await memoryTool.handler({}, noCtx)) as { content: string; count: number };
expect(result.count).toBe(2);
expect(result.content).toContain("aa");
expect(result.content).toContain("bb");
});
});
describe("memory_write tool (write/delete)", () => {
it("creates a typed fact file + index on write", async () => {
const result = await memoryWriteTool.handler(
{ action: "write", name: "react-stack", description: "uses react", type: "project", content: "Stack is React + vitest." },
noCtx,
);
expect(result).toEqual({ written: "react-stack", type: "project", indexUpdated: true });
expect(existsSync(path.join(userMemoryDir(), "react-stack.md"))).toBe(true);
expect(readFileSync(userMemoryIndexPath(), "utf-8")).toContain("react-stack.md");
});
it("overwrites an existing fact on write", async () => {
await memoryWriteTool.handler({ action: "write", name: "flip", description: "old", type: "user", content: "old body" }, noCtx);
await memoryWriteTool.handler({ action: "write", name: "flip", description: "new", type: "reference", content: "new body" }, noCtx);
const file = readFileSync(path.join(userMemoryDir(), "flip.md"), "utf-8");
expect(file).toContain("new body");
expect(file).toContain("description: new");
// Index has a single line for flip.
expect(readFileSync(userMemoryIndexPath(), "utf-8").match(/flip\.md/g)).toHaveLength(1);
});
it("rejects a write missing required fields", async () => {
const result = await memoryWriteTool.handler({ action: "write", name: "x", content: "y" }, noCtx);
expect(result).toEqual({ error: "write requires non-empty 'description', 'type', and 'content'." });
});
it("deletes an existing fact and its index line", async () => {
await memoryWriteTool.handler({ action: "write", name: "gone", description: "x", type: "user", content: "yy" }, noCtx);
const result = await memoryWriteTool.handler({ action: "delete", name: "gone" }, noCtx);
expect(result).toEqual({ deleted: "gone" });
expect(existsSync(path.join(userMemoryDir(), "gone.md"))).toBe(false);
expect(readFileSync(userMemoryIndexPath(), "utf-8")).not.toContain("gone");
});
it("reports an error deleting a missing fact", async () => {
const result = await memoryWriteTool.handler({ action: "delete", name: "nope" }, noCtx);
expect(result).toEqual({ error: "No memory fact named 'nope' to delete." });
});
it("produces a diff preview for a write", async () => {
const preview = await memoryWriteTool.preview!(
{ action: "write", name: "fresh", description: "d", type: "user", content: "body text" },
noCtx,
);
expect(preview).toContain("Create memory fact fresh.md");
expect(preview).toContain("body text");
});
it("produces a diff preview when overwriting an existing fact", async () => {
await memoryWriteTool.handler({ action: "write", name: "p", description: "d", type: "user", content: "old" }, noCtx);
const preview = await memoryWriteTool.preview!(
{ action: "write", name: "p", description: "d", type: "user", content: "new" },
noCtx,
);
expect(preview).toContain("@@");
expect(preview).toContain("+new");
expect(preview).toContain("-old");
});
it("refreshes the session user-memory cache via setUserMemory after a write", async () => {
const setUserMemory = vi.fn();
const ctx = { setUserMemory } as unknown as ToolContext;
await memoryWriteTool.handler({ action: "write", name: "cached", description: "d", type: "user", content: "remember this" }, ctx);
expect(setUserMemory).toHaveBeenCalledOnce();
// The refreshed cache is the bounded index form loadUserMemory produces, carrying the new line.
const arg = setUserMemory.mock.calls[0]![0] as string | null;
expect(arg).toContain("Personal memory index");
expect(arg).toContain("cached");
});
});
+131
View File
@@ -0,0 +1,131 @@
import { createPatch } from "diff";
import { z } from "zod";
import {
deleteMemoryEntry,
listMemoryEntries,
loadUserMemory,
MEMORY_TYPES,
readMemoryEntry,
writeMemoryEntry,
type MemoryType,
} from "../utils/userMemory.js";
import type { ToolDef } from "./types.js";
// The memory files live in the user's locode config dir (see utils/userMemory.ts), NOT under the
// project workspace. So unlike write_file/edit_file these tools do NOT call assertWithinWorkspace —
// they always operate on files under the fixed `memory/` dir resolved via userMemoryDir() (test-
// override-aware), and the only user-supplied identifier is a kebab-case `name` validated against
// [a-z0-9-]+, so there's no traversal surface.
//
// Split into two tools because ToolDef.mutating is a static per-tool flag (the confirmation/plan-mode
// gate keys off it): `memory` is a read-only no-prompt tool usable even in plan mode, while
// `memory_write` mutates the user-level files and goes through the normal confirm gate.
async function readIndexOrLegacy(): Promise<string> {
const loaded = await loadUserMemory();
return loaded ?? "(memory is empty)";
}
export const memoryTool: ToolDef<{ name?: string }> = {
name: "memory",
description:
"Read the user's personal memory (the typed per-fact files in the locode config dir's memory/ directory). " +
"With no arguments, lists every saved fact (name, type, description, and full body). With a 'name', reads just " +
"that one fact. The lightweight index is already in your system prompt each turn — use this tool to pull a " +
"fact's full body when the index hook tells you it's relevant. Read-only — use memory_write to save or delete.",
schema: z.object({
name: z
.string()
.optional()
.describe("The slug name of a specific fact to read. Omit to list all saved facts."),
}),
mutating: false,
handler: async (args) => {
if (args.name) {
const entry = await readMemoryEntry(args.name);
if (!entry) return { content: `No memory fact named '${args.name}'.` };
return { content: `---\nname: ${entry.name}\ndescription: ${entry.description}\ntype: ${entry.type}\n---\n\n${entry.body}` };
}
const entries = await listMemoryEntries();
if (entries.length === 0) {
// No typed facts — surface the legacy freeform file if one exists so the model isn't blind to it.
const legacy = await readIndexOrLegacy();
return { content: legacy, count: 0 };
}
const rendered = entries.map((e) => `${e.name} (${e.type}): ${e.description}\n${e.body}`).join("\n\n---\n\n");
return { content: rendered, count: entries.length };
},
};
const writeSchema = z.object({
action: z
.enum(["write", "delete"])
.describe("write: create or overwrite a typed memory fact; delete: remove one."),
name: z
.string()
.describe("The fact's slug (kebab-case, [a-z0-9-]+). Used as the filename and the frontmatter 'name'."),
description: z
.string()
.optional()
.describe("One-line summary used as the recall hook in the always-in-prompt index. Required for action='write'."),
type: z
.enum(MEMORY_TYPES as [MemoryType, ...MemoryType[]])
.optional()
.describe("Fact category: user (who the user is), feedback (working-style guidance), project (ongoing work), reference (external pointers). Required for action='write'."),
content: z
.string()
.optional()
.describe("The fact body. Required for action='write'. Keep it a concise, self-contained fact; for feedback/project include a 'Why:' and 'How to apply:' line."),
});
export const memoryWriteTool: ToolDef<z.infer<typeof writeSchema>> = {
name: "memory_write",
description:
"Write or delete a typed personal memory fact (file under the locode config dir's memory/ directory). Each fact " +
"is one file with frontmatter (name/description/type) + a body; a one-line index is folded into every future " +
"session's system prompt so you can recall it. Use 'write' to save a durable fact worth remembering across " +
"sessions (a stated preference, a correction, a project convention) and 'delete' to remove one. Do not use this " +
"for transient per-task notes.",
schema: writeSchema,
mutating: true,
preview: async (args) => {
if (args.action === "delete") {
const existing = await readMemoryEntry(args.name);
if (!existing) return `No memory fact named '${args.name}' — nothing to delete.`;
return createPatch(`${args.name}.md`, serializeForPreview(existing), "", "", "");
}
if (!args.description || !args.type || !args.content) {
return "write requires 'name', 'description', 'type', and 'content'.";
}
const existing = await readMemoryEntry(args.name);
const next = serializeForPreview({ name: args.name, description: args.description, type: args.type, body: args.content });
if (!existing) return `Create memory fact ${args.name}.md:\n${next}`;
return createPatch(`${args.name}.md`, serializeForPreview(existing), next, "", "");
},
handler: async (args, ctx) => {
if (args.action === "delete") {
const removed = await deleteMemoryEntry(args.name);
if (!removed) return { error: `No memory fact named '${args.name}' to delete.` };
if (ctx.setUserMemory) ctx.setUserMemory(await loadUserMemory());
return { deleted: args.name };
}
// write
if (!args.description || !args.type || !args.content) {
return { error: "write requires non-empty 'description', 'type', and 'content'." };
}
const content = args.content.trim();
if (!content) return { error: "'content' must not be empty." };
const entry = await writeMemoryEntry({
name: args.name,
description: args.description.trim(),
type: args.type,
body: content,
});
if (ctx.setUserMemory) ctx.setUserMemory(await loadUserMemory());
return { written: entry.name, type: entry.type, indexUpdated: true };
},
};
function serializeForPreview(entry: { name: string; description: string; type: MemoryType; body: string }): string {
return `---\nname: ${entry.name}\ndescription: ${entry.description}\ntype: ${entry.type}\n---\n\n${entry.body.trim()}\n`;
}
+117 -129
View File
@@ -1,153 +1,141 @@
import { mkdtempSync, readFileSync, rmSync, writeFileSync } from "node:fs";
import os from "node:os";
import { mkdtemp, readFile, rm, writeFile } from "node:fs/promises";
import { tmpdir } from "node:os";
import path from "node:path";
import { afterEach, beforeEach, describe, expect, it } from "vitest";
import { describe, expect, it } from "vitest";
import { multiEditTool } from "./multiEdit.js";
import type { ToolContext } from "./types.js";
describe("multiEdit tool", () => {
let cwd: string;
let ctx: ToolContext;
async function makeCwd(): Promise<string> {
return mkdtemp(path.join(tmpdir(), "locode-multiedit-"));
}
beforeEach(() => {
cwd = mkdtempSync(path.join(os.tmpdir(), "locode-multiedit-"));
ctx = { cwd };
describe("multiEditTool", () => {
it("rejects paths that escape the working directory", async () => {
const cwd = await makeCwd();
try {
await expect(
multiEditTool.handler({ path: "../outside.txt", edits: [{ old_string: "a", new_string: "b" }] }, { cwd }),
).rejects.toThrow("Path resolves outside the working directory");
} finally {
await rm(cwd, { recursive: true, force: true });
}
});
afterEach(() => {
rmSync(cwd, { recursive: true, force: true });
});
it("applies an ordered batch to one file in a single atomic write", async () => {
writeFileSync(path.join(cwd, "code.ts"), "const A = 1;\nconst B = 2;\nconst C = 3;\n");
const result = (await multiEditTool.handler(
{
path: "code.ts",
edits: [
{ old_string: "const A = 1;", new_string: "const A = 10;" },
{ old_string: "const C = 3;", new_string: "const C = 30;" },
],
},
ctx,
)) as { applied: number };
expect(result.applied).toBe(2);
expect(readFileSync(path.join(cwd, "code.ts"), "utf-8")).toBe(
"const A = 10;\nconst B = 2;\nconst C = 30;\n",
);
});
it("an earlier edit can change the text a later edit matches", async () => {
writeFileSync(path.join(cwd, "f.txt"), "alpha\n");
await multiEditTool.handler(
{
path: "f.txt",
edits: [
{ old_string: "alpha", new_string: "beta" },
{ old_string: "beta", new_string: "gamma" },
],
},
ctx,
);
expect(readFileSync(path.join(cwd, "f.txt"), "utf-8")).toBe("gamma\n");
});
it("errors on the first edit that doesn't match, naming the edit index", async () => {
writeFileSync(path.join(cwd, "f.txt"), "alpha\n");
await expect(
multiEditTool.handler(
it("applies several edits to one file in order, as a single atomic write", async () => {
const cwd = await makeCwd();
try {
const file = path.join(cwd, "src.txt");
await writeFile(file, "alpha\nbeta\ngamma\n");
const result = (await multiEditTool.handler(
{
path: "f.txt",
path: "src.txt",
edits: [
{ old_string: "alpha", new_string: "beta" },
{ old_string: "missing", new_string: "x" },
{ old_string: "alpha", new_string: "ALPHA" },
{ old_string: "beta", new_string: "BETA" },
{ old_string: "gamma", new_string: "GAMMA" },
],
},
ctx,
),
).rejects.toThrow(/Edit 2: old_string not found/);
{ cwd },
)) as { path: string; applied: number };
expect(result.applied).toBe(3);
expect(await readFile(file, "utf-8")).toBe("ALPHA\nBETA\nGAMMA\n");
} finally {
await rm(cwd, { recursive: true, force: true });
}
});
it("errors when an old_string is ambiguous and replace_all is not set", async () => {
writeFileSync(path.join(cwd, "f.txt"), "dup\ndup\n");
await expect(
multiEditTool.handler(
it("applies later edits to the result of earlier ones", async () => {
const cwd = await makeCwd();
try {
const file = path.join(cwd, "src.txt");
await writeFile(file, "foo\n");
// First edit renames the line; second edit matches the renamed text.
await multiEditTool.handler(
{
path: "f.txt",
edits: [{ old_string: "dup", new_string: "x" }],
path: "src.txt",
edits: [
{ old_string: "foo", new_string: "bar" },
{ old_string: "bar", new_string: "baz" },
],
},
ctx,
),
).rejects.toThrow(/Edit 1: old_string appears 2 times/);
{ cwd },
);
expect(await readFile(file, "utf-8")).toBe("baz\n");
} finally {
await rm(cwd, { recursive: true, force: true });
}
});
it("replace_all applies to all occurrences within the batch step", async () => {
writeFileSync(path.join(cwd, "f.txt"), "dup\ndup\n");
await multiEditTool.handler(
{
path: "f.txt",
edits: [{ old_string: "dup", new_string: "x", replace_all: true }],
},
ctx,
);
expect(readFileSync(path.join(cwd, "f.txt"), "utf-8")).toBe("x\nx\n");
it("fails on the first non-unique match and writes nothing", async () => {
const cwd = await makeCwd();
try {
const file = path.join(cwd, "src.txt");
const original = "dup\ndup\nunique\n";
await writeFile(file, original);
await expect(
multiEditTool.handler(
{
path: "src.txt",
edits: [
{ old_string: "dup", new_string: "x" }, // ambiguous, no replace_all
{ old_string: "unique", new_string: "UNIQUE" },
],
},
{ cwd },
),
).rejects.toThrow(/appears 2 times/);
// The failed batch must not have written anything — the file is unchanged.
expect(await readFile(file, "utf-8")).toBe(original);
} finally {
await rm(cwd, { recursive: true, force: true });
}
});
it("refuses to edit outside the working directory", async () => {
await expect(
multiEditTool.handler(
{ path: "../escape.txt", edits: [{ old_string: "a", new_string: "b" }] },
ctx,
),
).rejects.toThrow(/outside the working directory/);
it("respects replace_all within a batch edit", async () => {
const cwd = await makeCwd();
try {
const file = path.join(cwd, "src.txt");
await writeFile(file, "dup\ndup\n");
await multiEditTool.handler(
{ path: "src.txt", edits: [{ old_string: "dup", new_string: "x", replace_all: true }] },
{ cwd },
);
expect(await readFile(file, "utf-8")).toBe("x\nx\n");
} finally {
await rm(cwd, { recursive: true, force: true });
}
});
it("preview produces a unified diff of the full batch", async () => {
writeFileSync(path.join(cwd, "f.txt"), "one\ntwo\n");
const preview = await multiEditTool.preview!(
{
path: "f.txt",
edits: [
{ old_string: "one", new_string: "ONE" },
{ old_string: "two", new_string: "TWO" },
],
},
ctx,
);
expect(preview).toMatch(/-one/);
expect(preview).toMatch(/\+ONE/);
expect(preview).toMatch(/-two/);
expect(preview).toMatch(/\+TWO/);
it("produces a diff preview spanning all edits", async () => {
const cwd = await makeCwd();
try {
const file = path.join(cwd, "src.txt");
await writeFile(file, "a\nb\n");
const preview = await multiEditTool.preview!(
{ path: "src.txt", edits: [{ old_string: "a", new_string: "A" }, { old_string: "b", new_string: "B" }] },
{ cwd },
);
expect(preview).toContain("@@");
expect(preview).toContain("+A");
expect(preview).toContain("+B");
} finally {
await rm(cwd, { recursive: true, force: true });
}
});
it("preview warns when a batch edit will fail", async () => {
writeFileSync(path.join(cwd, "f.txt"), "one\n");
const preview = await multiEditTool.preview!(
{
path: "f.txt",
edits: [{ old_string: "missing", new_string: "x" }],
},
ctx,
);
expect(preview).toMatch(/Edit 1: old_string not found/);
});
// Same LF-vs-CRLF mismatch as editFile.test.ts: read_file always shows the model LF content, so
// a batch's old_string/new_string must match against a CRLF file's LF-normalized text, and the
// result written back must preserve the file's original CRLF convention.
it("matches LF old_strings against a CRLF file and preserves CRLF on write", async () => {
writeFileSync(path.join(cwd, "code.ts"), "const A = 1;\r\nconst B = 2;\r\nconst C = 3;\r\n");
await multiEditTool.handler(
{
path: "code.ts",
edits: [
{ old_string: "const A = 1;", new_string: "const A = 10;" },
{ old_string: "const C = 3;", new_string: "const C = 30;" },
],
},
ctx,
);
expect(readFileSync(path.join(cwd, "code.ts"), "utf-8")).toBe(
"const A = 10;\r\nconst B = 2;\r\nconst C = 30;\r\n",
);
it("records the pre-edit content for /undo via setLastEdit", async () => {
const cwd = await makeCwd();
try {
const file = path.join(cwd, "src.txt");
const original = "a\nb\n";
await writeFile(file, original);
let captured: { path: string; previousContent: string } | undefined;
await multiEditTool.handler(
{ path: "src.txt", edits: [{ old_string: "a", new_string: "A" }] },
{ cwd, setLastEdit: (e: { path: string; previousContent: string }) => (captured = e) } as any,
);
expect(captured).toEqual({ path: "src.txt", previousContent: original });
} finally {
await rm(cwd, { recursive: true, force: true });
}
});
});
+25 -37
View File
@@ -1,41 +1,32 @@
import { createPatch } from "diff";
import { randomBytes } from "node:crypto";
import { readFile as fsReadFile, rename as fsRename, unlink as fsUnlink, writeFile as fsWriteFile } from "node:fs/promises";
import path from "node:path";
import { z } from "zod";
import { resolveWithinCwd } from "./pathGuard.js";
import { applyEdit, countOccurrences, detectEol, fromLF, toLF } from "./editFile.js";
import { assertWithinWorkspace } from "../utils/path.js";
import { applyEdit, countOccurrences } from "./editFile.js";
import type { ToolDef } from "./types.js";
// A single edit within a multi_edit batch. Mirrors edit_file's args minus `path` (which is shared
// across the whole batch). Each edit is applied in array order to the result of the previous one, so
// an earlier edit can shift the text a later edit matches — that's why each old_string is checked
// against the running result, not the original file.
// across the whole batch). Each edit is applied in array order to the result of the previous one.
const editSchema = z.object({
old_string: z.string().describe("Exact text to replace. Must match the current file content exactly at this point in the batch — earlier edits may have shifted it."),
old_string: z.string().describe("Exact text to replace. Must match the current file content exactly at this point in the batch."),
new_string: z.string().describe("Replacement text."),
replace_all: z.boolean().optional().describe("Replace every occurrence instead of requiring a unique match."),
});
const schema = z.object({
path: z.string().describe("File path to edit, relative to the working directory or absolute."),
edits: z.array(editSchema).min(1).describe("Ordered list of edits to apply to the same file, one after another. Each edit sees the result of the previous one."),
edits: z.array(editSchema).min(1).describe("Ordered list of edits to apply to the same file, one after another."),
});
/** Applies a batch of edits to an in-memory string, validating each. Throws on the first edit that
* doesn't match uniquely (unless its replace_all is set) or doesn't match at all. Edits apply to the
* running result, so an earlier edit can change the text a later edit matches. `original` and every
* edit's old_string/new_string must already be LF-normalized (see editFile.ts's detectEol/toLF —
* read_file shows the model LF-only content regardless of the file's real line endings, so matching
* must happen in that same space). */
function applyBatch(
original: string,
edits: { old_string: string; new_string: string; replace_all?: boolean }[],
filePath: string,
): string {
* running result, so an earlier edit can change the text a later edit matches. */
function applyBatch(original: string, edits: { old_string: string; new_string: string; replace_all?: boolean }[], filePath: string): string {
let current = original;
edits.forEach((edit, i) => {
const oldLF = toLF(edit.old_string);
const occurrences = countOccurrences(current, oldLF);
const occurrences = countOccurrences(current, edit.old_string);
if (occurrences === 0) {
throw new Error(
`Edit ${i + 1}: old_string not found in ${filePath}. Earlier edits may have shifted the text — re-read the file and adjust. Make sure it matches exactly, including whitespace.`,
@@ -46,7 +37,7 @@ function applyBatch(
`Edit ${i + 1}: old_string appears ${occurrences} times in ${filePath}. Provide more surrounding context to make it unique, or set replace_all: true.`,
);
}
current = applyEdit(current, oldLF, toLF(edit.new_string), edit.replace_all);
current = applyEdit(current, edit.old_string, edit.new_string, edit.replace_all);
});
return current;
}
@@ -56,39 +47,33 @@ export const multiEditTool: ToolDef<z.infer<typeof schema>> = {
description:
"Apply several edits to the same file in one call, in order. Each edit is {old_string, new_string, replace_all?}. " +
"Use this instead of repeated edit_file calls when you have multiple distinct changes to one file — it's one confirmation " +
"and one atomic write instead of N round-trips. Each old_string must match uniquely at its point in the batch " +
"(unless replace_all is set). Read the file first.",
"and one atomic write. Each old_string must match uniquely at its point in the batch (unless replace_all is set). Read the file first.",
schema,
mutating: true,
preview: async ({ path: filePath, edits }, ctx) => {
let resolved: string;
const resolved = path.resolve(ctx.cwd, filePath);
assertWithinWorkspace(resolved, ctx.cwd, filePath);
let original: string;
try {
resolved = resolveWithinCwd(ctx.cwd, filePath);
} catch (err) {
return (err as Error).message;
}
let raw: string;
try {
raw = await fsReadFile(resolved, "utf-8");
original = await fsReadFile(resolved, "utf-8");
} catch {
return `File ${resolved} does not exist.`;
}
try {
const eol = detectEol(raw);
const updated = fromLF(applyBatch(toLF(raw), edits, filePath), eol);
return createPatch(resolved, raw, updated, "", "");
const updated = applyBatch(original, edits, filePath);
return createPatch(resolved, original, updated, "", "");
} catch (err) {
return `Warning: ${(err as Error).message} — this edit will fail.`;
}
},
handler: async ({ path: filePath, edits }, ctx) => {
const resolved = resolveWithinCwd(ctx.cwd, filePath);
const raw = await fsReadFile(resolved, "utf-8");
const eol = detectEol(raw);
const updated = fromLF(applyBatch(toLF(raw), edits, filePath), eol);
const resolved = path.resolve(ctx.cwd, filePath);
assertWithinWorkspace(resolved, ctx.cwd, filePath);
const original = await fsReadFile(resolved, "utf-8");
const updated = applyBatch(original, edits, filePath);
// Atomic write via temp+rename (same rationale as edit_file): a crash mid-write can't leave the
// user's source file half-overwritten — the live file stays intact until the rename swaps in the
// full new content. Clean up the temp file if anything fails so a stray `.tmp` doesn't accumulate.
// full new content. Clean up the temp file if anything fails.
const tmp = `${resolved}.locode-${randomBytes(4).toString("hex")}.tmp`;
try {
await fsWriteFile(tmp, updated, "utf-8");
@@ -97,6 +82,9 @@ export const multiEditTool: ToolDef<z.infer<typeof schema>> = {
await fsUnlink(tmp).catch(() => {});
throw err;
}
if (ctx.setLastEdit) {
ctx.setLastEdit({ path: filePath, previousContent: original });
}
return { path: resolved, applied: edits.length };
},
};
+32 -28
View File
@@ -6,21 +6,19 @@ import { notebookEditTool } from "./notebookEdit.js";
/** A minimal valid nbformat 4 notebook with two code cells. */
function minimalNotebook(): string {
return (
JSON.stringify(
{
nbformat: 4,
nbformat_minor: 5,
metadata: {},
cells: [
{ cell_type: "code", id: "c1", source: ["print('a')\n"], metadata: {}, outputs: [], execution_count: null },
{ cell_type: "code", id: "c2", source: ["print('b')\n"], metadata: {}, outputs: [], execution_count: null },
],
},
null,
2,
) + "\n"
);
return JSON.stringify(
{
nbformat: 4,
nbformat_minor: 5,
metadata: {},
cells: [
{ cell_type: "code", id: "c1", source: ["print('a')\n"], metadata: {}, outputs: [], execution_count: null },
{ cell_type: "code", id: "c2", source: ["print('b')\n"], metadata: {}, outputs: [], execution_count: null },
],
},
null,
2,
) + "\n";
}
async function makeCwd(): Promise<string> {
@@ -36,8 +34,11 @@ describe("notebookEditTool", () => {
const cwd = await makeCwd();
try {
await expect(
notebookEditTool.handler({ notebook_path: "../outside.ipynb", edit_mode: "delete" }, { cwd }),
).rejects.toThrow(/outside the working directory/);
notebookEditTool.handler(
{ notebook_path: "../outside.ipynb", edit_mode: "delete" },
{ cwd },
),
).rejects.toThrow("Path resolves outside the working directory");
} finally {
await rm(cwd, { recursive: true, force: true });
}
@@ -114,20 +115,22 @@ describe("notebookEditTool", () => {
}
});
it("deletes a cell by cell_id, and a missing id fails", async () => {
it("deletes a cell by cell_id and writes nothing if not found", async () => {
const cwd = await makeCwd();
try {
const file = path.join(cwd, "nb.ipynb");
await writeFile(file, minimalNotebook());
const original = minimalNotebook();
await writeFile(file, original);
await notebookEditTool.handler({ notebook_path: "nb.ipynb", cell_id: "c1", edit_mode: "delete" }, { cwd });
const cells = parseCells(await readFile(file, "utf-8"));
expect(cells).toHaveLength(1);
expect(cells[0]!.id).toBe("c2");
// A missing id fails.
// A missing id fails and leaves the file unchanged.
await expect(
notebookEditTool.handler({ notebook_path: "nb.ipynb", cell_id: "nope", edit_mode: "delete" }, { cwd }),
).rejects.toThrow(/not found/);
expect(await readFile(file, "utf-8")).not.toBe(original);
} finally {
await rm(cwd, { recursive: true, force: true });
}
@@ -159,6 +162,8 @@ describe("notebookEditTool", () => {
{ cwd },
);
expect(preview).toContain("@@");
// The diff is over the notebook JSON, so the changed source appears JSON-quoted/indented
// on +/- lines rather than as bare text — assert on the content, not the leading marker.
expect(preview).toContain("print('A')");
expect(preview).toContain("print('a')");
} finally {
@@ -166,19 +171,18 @@ describe("notebookEditTool", () => {
}
});
it("switching a code cell to markdown drops code-only fields", async () => {
it("records pre-edit content for /undo via setLastEdit", async () => {
const cwd = await makeCwd();
try {
const file = path.join(cwd, "nb.ipynb");
await writeFile(file, minimalNotebook());
const original = minimalNotebook();
await writeFile(file, original);
let captured: { path: string; previousContent: string } | undefined;
await notebookEditTool.handler(
{ notebook_path: "nb.ipynb", cell_index: 0, edit_mode: "replace", cell_type: "markdown", new_source: "prose" },
{ cwd },
{ notebook_path: "nb.ipynb", cell_index: 1, edit_mode: "delete" },
{ cwd, setLastEdit: (e: { path: string; previousContent: string }) => (captured = e) } as any,
);
const cell = parseCells(await readFile(file, "utf-8"))[0]!;
expect(cell.cell_type).toBe("markdown");
expect(cell).not.toHaveProperty("execution_count");
expect(cell).not.toHaveProperty("outputs");
expect(captured).toEqual({ path: "nb.ipynb", previousContent: original });
} finally {
await rm(cwd, { recursive: true, force: true });
}
+11 -12
View File
@@ -1,15 +1,15 @@
import { createPatch } from "diff";
import { randomBytes } from "node:crypto";
import { readFile as fsReadFile, rename as fsRename, unlink as fsUnlink, writeFile as fsWriteFile } from "node:fs/promises";
import path from "node:path";
import { z } from "zod";
import { resolveWithinCwd } from "./pathGuard.js";
import { assertWithinWorkspace } from "../utils/path.js";
import type { ToolDef } from "./types.js";
// .ipynb is a JSON document (nbformat 4): { nbformat, nbformat_minor, metadata, cells: Cell[] }.
// Each cell is { cell_type: "code"|"markdown"|"raw", id?, source, metadata, outputs?, execution_count? }.
// `source` is a list of strings where every line except the last carries a trailing "\n" (nbformat
// convention). We convert the model's single-string new_source to/from that array form, so the
// model never has to hand-write nbformat's line-array quirk — it just gives the full cell text.
// convention). We convert the model's single-string new_source to/from that array form.
type Notebook = { nbformat: number; nbformat_minor: number; metadata: Record<string, unknown>; cells: Cell[] };
type Cell = { cell_type: string; id?: string; source: string[]; metadata: Record<string, unknown>; outputs?: unknown[]; execution_count?: unknown };
@@ -108,12 +108,8 @@ export const notebookEditTool: ToolDef<z.infer<typeof schema>> = {
schema,
mutating: true,
preview: async (args, ctx) => {
let resolved: string;
try {
resolved = resolveWithinCwd(ctx.cwd, args.notebook_path);
} catch (err) {
return (err as Error).message;
}
const resolved = path.resolve(ctx.cwd, args.notebook_path);
assertWithinWorkspace(resolved, ctx.cwd, args.notebook_path);
let original: string;
try {
original = await fsReadFile(resolved, "utf-8");
@@ -134,7 +130,8 @@ export const notebookEditTool: ToolDef<z.infer<typeof schema>> = {
}
},
handler: async (args, ctx) => {
const resolved = resolveWithinCwd(ctx.cwd, args.notebook_path);
const resolved = path.resolve(ctx.cwd, args.notebook_path);
assertWithinWorkspace(resolved, ctx.cwd, args.notebook_path);
const original = await fsReadFile(resolved, "utf-8");
let notebook: Notebook;
try {
@@ -145,8 +142,7 @@ export const notebookEditTool: ToolDef<z.infer<typeof schema>> = {
if (!Array.isArray(notebook.cells)) throw new Error(`${resolved} has no cells array — not a valid .ipynb.`);
applyNotebookEdit(notebook, args, args.notebook_path);
const updated = JSON.stringify(notebook, null, 2) + "\n";
// Atomic write via temp+rename (same rationale as edit_file/multi_edit): a crash mid-write can't
// leave the notebook half-overwritten. Clean up the temp file if anything fails.
// Atomic write via temp+rename (same rationale as edit_file/multi_edit).
const tmp = `${resolved}.locode-${randomBytes(4).toString("hex")}.tmp`;
try {
await fsWriteFile(tmp, updated, "utf-8");
@@ -155,6 +151,9 @@ export const notebookEditTool: ToolDef<z.infer<typeof schema>> = {
await fsUnlink(tmp).catch(() => {});
throw err;
}
if (ctx.setLastEdit) {
ctx.setLastEdit({ path: args.notebook_path, previousContent: original });
}
return { path: resolved, edit_mode: args.edit_mode ?? "replace", cell_count: notebook.cells.length };
},
};
-36
View File
@@ -1,36 +0,0 @@
import path from "node:path";
import { describe, expect, it } from "vitest";
import { PathOutsideCwdError, resolveWithinCwd } from "./pathGuard.js";
describe("resolveWithinCwd", () => {
const cwd = path.resolve("/project");
it("resolves a plain relative path inside cwd", () => {
expect(resolveWithinCwd(cwd, "src/index.ts")).toBe(path.join(cwd, "src", "index.ts"));
});
it("resolves an absolute path that happens to already be inside cwd", () => {
const inside = path.join(cwd, "foo.txt");
expect(resolveWithinCwd(cwd, inside)).toBe(inside);
});
it("resolves cwd itself", () => {
expect(resolveWithinCwd(cwd, ".")).toBe(cwd);
});
it("rejects a ../ escape", () => {
expect(() => resolveWithinCwd(cwd, "../outside.txt")).toThrow(PathOutsideCwdError);
});
it("rejects a deeper ../../ escape", () => {
expect(() => resolveWithinCwd(cwd, "sub/../../outside.txt")).toThrow(PathOutsideCwdError);
});
it("rejects an absolute path outside cwd", () => {
expect(() => resolveWithinCwd(cwd, path.resolve("/etc/passwd"))).toThrow(PathOutsideCwdError);
});
it("rejects the filesystem root", () => {
expect(() => resolveWithinCwd(cwd, path.parse(cwd).root)).toThrow(PathOutsideCwdError);
});
});
-27
View File
@@ -1,27 +0,0 @@
import path from "node:path";
/** Thrown by resolveWithinCwd — kept as its own class only so callers can recognize it (via
* instanceof) if they ever need to react differently than a plain thrown Error. */
export class PathOutsideCwdError extends Error {}
/** Resolves `targetPath` against `cwd` and hard-blocks the result if it would land outside the
* project root (the working directory locode was launched in) — an absolute path elsewhere on
* disk, a `../` escape, or (on Windows) a path on a different drive all reject. This applies
* unconditionally, regardless of permission mode: even `auto-accept` skips a tool's `preview`
* entirely (see gateAndRun in agent/loop.ts), so this check has to live in each tool's `handler`
* — which always runs — to actually hold as a floor rather than just a confirmation-dialog hint.
* It's deliberately not configurable; a model tricked (or simply mistaken) into targeting a path
* outside the project shouldn't be one auto-approved call away from touching it. */
export function resolveWithinCwd(cwd: string, targetPath: string): string {
const resolved = path.resolve(cwd, targetPath);
const rel = path.relative(cwd, resolved);
// rel === "" is targetPath resolving to cwd itself — fine. Anything starting with ".." walked
// upward out of cwd; an absolute rel (Windows: a different drive, e.g. "D:\foo") never went
// through cwd's tree in the first place. Either way, it's outside.
if (rel !== "" && (rel.startsWith(`..${path.sep}`) || rel === ".." || path.isAbsolute(rel))) {
throw new PathOutsideCwdError(
`Refusing to write outside the working directory: "${targetPath}" resolves to ${resolved}, which is not inside ${cwd}.`,
);
}
return resolved;
}
+6 -47
View File
@@ -2,25 +2,9 @@ import { readFile as fsReadFile, stat as fsStat } from "node:fs/promises";
import path from "node:path";
import { z } from "zod";
import { imageMimeType, MAX_IMAGE_BYTES } from "../utils/image.js";
import { assertWithinWorkspace } from "../utils/path.js";
import type { ToolDef } from "./types.js";
function withinWorkspace(resolved: string, workspace: string): boolean {
const rel = path.relative(workspace, resolved);
return !rel.startsWith("..") && !path.isAbsolute(rel);
}
// Reject files larger than this so we never accidentally OOM on a huge binary or log file.
const MAX_FILE_SIZE = 10 * 1024 * 1024; // 10 MB
// Common binary extensions — if the file extension matches, reject without reading.
const BINARY_EXTENSIONS = new Set([
".exe", ".dll", ".so", ".dylib", ".bin", ".dat", ".o", ".obj", ".pyc", ".pyo",
".class", ".jar", ".war", ".zip", ".tar", ".gz", ".bz2", ".7z", ".rar",
".iso", ".dmg", ".pdb", ".lib", ".a", ".woff", ".woff2", ".eot", ".ttf", ".otf",
".pdf", ".doc", ".docx", ".xls", ".xlsx", ".ppt", ".pptx",
".sqlite", ".db", ".ico", ".cur",
]);
// Cap on how much text a single read_file call returns, so a huge file can't blow up the context
// in one call. Cut on a line boundary (never mid-line) and report the exact next offset, so the
// model can page through the rest with `offset` instead of re-reading the same truncated prefix in
@@ -39,25 +23,17 @@ export const readFileTool: ToolDef<z.infer<typeof schema>> = {
description:
"Read a local file. Text files return 1-indexed lines; large files are paginated (use nextOffset for next page). " +
"Image files (png, jpg, jpeg, gif, webp, bmp) are returned as image content (requires vision-capable model). " +
"Use to inspect file contents before editing, or to understand existing code. Prefer this over bash cat for files.",
"When you need to inspect many files, use grep first to find the relevant ones and only read_file the files or page ranges you actually need — " +
"the session has a per-turn tool-call budget, and unnecessary full-file reads burn through it quickly.",
schema,
mutating: false,
handler: async ({ path: filePath, offset, limit }, ctx) => {
const resolved = path.resolve(ctx.cwd, filePath);
if (!withinWorkspace(resolved, ctx.cwd)) {
throw new Error(`File ${filePath} resolves outside the workspace.`);
}
// --- Size guard: reject files over MAX_FILE_SIZE before reading ---
const stats = await fsStat(resolved);
if (stats.size > MAX_FILE_SIZE) {
throw new Error(
`${filePath} is ${(stats.size / 1_048_576).toFixed(1)}MB, over the ${MAX_FILE_SIZE / 1_048_576}MB read limit.`,
);
}
assertWithinWorkspace(resolved, ctx.cwd, filePath);
const mimeType = imageMimeType(resolved);
if (mimeType) {
const stats = await fsStat(resolved);
if (stats.size > MAX_IMAGE_BYTES) {
throw new Error(
`${filePath} is ${(stats.size / 1_048_576).toFixed(1)}MB, over the ${MAX_IMAGE_BYTES / 1_048_576}MB limit for image reads.`,
@@ -67,25 +43,8 @@ export const readFileTool: ToolDef<z.infer<typeof schema>> = {
return { path: resolved, image: true, mimeType, bytes: buffer.byteLength, base64: buffer.toString("base64") };
}
// --- Binary guard: reject by extension ---
const ext = path.extname(resolved).toLowerCase();
if (BINARY_EXTENSIONS.has(ext)) {
throw new Error(
`${filePath} looks like a binary file (${ext}). Use bash for binary inspection.`,
);
}
const content = await fsReadFile(resolved, "utf-8");
// --- Binary guard: null-byte heuristic (catches extensionless binaries) ---
const nullIndex = content.indexOf("\0");
if (nullIndex !== -1) {
throw new Error(
`${filePath} appears to be a binary file (null byte at position ${nullIndex}). Use bash for binary inspection.`,
);
}
const lines = content.split(/\r?\n/);
const lines = content.split("\n");
const start = offset ? offset - 1 : 0;
const requestedEnd = limit ? Math.min(start + limit, lines.length) : lines.length;
+72
View File
@@ -0,0 +1,72 @@
import { z } from "zod";
import type { ToolDef } from "./types.js";
const schema = z
.object({
agentId: z
.string()
.optional()
.describe(
"The agentId returned by a prior 'agent'/'agent__*' call. Only resumable agents (single, " +
"shared-cwd delegations) return an agentId — parallel (worktree-isolated) agents do not. " +
"Provide this OR `name`.",
),
name: z
.string()
.optional()
.describe(
"The name of a teammate you created by passing `name` to a prior `agent` call. Addressing a " +
"teammate by name avoids tracking its agentId. Provide this OR `agentId`. Use list_teammates " +
"to see your named teammates.",
),
message: z
.string()
.describe("The follow-up instruction for the sub-agent. It retains the context of the original delegation."),
})
.refine((d) => d.agentId || d.name, {
message: "Provide either an `agentId` (returned by a prior agent call) or a `name` (of a named teammate).",
});
/** Continues a previously-spawned resumable sub-agent with a follow-up message, preserving its
* context — the cheaper alternative to re-delegating from scratch when a sub-agent's first answer
* was close but needs a correction, or when it ran out of budget mid-task. Only sub-agents that ran
* in the shared cwd (single/sequential delegations) are resumable and return an agentId; parallel
* worktree-isolated agents are fire-and-forget. The target may be identified by its agentId OR, if
* it was spawned with a `name`, by that name. Read-only: it has no filesystem side effects beyond
* what the continued sub-agent itself does (and those still go through the normal confirm gate). */
export const sendMessageTool: ToolDef<z.infer<typeof schema>> = {
name: "send_message",
description:
"Continue a previously-spawned resumable sub-agent (one that returned an agentId, or one you named via " +
"the `agent` tool's `name` arg) with a follow-up message, preserving its context. Cheaper than " +
"re-delegating from scratch. Use it to refine a sub-agent's answer, ask a follow-up, or continue one that " +
"ran out of its step budget. Identify it by `agentId` OR by `name` (a teammate name). Only shared-cwd agents " +
"are resumable; parallel worktree-isolated agents don't expose an agentId.",
schema,
mutating: false,
handler: async (args, ctx) => {
if (!ctx.resumeSubAgent) {
throw new Error("Sub-agent continuation is not available in this context.");
}
// Resolve the target: a name takes precedence (it's the human-friendly handle), but fall back to
// an explicit agentId. If a name is given but not on the roster, return a clean error instead of
// calling resumeSubAgent with an undefined agentId.
let agentId = args.agentId;
if (args.name) {
const resolved = ctx.resolveTeammate?.(args.name);
if (!resolved) {
return {
error: `No teammate named "${args.name}" was found. Use list_teammates to see named teammates, or pass the agentId returned by the original agent call.`,
};
}
agentId = resolved;
}
if (!agentId) {
return {
error: "Provide either a `name` (of a named teammate) or an `agentId` (returned by a prior agent call) to identify the sub-agent to continue.",
};
}
const result = await ctx.resumeSubAgent(agentId, args.message);
return { agentId, result, ...(args.name ? { name: args.name } : {}) };
},
};
+196 -92
View File
@@ -1,116 +1,220 @@
import { describe, it, expect } from "vitest";
import { taskCreateTool, taskListTool, taskGetTool, taskUpdateTool, TaskStore } from "./task.js";
import { describe, expect, it, vi } from "vitest";
import { TaskStore, taskCreateTool, taskGetTool, taskListTool, taskUpdateTool } from "./task.js";
import type { Task, TaskSummary } from "./task.js";
import type { ToolContext } from "./types.js";
function ctxWithStore(): { ctx: ToolContext; store: TaskStore } {
const store = new TaskStore();
return { ctx: { taskStore: store } as ToolContext, store };
function ctxWith(store?: TaskStore): ToolContext {
return { cwd: "/x", ...(store ? { taskStore: store } : {}) };
}
describe("task tools", () => {
it("task_create creates a pending task and returns it with an id", async () => {
const { ctx, store } = ctxWithStore();
const out = (await taskCreateTool.handler({ subject: "Fix bug", description: "details" }, ctx)) as {
id: string;
task: { status: string; blocks: string[]; blockedBy: string[] };
describe("TaskStore", () => {
it("creates tasks with sequential ids starting at t1 and pending status", () => {
const s = new TaskStore();
const a = s.create({ subject: "A", description: "do A" });
const b = s.create({ subject: "B", description: "do B", activeForm: "doing B" });
expect(a.id).toBe("t1");
expect(b.id).toBe("t2");
expect(a.status).toBe("pending");
expect(b.activeForm).toBe("doing B");
expect(a.blocks).toEqual([]);
expect(a.blockedBy).toEqual([]);
});
it("emits a snapshot via the change emitter after each mutation", () => {
const s = new TaskStore();
const snaps: ReturnType<TaskStore["list"]>[] = [];
s.setEmitter((tasks) => snaps.push(tasks));
s.create({ subject: "A", description: "x" });
s.create({ subject: "B", description: "y" });
expect(snaps).toHaveLength(2);
expect(snaps[1]!.map((t) => t.id)).toEqual(["t1", "t2"]);
});
it("updates status, subject, owner, and activeForm", () => {
const s = new TaskStore();
const t = s.create({ subject: "A", description: "x" });
const u = s.update(t.id, { status: "in_progress", owner: "agent-1", activeForm: "working" });
expect(u?.status).toBe("in_progress");
expect(u?.owner).toBe("agent-1");
expect(u?.activeForm).toBe("working");
});
it("links dependencies via addBlocks/addBlockedBy, ignoring self-refs, unknown ids, and duplicates", () => {
const s = new TaskStore();
const a = s.create({ subject: "A", description: "x" });
const b = s.create({ subject: "B", description: "y" });
// A blocks B: add B's blockedBy=[A] and A's blocks=[B].
s.update(b.id, { addBlockedBy: [a.id] });
s.update(a.id, { addBlocks: [b.id] });
expect(s.get(a.id)!.blocks).toEqual([b.id]);
expect(s.get(b.id)!.blockedBy).toEqual([a.id]);
// Self-ref, unknown id, and duplicate are all ignored.
s.update(a.id, { addBlocks: [a.id, "t99", b.id] });
expect(s.get(a.id)!.blocks).toEqual([b.id]);
});
it("prevents a direct 2-cycle when adding a blockedBy dependency", () => {
const s = new TaskStore();
const a = s.create({ subject: "A", description: "x" });
const b = s.create({ subject: "B", description: "y" });
s.update(a.id, { addBlockedBy: [b.id] }); // A waits on B
// Now B waiting on A would create a 2-cycle — skipped silently.
s.update(b.id, { addBlockedBy: [a.id] });
expect(s.get(b.id)!.blockedBy).toEqual([]);
});
it("merge-patches metadata, deleting keys set to null", () => {
const s = new TaskStore();
const t = s.create({ subject: "A", description: "x", metadata: { keep: 1, drop: 2 } });
s.update(t.id, { metadata: { added: 3, drop: null } });
expect(s.get(t.id)!.metadata).toEqual({ keep: 1, added: 3 });
});
it("deletes a task (status: 'deleted') and prunes dangling block/blockedBy refs", () => {
const s = new TaskStore();
const a = s.create({ subject: "A", description: "x" });
const b = s.create({ subject: "B", description: "y" });
s.update(b.id, { addBlockedBy: [a.id] });
s.update(a.id, { addBlocks: [b.id] });
expect(s.update(a.id, { status: "deleted" })).toBeUndefined();
expect(s.get(a.id)).toBeUndefined();
// B's blockedBy no longer references the deleted A.
expect(s.get(b.id)!.blockedBy).toEqual([]);
expect(s.get(b.id)!.blocks).toEqual([]);
});
it("update returns undefined for an unknown id", () => {
const s = new TaskStore();
expect(s.update("t99", { status: "in_progress" })).toBeUndefined();
});
});
describe("task_create tool", () => {
it("is non-mutating (no confirmation prompt)", () => {
expect(taskCreateTool.mutating).toBe(false);
});
it("creates a task and returns its id + snapshot", async () => {
const store = new TaskStore();
const result = (await taskCreateTool.handler(
{ subject: "Fix bug", description: "root-cause then patch", activeForm: "Fixing bug" },
ctxWith(store),
)) as { id: string; task: Task };
expect(result.id).toBe("t1");
expect(result.task.subject).toBe("Fix bug");
expect(result.task.activeForm).toBe("Fixing bug");
// Mutating the returned snapshot must not affect the store.
result.task.blocks.push("t99");
expect(store.get("t1")!.blocks).toEqual([]);
});
it("returns a clear error when no task store is available", async () => {
const result = (await taskCreateTool.handler({ subject: "x", description: "y" }, ctxWith(undefined))) as {
error: string;
};
expect(out.id).toMatch(/^t\d+$/);
expect(out.task.status).toBe("pending");
expect(out.task.blocks).toEqual([]);
expect(out.task.blockedBy).toEqual([]);
expect(store.list()).toHaveLength(1);
expect(result.error).toMatch(/not available/i);
});
});
describe("task_list tool", () => {
it("returns summaries of all tasks", async () => {
const store = new TaskStore();
store.create({ subject: "A", description: "x" });
store.create({ subject: "B", description: "y" });
const result = (await taskListTool.handler({}, ctxWith(store))) as { tasks: TaskSummary[] };
expect(result.tasks).toHaveLength(2);
expect(result.tasks.map((t) => t.subject)).toEqual(["A", "B"]);
});
it("task_list returns all task summaries", async () => {
const { ctx } = ctxWithStore();
await taskCreateTool.handler({ subject: "A", description: "x" }, ctx);
await taskCreateTool.handler({ subject: "B", description: "y" }, ctx);
const out = (await taskListTool.handler({}, ctx)) as { tasks: { subject: string }[] };
expect(out.tasks.map((t) => t.subject)).toEqual(["A", "B"]);
it("returns an empty list (not an error) when no store is available", async () => {
const result = (await taskListTool.handler({}, ctxWith(undefined))) as { tasks: TaskSummary[] };
expect(result.tasks).toEqual([]);
});
});
describe("task_get tool", () => {
it("returns the full task including dependencies and metadata", async () => {
const store = new TaskStore();
const a = store.create({ subject: "A", description: "x" });
const b = store.create({ subject: "B", description: "y" });
store.update(b.id, { addBlockedBy: [a.id] });
const result = (await taskGetTool.handler({ taskId: b.id }, ctxWith(store))) as { task?: Task };
expect(result.task!.blockedBy).toEqual([a.id]);
});
it("task_get returns full details and errors on unknown id", async () => {
const { ctx } = ctxWithStore();
const created = (await taskCreateTool.handler({ subject: "A", description: "long desc" }, ctx)) as {
id: string;
task: { description: string };
it("returns an error for an unknown id", async () => {
const store = new TaskStore();
const result = (await taskGetTool.handler({ taskId: "t99" }, ctxWith(store))) as { error: string };
expect(result.error).toMatch(/not found/i);
});
});
describe("task_update tool", () => {
it("updates status and returns the updated task", async () => {
const store = new TaskStore();
const t = store.create({ subject: "A", description: "x" });
const result = (await taskUpdateTool.handler({ taskId: t.id, status: "completed" }, ctxWith(store))) as {
task?: Task;
};
const got = (await taskGetTool.handler({ taskId: created.id }, ctx)) as { task: { description: string } };
expect(got.task.description).toBe("long desc");
const miss = (await taskGetTool.handler({ taskId: "nope" }, ctx)) as { error: string };
expect(miss.error).toMatch(/not found/);
expect(result.task!.status).toBe("completed");
});
it("task_update sets status and marks in_progress/completed", async () => {
const { ctx } = ctxWithStore();
const created = (await taskCreateTool.handler({ subject: "A", description: "x" }, ctx)) as { id: string };
const upd = (await taskUpdateTool.handler({ taskId: created.id, status: "in_progress" }, ctx)) as {
task: { status: string };
it("deletes when status is 'deleted' and returns { deleted }", async () => {
const store = new TaskStore();
const t = store.create({ subject: "A", description: "x" });
const result = (await taskUpdateTool.handler({ taskId: t.id, status: "deleted" }, ctxWith(store))) as {
deleted: string;
};
expect(upd.task.status).toBe("in_progress");
const done = (await taskUpdateTool.handler({ taskId: created.id, status: "completed" }, ctx)) as {
task: { status: string };
expect(result.deleted).toBe(t.id);
expect(store.get(t.id)).toBeUndefined();
});
it("adds dependencies via addBlockedBy", async () => {
const store = new TaskStore();
const a = store.create({ subject: "A", description: "x" });
const b = store.create({ subject: "B", description: "y" });
await taskUpdateTool.handler({ taskId: b.id, addBlockedBy: [a.id] }, ctxWith(store));
expect(store.get(b.id)!.blockedBy).toEqual([a.id]);
});
it("returns an error for an unknown id", async () => {
const store = new TaskStore();
const result = (await taskUpdateTool.handler({ taskId: "t99", status: "in_progress" }, ctxWith(store))) as {
error: string;
};
expect(done.task.status).toBe("completed");
expect(result.error).toMatch(/not found/i);
});
it("addBlocks/addBlockedBy link two tasks both ways", async () => {
const { ctx } = ctxWithStore();
const a = (await taskCreateTool.handler({ subject: "A", description: "x" }, ctx)) as { id: string };
const b = (await taskCreateTool.handler({ subject: "B", description: "y" }, ctx)) as { id: string };
// B is blocked by A (one-directional: sets B.blockedBy, not A.blocks)
await taskUpdateTool.handler({ taskId: b.id, addBlockedBy: [a.id] }, ctx);
const bAfter = (await taskGetTool.handler({ taskId: b.id }, ctx)) as { task: { blockedBy: string[] } };
expect(bAfter.task.blockedBy).toContain(a.id);
// Add the back-ref explicitly: A blocks B.
await taskUpdateTool.handler({ taskId: a.id, addBlocks: [b.id] }, ctx);
const aAfter = (await taskGetTool.handler({ taskId: a.id }, ctx)) as { task: { blocks: string[] } };
expect(aAfter.task.blocks).toContain(b.id);
it("returns a clear error when no task store is available", async () => {
const result = (await taskUpdateTool.handler({ taskId: "t1", status: "in_progress" }, ctxWith(undefined))) as {
error: string;
};
expect(result.error).toMatch(/not available/i);
});
it("ignores self-refs, unknown ids, and direct 2-cycles", async () => {
const { ctx } = ctxWithStore();
const a = (await taskCreateTool.handler({ subject: "A", description: "x" }, ctx)) as { id: string };
const b = (await taskCreateTool.handler({ subject: "B", description: "y" }, ctx)) as { id: string };
// self-ref ignored
await taskUpdateTool.handler({ taskId: a.id, addBlockedBy: [a.id] }, ctx);
expect(((await taskGetTool.handler({ taskId: a.id }, ctx)) as { task: { blockedBy: string[] } }).task.blockedBy).toEqual([]);
// unknown id ignored
await taskUpdateTool.handler({ taskId: a.id, addBlockedBy: ["zzz"] }, ctx);
expect(((await taskGetTool.handler({ taskId: a.id }, ctx)) as { task: { blockedBy: string[] } }).task.blockedBy).toEqual([]);
// B waits on A; now make A wait on B — should be skipped (2-cycle)
await taskUpdateTool.handler({ taskId: b.id, addBlockedBy: [a.id] }, ctx);
await taskUpdateTool.handler({ taskId: a.id, addBlockedBy: [b.id] }, ctx);
expect(((await taskGetTool.handler({ taskId: a.id }, ctx)) as { task: { blockedBy: string[] } }).task.blockedBy).toEqual([]);
it("the store emitter fires after a tool-driven update", async () => {
const store = new TaskStore();
const fired = vi.fn();
store.setEmitter(fired);
const t = store.create({ subject: "A", description: "x" });
fired.mockClear(); // create already fired once; isolate the update
await taskUpdateTool.handler({ taskId: t.id, status: "in_progress" }, ctxWith(store));
expect(fired).toHaveBeenCalledTimes(1);
});
});
it("status deleted removes the task and prunes dangling refs", async () => {
const { ctx } = ctxWithStore();
const a = (await taskCreateTool.handler({ subject: "A", description: "x" }, ctx)) as { id: string };
const b = (await taskCreateTool.handler({ subject: "B", description: "y" }, ctx)) as { id: string };
await taskUpdateTool.handler({ taskId: b.id, addBlockedBy: [a.id] }, ctx);
const del = (await taskUpdateTool.handler({ taskId: a.id, status: "deleted" }, ctx)) as { deleted: string };
expect(del.deleted).toBe(a.id);
// B no longer blocked by the removed A
const bAfter = (await taskGetTool.handler({ taskId: b.id }, ctx)) as { task: { blockedBy: string[] } };
expect(bAfter.task.blockedBy).toEqual([]);
expect(((await taskListTool.handler({}, ctx)) as { tasks: unknown[] }).tasks).toHaveLength(1);
describe("task tool schema validation", () => {
it("rejects an empty subject on task_create", () => {
expect(() => taskCreateTool.schema.parse({ subject: "", description: "x" })).toThrow();
});
it("metadata merge-patch: set keys, null deletes", async () => {
const { ctx } = ctxWithStore();
const a = (await taskCreateTool.handler({ subject: "A", description: "x", metadata: { k: 1 } }, ctx)) as { id: string };
await taskUpdateTool.handler({ taskId: a.id, metadata: { k2: "v" } }, ctx);
let t = (await taskGetTool.handler({ taskId: a.id }, ctx)) as { task: { metadata: Record<string, unknown> } };
expect(t.task.metadata).toEqual({ k: 1, k2: "v" });
await taskUpdateTool.handler({ taskId: a.id, metadata: { k: null } }, ctx);
t = (await taskGetTool.handler({ taskId: a.id }, ctx)) as { task: { metadata: Record<string, unknown> } };
expect(t.task.metadata).toEqual({ k2: "v" });
it("rejects an empty taskId on task_get/task_update", () => {
expect(() => taskGetTool.schema.parse({ taskId: "" })).toThrow();
expect(() => taskUpdateTool.schema.parse({ taskId: "" })).toThrow();
});
it("returns an error when taskStore is absent", async () => {
const ctx = {} as ToolContext;
const out = (await taskCreateTool.handler({ subject: "A", description: "x" }, ctx)) as { error: string };
expect(out.error).toMatch(/not available/);
it("rejects an unknown status value on task_update", () => {
expect(() => taskUpdateTool.schema.parse({ taskId: "t1", status: "done" })).toThrow();
});
it("accepts 'deleted' as a status on task_update", () => {
expect(() => taskUpdateTool.schema.parse({ taskId: "t1", status: "deleted" })).not.toThrow();
});
});
+63 -72
View File
@@ -7,8 +7,8 @@ export type TaskStatus = "pending" | "in_progress" | "completed";
/** A structured, trackable unit of work. Tasks form a dependency graph via `blocks`/`blockedBy`
* (each lists the other's task ids), can be owned/claimed by a named agent, and carry free-form
* metadata. Unlike the flat todo list, tasks are created and updated incrementally (not replaced
* wholesale) so dependencies and ownership can be expressed. */
* metadata. Unlike the old flat todo list, tasks are created and updated incrementally (not
* replaced wholesale) so dependencies and ownership can be expressed. */
export interface Task {
id: string;
subject: string;
@@ -36,21 +36,15 @@ export interface TaskSummary {
blockedBy: string[];
}
/** JSON-serializable shape for persisting a TaskStore. */
export interface TaskStoreSnapshot {
seq: number;
tasks: Task[];
}
/** In-memory task store. The task tools operate on it via `ctx.taskStore`. Mutations emit a
* snapshot through an optional `onChange` callback the loop wires up, so each create/update can
* refresh a UI checklist. Purely in-memory (per-session); not persisted. */
/** In-memory task store held by the Session. The tools below operate on it via `ctx.taskStore`.
* Mutations emit a snapshot to the UI through the `onChange` emitter the loop wires up, so each
* create/update renders a fresh checklist item in the scrollback (matching the old todo behavior). */
export class TaskStore {
private tasks = new Map<string, Task>();
private seq = 0;
private emitter: ((tasks: TaskSummary[]) => void) | undefined;
/** Wired by the host so the store can broadcast a snapshot after each mutation. */
/** Wired by the agent loop so the store can broadcast a snapshot after each mutation. */
setEmitter(emit: (tasks: TaskSummary[]) => void): void {
this.emitter = emit;
}
@@ -160,24 +154,6 @@ export class TaskStore {
}
this.emit();
}
/** Serialize the entire store for persistence. */
toJSON(): TaskStoreSnapshot {
return {
seq: this.seq,
tasks: [...this.tasks.values()].map(serializeTask),
};
}
/** Restore a store from a previously-serialized snapshot. */
static fromJSON(snapshot: TaskStoreSnapshot): TaskStore {
const store = new TaskStore();
store.seq = snapshot.seq;
for (const task of snapshot.tasks) {
store.tasks.set(task.id, { ...task, blocks: [...task.blocks], blockedBy: [...task.blockedBy], metadata: task.metadata ? { ...task.metadata } : undefined });
}
return store;
}
}
/** Merge-patches metadata: a null value deletes the key, any other value sets it. */
@@ -206,23 +182,21 @@ function serializeTask(t: Task): Task {
const metadataSchema = z.record(z.string(), z.any()).optional();
const taskCreateSchema = z.object({
subject: z.string().min(1).describe("A brief, actionable title in imperative form (e.g. 'Fix authentication bug')."),
description: z.string().describe("What needs to be done, in enough detail to act on."),
activeForm: z
.string()
.optional()
.describe("Present-continuous label shown in the spinner while in_progress (e.g. 'Running tests'). Optional."),
metadata: metadataSchema,
});
export const taskCreateTool: ToolDef<z.infer<typeof taskCreateSchema>> = {
name: "task_create",
description:
"Create a structured task to track a unit of multi-step work. Use for non-trivial work (3+ steps) so progress is " +
"visible and dependencies can be expressed. Returns the new task with its id. Call task_list to see all tasks, " +
"task_get for full details, and task_update to set status, add dependencies (addBlocks/addBlockedBy), or claim ownership.",
schema: taskCreateSchema,
"visible and dependencies can be expressed. Returns the new task with its id. Call task_list to see all tasks, " +
"task_get for full details, and task_update to set status, add dependencies (addBlocks/addBlockedBy), or claim ownership.",
schema: z.object({
subject: z.string().min(1).describe("A brief, actionable title in imperative form (e.g. 'Fix authentication bug')."),
description: z.string().describe("What needs to be done, in enough detail to act on."),
activeForm: z
.string()
.optional()
.describe("Present-continuous label shown in the spinner while in_progress (e.g. 'Running tests'). Optional."),
metadata: metadataSchema,
}),
// Purely informational (tracks state in-memory, never touches the filesystem) — no confirmation prompt.
mutating: false,
handler: async (args, ctx) => {
@@ -232,28 +206,33 @@ export const taskCreateTool: ToolDef<z.infer<typeof taskCreateSchema>> = {
},
};
const taskListSchema = z.object({});
const taskCreateSchema = z.object({
subject: z.string().min(1),
description: z.string(),
activeForm: z.string().optional(),
metadata: metadataSchema,
});
export const taskListTool: ToolDef<z.infer<typeof taskListSchema>> = {
name: "task_list",
description:
"List all tasks with their id, subject, status, owner, and what blocks them. Use this to see overall progress and " +
"find the next available task to claim.",
schema: taskListSchema,
"find the next available task to claim.",
schema: z.object({}),
mutating: false,
handler: async (_args, ctx) => {
return { tasks: ctx.taskStore?.list() ?? [] };
},
};
const taskGetSchema = z.object({ taskId: z.string().min(1) });
const taskListSchema = z.object({});
export const taskGetTool: ToolDef<z.infer<typeof taskGetSchema>> = {
name: "task_get",
description:
"Get a task's full details (description, activeForm, blocks, blockedBy, metadata). Use before starting a task to " +
"verify its blockedBy list is empty — if it isn't, the blocking tasks must complete first.",
schema: taskGetSchema,
"verify its blockedBy list is empty — if it isn't, the blocking tasks must complete first.",
schema: z.object({ taskId: z.string().min(1) }),
mutating: false,
handler: async (args, ctx) => {
const task = ctx.taskStore?.get(args.taskId);
@@ -261,6 +240,40 @@ export const taskGetTool: ToolDef<z.infer<typeof taskGetSchema>> = {
},
};
const taskGetSchema = z.object({ taskId: z.string().min(1) });
export const taskUpdateTool: ToolDef<z.infer<typeof taskUpdateSchema>> = {
name: "task_update",
description:
"Update a task: set status (pending|in_progress|completed — or 'deleted' to remove it), rename subject/description, " +
"set owner to claim it, add dependencies via addBlocks/addBlockedBy (task ids), or merge-patch metadata (set a key " +
"to null to delete it). Mark a task in_progress when starting it and completed when done. Verify blockedBy is empty " +
"before starting. Returns the updated task, or { deleted: id } when status is 'deleted'.",
schema: z.object({
taskId: z.string().min(1),
status: z.enum(["pending", "in_progress", "completed", "deleted"]).optional(),
subject: z.string().optional(),
description: z.string().optional(),
activeForm: z.string().optional(),
owner: z.string().optional(),
addBlocks: z.array(z.string()).optional(),
addBlockedBy: z.array(z.string()).optional(),
metadata: metadataSchema,
}),
mutating: false,
handler: async (args, ctx) => {
const store = ctx.taskStore;
if (!store) return { error: "Task tracking is not available in this context." };
if (!store.get(args.taskId)) return { error: `Task ${args.taskId} not found.` };
if (args.status === "deleted") {
store.update(args.taskId, args);
return { deleted: args.taskId };
}
const task = store.update(args.taskId, args);
return task ? { task: serializeTask(task) } : { deleted: args.taskId };
},
};
const taskUpdateSchema = z.object({
taskId: z.string().min(1),
status: z.enum(["pending", "in_progress", "completed", "deleted"]).optional(),
@@ -271,26 +284,4 @@ const taskUpdateSchema = z.object({
addBlocks: z.array(z.string()).optional(),
addBlockedBy: z.array(z.string()).optional(),
metadata: metadataSchema,
});
export const taskUpdateTool: ToolDef<z.infer<typeof taskUpdateSchema>> = {
name: "task_update",
description:
"Update a task: set status (pending|in_progress|completed — or 'deleted' to remove it), rename subject/description, " +
"set owner to claim it, add dependencies via addBlocks/addBlockedBy (task ids), or merge-patch metadata (set a key " +
"to null to delete it). Mark a task in_progress when starting it and completed when done. Verify blockedBy is empty " +
"before starting. Returns the updated task, or { deleted: id } when status is 'deleted'.",
schema: taskUpdateSchema,
mutating: false,
handler: async (args, ctx) => {
const store = ctx.taskStore;
if (!store) return { error: "Task tracking is not available in this context." };
if (!store.get(args.taskId)) return { error: `Task ${args.taskId} not found.` };
if (args.status === "deleted") {
store.update(args.taskId, args);
return { deleted: args.taskId };
}
const task = store.update(args.taskId, args);
return task ? { task: serializeTask(task) } : { deleted: args.taskId };
},
};
});
+165
View File
@@ -0,0 +1,165 @@
import { describe, expect, it, vi } from "vitest";
import { agentTool } from "./agentTool.js";
import { sendMessageTool } from "./sendMessage.js";
import { listTeammatesTool } from "./teammates.js";
import type { ToolContext } from "./types.js";
// A mock ctx whose teammate roster is a real Map, so registration/resolution/list behave like the
// session-backed wiring in loop.ts (which sets registerTeammate/resolveTeammate/listTeammates over
// session.namedAgents). runSubAgent/resumeSubAgent are vi mocks the individual tests program.
function ctxWithRoster(opts?: {
runSubAgent?: ToolContext["runSubAgent"];
resumeSubAgent?: ToolContext["resumeSubAgent"];
}): { ctx: ToolContext; roster: Map<string, string> } {
const roster = new Map<string, string>();
const ctx: ToolContext = {
cwd: "/x",
registerTeammate: (name, agentId) => {
roster.set(name, agentId);
},
resolveTeammate: (name) => roster.get(name),
listTeammates: () => [...roster.entries()].map(([name, agentId]) => ({ name, agentId })),
...(opts?.runSubAgent ? { runSubAgent: opts.runSubAgent } : {}),
...(opts?.resumeSubAgent ? { resumeSubAgent: opts.resumeSubAgent } : {}),
};
return { ctx, roster };
}
describe("agent tool — named teammates", () => {
it("registers a name on the roster when the delegation is resumable, and echoes name + agentId", async () => {
const runSubAgent = vi.fn().mockResolvedValue({ agentId: "agent-1", result: "did it", resumable: true });
const { ctx, roster } = ctxWithRoster({ runSubAgent });
const result = (await agentTool.handler(
{ description: "research X", prompt: "find X", name: "researcher" },
ctx,
)) as { name: string; agentId: string; result: string };
expect(result.name).toBe("researcher");
expect(result.agentId).toBe("agent-1");
expect(roster.get("researcher")).toBe("agent-1");
expect(runSubAgent).toHaveBeenCalledOnce();
});
it("does NOT register a name (and omits name/agentId) when the delegation is NOT resumable (parallel/isolated)", async () => {
const runSubAgent = vi.fn().mockResolvedValue({ agentId: "agent-2", result: "analyzed", resumable: false });
const { ctx, roster } = ctxWithRoster({ runSubAgent });
const result = (await agentTool.handler(
{ description: "review Y", prompt: "review Y", name: "reviewer" },
ctx,
)) as { result: string; name?: string; agentId?: string };
expect(result.result).toBe("analyzed");
expect(result.name).toBeUndefined();
expect(result.agentId).toBeUndefined();
expect(roster.has("reviewer")).toBe(false);
});
it("returns an error WITHOUT running when the name is already a teammate (no clobber)", async () => {
const runSubAgent = vi.fn();
const { ctx, roster } = ctxWithRoster({ runSubAgent });
roster.set("researcher", "agent-1");
const result = (await agentTool.handler(
{ description: "research Z", prompt: "find Z", name: "researcher" },
ctx,
)) as { error: string };
expect(result.error).toMatch(/already exists/i);
expect(runSubAgent).not.toHaveBeenCalled();
// The existing teammate is untouched.
expect(roster.get("researcher")).toBe("agent-1");
});
it("still runs (and silently ignores the name) when no roster is wired (non-session context)", async () => {
// No resolveTeammate/registerTeammate on ctx — simulates a hand-built, non-session context.
const runSubAgent = vi.fn().mockResolvedValue({ agentId: "agent-9", result: "ok", resumable: true });
const ctx: ToolContext = { cwd: "/x", runSubAgent };
const result = (await agentTool.handler(
{ description: "solo", prompt: "do solo", name: "lonely" },
ctx,
)) as { agentId: string; name?: string };
expect(runSubAgent).toHaveBeenCalledOnce();
expect(result.agentId).toBe("agent-9");
// No roster to register on → name simply not echoed as a teammate handle.
expect(result.name).toBeUndefined();
});
});
describe("send_message — name vs agentId", () => {
it("resolves a teammate by name, then calls resumeSubAgent with its agentId", async () => {
const resumeSubAgent = vi.fn().mockResolvedValue("followed up");
const { ctx, roster } = ctxWithRoster({ resumeSubAgent });
roster.set("researcher", "agent-1");
const result = (await sendMessageTool.handler({ name: "researcher", message: "go deeper" }, ctx)) as {
agentId: string;
name: string;
result: string;
};
expect(resumeSubAgent).toHaveBeenCalledWith("agent-1", "go deeper");
expect(result.agentId).toBe("agent-1");
expect(result.name).toBe("researcher");
expect(result.result).toBe("followed up");
});
it("returns a clean error (without resuming) for an unknown name", async () => {
const resumeSubAgent = vi.fn();
const { ctx } = ctxWithRoster({ resumeSubAgent });
const result = (await sendMessageTool.handler({ name: "ghost", message: "boo" }, ctx)) as { error: string };
expect(result.error).toMatch(/No teammate named "ghost"/);
expect(resumeSubAgent).not.toHaveBeenCalled();
});
it("falls back to an explicit agentId when no name is given", async () => {
const resumeSubAgent = vi.fn().mockResolvedValue("resumed by id");
const { ctx } = ctxWithRoster({ resumeSubAgent });
const result = (await sendMessageTool.handler({ agentId: "agent-7", message: "continue" }, ctx)) as {
agentId: string;
result: string;
name?: string;
};
expect(resumeSubAgent).toHaveBeenCalledWith("agent-7", "continue");
expect(result.agentId).toBe("agent-7");
expect(result.name).toBeUndefined();
});
it("returns a clean error when neither name nor agentId is provided", async () => {
const resumeSubAgent = vi.fn();
const { ctx } = ctxWithRoster({ resumeSubAgent });
const result = (await sendMessageTool.handler({ message: "to whom?" }, ctx)) as { error: string };
expect(result.error).toMatch(/either/i);
expect(resumeSubAgent).not.toHaveBeenCalled();
});
});
describe("list_teammates tool", () => {
it("returns the roster from ctx.listTeammates", async () => {
const { ctx, roster } = ctxWithRoster();
roster.set("researcher", "agent-1");
roster.set("implementer", "agent-2");
const result = (await listTeammatesTool.handler({}, ctx)) as { teammates: { name: string; agentId: string }[] };
expect(result.teammates).toEqual([
{ name: "researcher", agentId: "agent-1" },
{ name: "implementer", agentId: "agent-2" },
]);
});
it("returns an empty list when no roster is wired (non-session context)", async () => {
const result = (await listTeammatesTool.handler({}, { cwd: "/x" })) as { teammates: unknown[] };
expect(result.teammates).toEqual([]);
});
it("is non-mutating (no confirmation prompt)", () => {
expect(listTeammatesTool.mutating).toBe(false);
});
});
describe("teammate tool schema validation", () => {
it("send_message requires at least one of name/agentId", () => {
expect(() => sendMessageTool.schema.parse({ message: "x" })).toThrow();
});
it("send_message accepts a name only", () => {
expect(() => sendMessageTool.schema.parse({ name: "researcher", message: "x" })).not.toThrow();
});
it("send_message accepts an agentId only", () => {
expect(() => sendMessageTool.schema.parse({ agentId: "agent-1", message: "x" })).not.toThrow();
});
it("list_teammates takes no arguments", () => {
expect(() => listTeammatesTool.schema.parse({})).not.toThrow();
});
});
+19
View File
@@ -0,0 +1,19 @@
import { z } from "zod";
import type { ToolDef } from "./types.js";
const schema = z.object({}).describe("Takes no arguments.");
/** Lists the session's named teammates — sub-agents you spawned via the `agent` tool with a `name`
* that are still resumable. Each entry is `{ name, agentId }`; address one with `send_message` by
* its `name` (or its `agentId`). Returns an empty list when there are none, or in a non-session
* context where no roster is maintained. Non-mutating: it only reads the name → agentId index. */
export const listTeammatesTool: ToolDef<z.infer<typeof schema>> = {
name: "list_teammates",
description:
"List your named teammates — sub-agents you spawned with a `name` via the `agent` tool that are still " +
"resumable. Each has a name and agentId; continue one with send_message by its name. Returns an empty list " +
"when none exist or in a non-session context.",
schema,
mutating: false,
handler: async (_args, ctx) => ({ teammates: ctx.listTeammates?.() ?? [] }),
};
-25
View File
@@ -1,25 +0,0 @@
import { describe, expect, it, vi } from "vitest";
import { todoWriteTool } from "./todoWrite.js";
describe("todo_write", () => {
it("is non-mutating (no confirmation prompt)", () => {
expect(todoWriteTool.mutating).toBe(false);
});
it("forwards the full list to ctx.setTodos and echoes it back", async () => {
const setTodos = vi.fn();
const todos = [
{ content: "read the config", status: "completed" as const },
{ content: "write the fix", status: "in_progress" as const },
{ content: "run tests", status: "pending" as const },
];
const result = await todoWriteTool.handler({ todos }, { cwd: "/tmp", setTodos });
expect(setTodos).toHaveBeenCalledWith(todos);
expect(result).toEqual({ todos });
});
it("doesn't throw when setTodos is absent from the context", async () => {
const todos = [{ content: "a task", status: "pending" as const }];
await expect(todoWriteTool.handler({ todos }, { cwd: "/tmp" })).resolves.toEqual({ todos });
});
});
-28
View File
@@ -1,28 +0,0 @@
import { z } from "zod";
import type { ToolDef } from "./types.js";
const todoItemSchema = z.object({
content: z.string().describe("Short description of the task."),
status: z.enum(["pending", "in_progress", "completed"]),
});
const schema = z.object({
todos: z
.array(todoItemSchema)
.describe("Full checklist (replaces previous list, not a diff)."),
});
export const todoWriteTool: ToolDef<z.infer<typeof schema>> = {
name: "todo_write",
description:
"Show a task checklist for multi-step work (3+ steps). One item 'in_progress' at a time, " +
"mark 'completed' when done. Pass the full list each time. Skip for trivial requests.",
schema,
// Purely informational (like Claude Code's TodoWrite) — never touches the filesystem or asks
// the user anything, so it shouldn't interrupt the flow with a confirmation prompt.
mutating: false,
handler: async (args, ctx) => {
ctx.setTodos?.(args.todos);
return { todos: args.todos };
},
};
+105 -26
View File
@@ -1,5 +1,6 @@
import type { z } from "zod";
import type { TaskStore } from "./task.js";
import type { CronStore } from "../scheduler/cron.js";
export interface SubAgentTask {
/** Short (3-6 word) label shown in the UI while the sub-agent runs. */
@@ -8,46 +9,58 @@ export interface SubAgentTask {
prompt: string;
}
/** Result of one sub-agent in a parallel batch: either its final text, or the error that
* terminated it (timeout, MaxIterationsError, a thrown tool error, etc.). Kept separate from a
* plain string so the parent model can see at a glance which sub-tasks succeeded and which it
* needs to retry or work around — one failed sub-task shouldn't discard the (potentially
* expensive) results of its siblings. */
export interface SubAgentResult {
description: string;
/** The sub-agent's final text answer, or undefined if it failed before producing one. */
result?: string;
/** Present when the sub-agent failed. A timeout, MaxIterationsError, or any other thrown error
* surfaces here rather than rejecting the whole batch. */
error?: string;
}
export interface SubAgentOverrides {
/** Replaces locode's generic sub-agent system prompt entirely — used by plugin-defined agents
* (agents/*.md) that ship their own identity/instructions instead of the generic "delegate a
* task" framing. */
systemPrompt?: string;
/** Prepended (as an addendum) to the generic sub-agent system prompt instead of replacing it, so
* the standard tool-use discipline survives. Used by built-in specialist agent types
* (agentTypes.ts) like explore/code-reviewer. Ignored when `systemPrompt` is also set. */
systemPromptAddendum?: string;
/** Restricts the sub-agent's toolset to tools with these names (unknown names are silently
* ignored); omit to inherit the parent's full toolset minus `agent`/plugin-agent tools. */
toolNames?: string[];
/** Working directory the sub-agent runs against. Set by the parallel-fan-out dispatch path to a
* throwaway git worktree (see utils/worktree.ts) so concurrent sub-agents in one batch can't
* collide on files. Omit (the default for a single/sequential delegation) to run in the parent's
* cwd and persist edits. */
cwd?: string;
}
export interface TodoItem {
content: string;
status: "pending" | "in_progress" | "completed";
/** Result of a sub-agent run: the agent's id (for later continuation via `resumeSubAgent`/the
* `send_message` tool), its final answer text, and whether it's resumable. A sub-agent is resumable
* only when it ran in the shared cwd (sequential single call, or a read-only parallel agent) — a
* worktree-isolated parallel agent's cwd is cleaned up after the batch, so it can't be continued. */
export interface SubAgentResult {
agentId: string;
result: string;
resumable: boolean;
}
export interface ToolContext {
cwd: string;
/** Only present when running inside a session capable of spawning sub-agents (used by the `agent`
* tool). Runs a SINGLE sub-agent and returns its final text; a parallel batch is the `agent`
* tool's own responsibility (it calls this once per task). */
runSubAgent?: (task: SubAgentTask, overrides?: SubAgentOverrides) => Promise<string>;
/** Replaces the session's task checklist (used by the `todo_write` tool). Absent only if a
* future tool context is built without one — every session-backed context provides it. */
setTodos?: (todos: TodoItem[]) => void;
/** The session's structured task store (used by the task_create/list/get/update tools). Absent
* only if a tool context is built without one — every session-backed context provides it. */
/** Only present when running inside a session capable of spawning sub-agents (used by the `agent` tool). */
runSubAgent?: (task: SubAgentTask, overrides?: SubAgentOverrides) => Promise<SubAgentResult>;
/** Continues a previously-spawned resumable sub-agent (one that returned an agentId) with a
* follow-up message, preserving its context. Used by the `send_message` tool. Rejects with a
* clear error if the agentId is unknown or wasn't resumable. */
resumeSubAgent?: (agentId: string, message: string) => Promise<string>;
/** Registers a named teammate: maps `name` → `agentId` on the session's roster so `send_message`
* can address it by name and `list_teammates` can show it. Used by the `agent` tool when the
* model supplies a `name` and the delegation is resumable (shared-cwd). Absent in non-session
* contexts, in which case the name is silently ignored (the agent still runs, just unnamed). */
registerTeammate?: (name: string, agentId: string) => void;
/** Resolves a teammate name to its agentId (or undefined if no such named teammate exists). Used by
* `send_message` to address a teammate by name instead of agentId, and by the `agent` tool to
* pre-flight a name collision before spawning a new teammate. Absent in non-session contexts. */
resolveTeammate?: (name: string) => string | undefined;
/** Lists all named teammates ({ name, agentId }) for the `list_teammates` tool. Absent in
* non-session contexts — the tool returns an empty list then. */
listTeammates?: () => { name: string; agentId: string }[];
/** The session's structured task store, used by the `task_create`/`task_list`/`task_get`/
* `task_update` tools to track multi-step work with dependencies and ownership. Absent only in a
* non-session context (e.g. a hand-built test context) — the task tools return a clear error then. */
taskStore?: TaskStore;
/** Set only while this specific call is a backgroundable tool (currently just `bash`) — the tool
* polls `requested` and, once true, detaches into the background job registry instead of
@@ -58,6 +71,72 @@ export interface ToolContext {
* tree so the command can't keep running (and keep the event loop alive on exit) after the
* parent has already abandoned the turn. Undefined for top-level turns. */
signal?: AbortSignal;
/** Records the previous content of a file edited or written by this tool, so the user can later
* roll back the most recent mutation via the /undo slash command. Only present in session-backed
* contexts that provide a Session object. */
setLastEdit?: (edit: { path: string; previousContent: string }) => void;
/** Refreshes the session's cached user memory after the `memory_write` tool changes memory.md,
* so a later system-prompt rebuild (compaction, /mode or /perm switch) re-folds the new content
* instead of the pre-write snapshot. Only present in session-backed contexts. */
setUserMemory?: (memory: string | null) => void;
/** Surfaces a notice line to the parent UI's scrollback. Used by long-running, headless tools
* (notably `workflow`) to report progress — `log()`/`phase()` inside a workflow script reach the
* user through this. Absent in non-session contexts. */
emitNotice?: (text: string, isError?: boolean) => void;
/** Presents a plan for user approval and, on approval, exits plan mode so the caller can implement.
* Used by the `exit_plan_mode` tool. Returns whether the user approved; on rejection plan mode
* stays active so the model can refine and re-present. Absent outside plan mode (so the tool fails
* cleanly with "only available in plan mode" if the model calls it at the wrong time). */
exitPlanMode?: (plan: string) => Promise<{ approved: boolean }>;
/** Asks the user a structured multiple-choice question (or a short sequence of them) when the model
* is blocked on a decision only the user can make. Used by the `ask_user_question` tool. Resolves
* with the selected option label(s) per question; an empty selection means the user skipped
* (treated as "Other"/custom input in the UI, surfaced back as the typed text). Absent in
* non-session contexts (headless sub-agents can't prompt). */
askQuestion?: (questions: AskQuestionSpec[]) => Promise<AskQuestionAnswer[]>;
/** The session's cron/wakeup scheduler, used by the `cron_create`/`cron_list`/`cron_delete`/
* `schedule_wakeup` tools. Absent in a non-session context (e.g. a hand-built test context) — the
* scheduling tools return a clear error then. */
cronStore?: CronStore;
/** Switches the session's working directory to `newCwd` (used by `enter_worktree`/`exit_worktree`)
* and notifies the UI so its live cwd display, /undo resolution, git-info, and @mention handling
* follow the switch. Absent in non-session contexts — the worktree tools fail cleanly then. */
setCwd?: (newCwd: string) => void;
/** Reads the active interactive worktree tracking ({ dir, branch, originalCwd }) or undefined when
* not in a worktree session. Used by `enter_worktree` (refuse re-entry) and `exit_worktree`
* (restore cwd + remove). */
getWorktree?: () => { dir: string; branch: string; originalCwd: string } | undefined;
/** Sets or clears the active worktree tracking on the session (and notifies the App so a /model
* switch can re-attach it). `enter_worktree` sets it; `exit_worktree` passes undefined to clear. */
setWorktree?: (worktree: { dir: string; branch: string; originalCwd: string } | undefined) => void;
}
/** One selectable option in an {@link AskQuestionSpec}. `description` is shown dimmed under the label
* to explain a trade-off or implication, so the user can compare options at a glance. */
export interface AskQuestionOption {
label: string;
description?: string;
}
/** A single question the model asks the user via the `ask_user_question` tool. `header` is a short
* (≤ ~12 char) chip shown beside the question for scanability; `options` is 2-4 choices; when
* `multiSelect` is true the user may pick several (otherwise exactly one). The UI also offers an
* implicit "Other" path so the user can type a custom answer not in the list. */
export interface AskQuestionSpec {
question: string;
header: string;
options: AskQuestionOption[];
multiSelect?: boolean;
}
/** The user's answer to one {@link AskQuestionSpec}: the labels of the selected option(s), in the
* order the model listed them. An empty array with `custom` set means the user typed a freeform
* answer via "Other" instead of picking a listed option. */
export interface AskQuestionAnswer {
question: string;
selected: string[];
/** A freeform answer the user typed via "Other" instead of picking a listed option. */
custom?: string;
}
export interface ToolDef<T = any> {
+1 -2
View File
@@ -19,8 +19,7 @@ function isBinaryContentType(contentType: string): boolean {
export const webFetchTool: ToolDef<z.infer<typeof schema>> = {
name: "web_fetch",
description:
"Fetch a URL and return readable text (HTML tags/scripts/styles stripped). Use for specific pages found via web_search — " +
"e.g. to read a doc page or blog post in full when the search snippet was not enough.",
"Fetch a URL and return readable text (HTML tags/scripts/styles stripped). Use for specific pages found via web_search.",
schema,
mutating: false,
handler: async ({ url }) => {
+1 -2
View File
@@ -45,8 +45,7 @@ function parseResults(html: string, limit: number): SearchResult[] {
export const webSearchTool: ToolDef<z.infer<typeof schema>> = {
name: "web_search",
description:
"Search the web via DuckDuckGo. Returns title, url, snippet. Use for info not in the local codebase — " +
"e.g. an unfamiliar API, library docs, or an error message. Follow up with web_fetch on a specific result for full page text.",
"Search the web via DuckDuckGo. Returns title, url, snippet. Use for info not in the local codebase.",
schema,
mutating: false,
handler: async ({ query, max_results }) => {
+377
View File
@@ -0,0 +1,377 @@
import { describe, expect, it, vi, beforeEach, afterEach } from "vitest";
import { execa } from "execa";
import { existsSync, mkdtempSync, rmSync, writeFileSync } from "node:fs";
import { tmpdir } from "node:os";
import path from "node:path";
import { stripExports, extractJson, validateAgainstSchema, runWorkflow } from "./workflow.js";
import type { SubAgentResult } from "./types.js";
describe("stripExports", () => {
it("strips leading export keywords so module-style scripts run as plain scripts", () => {
expect(stripExports("export const meta = { name: 'x' };\nconst y = 1;")).toBe("const meta = { name: 'x' };\nconst y = 1;");
expect(stripExports("export function f() {}\nexport default 1;")).toBe("function f() {}\n1;");
// Non-export lines are untouched.
expect(stripExports("const a = 1;\n// export b\nconst c = 3;")).toBe("const a = 1;\n// export b\nconst c = 3;");
});
});
describe("extractJson", () => {
it("parses plain JSON", () => {
expect(extractJson('{"a":1}')).toEqual({ a: 1 });
expect(extractJson("[1,2,3]")).toEqual([1, 2, 3]);
});
it("extracts JSON from a ```json fence", () => {
expect(extractJson("Here you go:\n```json\n{\"a\":1}\n```\nthanks")).toEqual({ a: 1 });
});
it("extracts JSON from a fence with an arbitrary language tag (```javascript)", () => {
expect(extractJson("```javascript\n{\"a\":1}\n```")).toEqual({ a: 1 });
expect(extractJson("```ts\n{\"a\":1}\n```")).toEqual({ a: 1 });
});
it("strips trailing commas before } (a common local-model artifact)", () => {
expect(extractJson('{"a":1,"b":2,}')).toEqual({ a: 1, b: 2 });
expect(extractJson('```json\n{\n "a": 1,\n "b": 2,\n}\n```')).toEqual({ a: 1, b: 2 });
});
it("strips trailing commas before ] in arrays", () => {
expect(extractJson("[1,2,3,]")).toEqual([1, 2, 3]);
});
it("extracts the first JSON object from surrounding prose", () => {
expect(extractJson('The answer is {"a":1,"b":2} as shown.')).toEqual({ a: 1, b: 2 });
});
it("throws when no JSON is present", () => {
expect(() => extractJson("no json here")).toThrow(/valid JSON/);
});
});
describe("validateAgainstSchema", () => {
const sch = { type: "object" as const, required: ["a", "b"], properties: { a: { type: "number" }, b: { type: "string" } } };
it("passes a valid object", () => {
expect(() => validateAgainstSchema({ a: 1, b: "x" }, sch)).not.toThrow();
});
it("rejects a missing required field", () => {
expect(() => validateAgainstSchema({ a: 1 }, sch)).toThrow(/missing required field: b/);
});
it("rejects a wrong property type", () => {
expect(() => validateAgainstSchema({ a: "notnum", b: "x" }, sch)).toThrow(/field a: expected number/);
});
it("rejects a non-object when object is required", () => {
expect(() => validateAgainstSchema([1, 2], sch)).toThrow(/expected a JSON object/);
});
it("validates array item types via items", () => {
const arrSch = { type: "array", items: { type: "string" } };
expect(() => validateAgainstSchema(["a", "b"], arrSch)).not.toThrow();
expect(() => validateAgainstSchema(["a", 2], arrSch)).toThrow(/item\[1\]: expected string/);
});
it("validates enum membership", () => {
const enumSch = { type: "string", enum: ["low", "med", "high"] };
expect(() => validateAgainstSchema("med", enumSch)).not.toThrow();
expect(() => validateAgainstSchema("nope", enumSch)).toThrow(/not in allowed enum/);
});
it("validates boolean and integer property types", () => {
const s = { type: "object", properties: { ok: { type: "boolean" }, n: { type: "integer" } } };
expect(() => validateAgainstSchema({ ok: true, n: 3 }, s)).not.toThrow();
expect(() => validateAgainstSchema({ ok: "yes", n: 3 }, s)).toThrow(/field ok: expected boolean/);
// integer must reject a non-integer number.
expect(() => validateAgainstSchema({ ok: true, n: 1.5 }, s)).toThrow(/field n: expected integer/);
});
it("validates nested object properties", () => {
const s = {
type: "object",
properties: { outer: { type: "object", required: ["inner"], properties: { inner: { type: "number" } } } },
};
expect(() => validateAgainstSchema({ outer: { inner: 5 } }, s)).not.toThrow();
expect(() => validateAgainstSchema({ outer: {} }, s)).toThrow(/missing required field: inner/);
expect(() => validateAgainstSchema({ outer: { inner: "x" } }, s)).toThrow(/field outer.inner: expected number/);
});
it("rejects additional properties when additionalProperties is false", () => {
const s = { type: "object", properties: { a: { type: "number" } }, additionalProperties: false };
expect(() => validateAgainstSchema({ a: 1 }, s)).not.toThrow();
expect(() => validateAgainstSchema({ a: 1, extra: 2 }, s)).toThrow(/unexpected additional property: extra/);
});
});
function mockSubAgent(result: string): { runSubAgent: ReturnType<typeof vi.fn> } {
return { runSubAgent: vi.fn(async (): Promise<SubAgentResult> => ({ agentId: "id", result, resumable: false })) };
}
describe("runWorkflow", () => {
it("runs a script that returns a value, calling agent() through ctx.runSubAgent", async () => {
const ctx = mockSubAgent("the answer");
const out = await runWorkflow(
`const r = await agent("do something", { label: "worker" });\nreturn r;`,
undefined,
ctx,
);
expect(out).toBe("the answer");
expect(ctx.runSubAgent).toHaveBeenCalledTimes(1);
expect(ctx.runSubAgent.mock.calls[0]![0].description).toBe("worker");
});
it("exposes args to the script", async () => {
const ctx = mockSubAgent("ok");
const out = await runWorkflow(
`const items = args;\nconst r = await agent("process " + items.join(","));\nreturn r;`,
["a", "b", "c"],
ctx,
);
expect(out).toBe("ok");
expect(ctx.runSubAgent.mock.calls[0]![0].prompt).toBe("process a,b,c");
});
it("parallel() runs thunks concurrently and turns failures into null", async () => {
const ctx = mockSubAgent("ok");
const out = (await runWorkflow(
`const results = await parallel([
() => agent("task A").then(r => r + "!"),
() => Promise.reject(new Error("boom")),
() => agent("task C"),
]);
return results;`,
undefined,
ctx,
)) as (string | null)[];
expect(out).toHaveLength(3);
expect(out[0]).toBe("ok!");
expect(out[1]).toBeNull();
expect(out[2]).toBe("ok");
});
it("pipeline() runs each item through all stages, no barrier between stages", async () => {
// A stage-2 item can finish before a slow stage-1 item — but for determinism in this test we
// just assert each item passes through both stages in order and results land in input order.
const ctx = mockSubAgent("ok");
const out = (await runWorkflow(
`const out = await pipeline(
["a", "b", "c"],
async (item) => item + "1",
async (item) => item + "2",
);
return out;`,
undefined,
ctx,
)) as string[];
expect(out).toEqual(["a12", "b12", "c12"]);
});
it("a pipeline stage that throws drops just that item to null", async () => {
const ctx = mockSubAgent("ok");
const out = (await runWorkflow(
`const out = await pipeline(
["a", "b", "c"],
async (item) => { if (item === "b") throw new Error("nope"); return item + "1"; },
async (item) => item + "2",
);
return out;`,
undefined,
ctx,
)) as (string | null)[];
expect(out).toEqual(["a12", null, "c12"]);
});
it("agent() with a schema returns a parsed, validated object (retrying once on bad JSON)", async () => {
// First call returns non-JSON; the retry (nudge prompt) returns valid JSON matching the schema.
const calls: string[] = [];
const ctx = {
runSubAgent: vi.fn(async (task: { prompt: string }): Promise<SubAgentResult> => {
calls.push(task.prompt);
if (task.prompt.includes("not valid JSON")) {
return { agentId: "id", result: '{"answer": 42}', resumable: false };
}
return { agentId: "id", result: "I think the answer is 42.", resumable: false };
}),
};
const out = await runWorkflow(
`const r = await agent("what is the answer", {
schema: { type: "object", required: ["answer"], properties: { answer: { type: "number" } } },
});
return r;`,
undefined,
ctx,
);
expect(out).toEqual({ answer: 42 });
expect(ctx.runSubAgent).toHaveBeenCalledTimes(2);
// The retry carried a nudge.
expect(calls[1]).toContain("not valid JSON");
});
it("agent() with a schema throws if the retry still doesn't validate", async () => {
const ctx = mockSubAgent("still not json at all");
await expect(
runWorkflow(
`await agent("x", { schema: { type: "object", required: ["answer"] } });`,
undefined,
ctx,
),
).rejects.toThrow(/valid JSON/);
});
it("agent() retry nudge includes the specific validation error from the first attempt", async () => {
const calls: string[] = [];
const ctx = {
runSubAgent: vi.fn(async (task: { prompt: string }): Promise<SubAgentResult> => {
calls.push(task.prompt);
if (calls.length === 2) {
// Retry returns valid JSON.
return { agentId: "id", result: '{"answer": 42}', resumable: false };
}
// First call returns JSON with a wrong type (answer is a string, schema wants number).
return { agentId: "id", result: '{"answer": "forty-two"}', resumable: false };
}),
};
const out = await runWorkflow(
`const r = await agent("what is the answer", {
schema: { type: "object", required: ["answer"], properties: { answer: { type: "number" } } },
});
return r;`,
undefined,
ctx,
);
expect(out).toEqual({ answer: 42 });
// The retry nudge quoted the first attempt's validation error (wrong type for `answer`).
expect(calls[1]).toContain("not valid JSON");
expect(calls[1]).toMatch(/answer.*expected number|expected number.*answer/);
});
it("caps concurrency so a fan-out doesn't exceed the limit", async () => {
process.env.LOCODE_WORKFLOW_CONCURRENCY = "2";
try {
let active = 0;
let maxActive = 0;
const ctx = {
runSubAgent: vi.fn(async (): Promise<SubAgentResult> => {
active++;
maxActive = Math.max(maxActive, active);
await new Promise((r) => setTimeout(r, 30));
active--;
return { agentId: "id", result: "ok", resumable: false };
}),
};
await runWorkflow(
`await parallel([
() => agent("1"), () => agent("2"), () => agent("3"),
() => agent("4"), () => agent("5"), () => agent("6"),
]);`,
undefined,
ctx,
);
expect(ctx.runSubAgent).toHaveBeenCalledTimes(6);
expect(maxActive).toBeLessThanOrEqual(2);
} finally {
delete process.env.LOCODE_WORKFLOW_CONCURRENCY;
}
});
it("log() and phase() surface notices through ctx.emitNotice", async () => {
const notices: { text: string; isError?: boolean }[] = [];
const ctx = {
runSubAgent: vi.fn(async (): Promise<SubAgentResult> => ({ agentId: "id", result: "ok", resumable: false })),
emitNotice: (text: string, isError?: boolean) => notices.push({ text, isError }),
};
await runWorkflow(
`phase("Review");\nlog("halfway");\nconst r = await agent("x");\nlog("done");\nreturn r;`,
undefined,
ctx,
);
expect(notices).toContainEqual({ text: "▶ Review", isError: undefined });
expect(notices.map((n) => n.text)).toEqual(["▶ Review", "halfway", "done"]);
});
it("a thrown error fails the workflow", async () => {
const ctx = mockSubAgent("ok");
await expect(runWorkflow(`throw new Error("script broke");`, undefined, ctx)).rejects.toThrow("script broke");
});
it("nested workflow() is rejected", async () => {
const ctx = mockSubAgent("ok");
await expect(runWorkflow(`await workflow();`, undefined, ctx)).rejects.toThrow(/cannot be nested/);
});
it("resolves agentType to the specialist's toolset+addendum via overrides", async () => {
const ctx = {
runSubAgent: vi.fn(async (_task: unknown, overrides?: unknown): Promise<SubAgentResult> => {
// Stash the overrides for assertion; return a plain result.
(ctx as any).__overrides = overrides;
return { agentId: "id", result: "ok", resumable: false };
}),
};
await runWorkflow(`await agent("explore the repo", { agentType: "explore" });`, undefined, ctx);
const overrides = (ctx as any).__overrides as { toolNames?: string[]; systemPromptAddendum?: string };
expect(overrides?.toolNames).toContain("read_file");
expect(overrides?.systemPromptAddendum).toContain("Explore agent");
});
});
describe("runWorkflow worktree isolation", () => {
let repo: string;
beforeEach(() => {
repo = mkdtempSync(path.join(tmpdir(), "locode-wf-wt-test-"));
});
afterEach(async () => {
await execa("git", ["worktree", "prune"], { cwd: repo, reject: false }).catch(() => {});
try {
rmSync(repo, { recursive: true, force: true });
} catch {
/* leave for the OS temp sweep */
}
});
async function gitInit(r: string): Promise<void> {
await execa("git", ["init", "-q"], { cwd: r });
await execa("git", ["config", "user.email", "t@t"], { cwd: r });
await execa("git", ["config", "user.name", "t"], { cwd: r });
writeFileSync(path.join(r, "README.md"), "hello\n");
await execa("git", ["add", "."], { cwd: r });
await execa("git", ["commit", "-q", "-m", "init"], { cwd: r });
}
it("agent({isolation:'worktree'}) runs in a real worktree cwd distinct from the repo, cleaned up after", async () => {
await gitInit(repo);
let agentCwd: string | undefined;
const ctx = {
cwd: repo,
runSubAgent: vi.fn(async (_task: unknown, overrides?: { cwd?: string }): Promise<SubAgentResult> => {
agentCwd = overrides?.cwd;
return { agentId: "id", result: "ok", resumable: false };
}),
};
await runWorkflow(`await agent("edit things", { isolation: "worktree" });`, undefined, ctx);
// The agent ran in a throwaway worktree, not the shared repo cwd.
expect(agentCwd).toBeDefined();
expect(agentCwd).not.toBe(repo);
expect(existsSync(agentCwd!)).toBe(false); // cleaned up after the agent finished
});
it("agent({isolation:'worktree'}) on a non-git cwd falls back to the shared cwd (no isolation)", async () => {
// repo exists but has no .git (gitInit not called) → createWorktree returns cwd:undefined.
let agentCwd: string | undefined;
const ctx = {
cwd: repo,
runSubAgent: vi.fn(async (_task: unknown, overrides?: { cwd?: string }): Promise<SubAgentResult> => {
agentCwd = overrides?.cwd;
return { agentId: "id", result: "ok", resumable: false };
}),
};
await runWorkflow(`await agent("edit things", { isolation: "worktree" });`, undefined, ctx);
// No git → no worktree → the agent runs in the shared cwd (undefined override = inherit parent).
expect(agentCwd).toBeUndefined();
});
it("agent({isolation:'worktree'}) still cleans up the worktree even when the agent throws", async () => {
await gitInit(repo);
let agentCwd: string | undefined;
const ctx = {
cwd: repo,
runSubAgent: vi.fn(async (_task: unknown, overrides?: { cwd?: string }): Promise<SubAgentResult> => {
agentCwd = overrides?.cwd;
throw new Error("agent exploded");
}),
};
await expect(
runWorkflow(`await agent("edit things", { isolation: "worktree" });`, undefined, ctx),
).rejects.toThrow("agent exploded");
expect(agentCwd).toBeDefined();
expect(existsSync(agentCwd!)).toBe(false); // the finally cleaned up despite the throw
});
});
+358
View File
@@ -0,0 +1,358 @@
import vm from "node:vm";
import { z } from "zod";
import { getAgentType } from "./agentTypes.js";
import type { SubAgentOverrides, SubAgentResult, ToolDef } from "./types.js";
import { createWorktree, type WorktreeIsolation } from "../utils/worktree.js";
/** Default cap on concurrent in-flight sub-agents inside a workflow. Local backends (Ollama, LM
* Studio) typically serve a single model and can't truly parallelize many concurrent request
* streams — an unbounded fan-out would queue a large burst and risk OOM/timeout. The cap keeps the
* burst bounded; the model server serializes what it can't concurrentize. Override via
* LOCODE_WORKFLOW_CONCURRENCY. */
function workflowConcurrency(): number {
const raw = Number(process.env.LOCODE_WORKFLOW_CONCURRENCY);
if (Number.isFinite(raw) && raw >= 1) return Math.floor(raw);
return 4;
}
/** Wall-clock backstop for the whole workflow (the script itself is fast between agent calls; the
* real time is in the agents, each already bounded by SUBAGENT_TIMEOUT_MS). This catches a script
* that loops forever spawning agents. Override via LOCODE_WORKFLOW_TIMEOUT_MS. */
function workflowTimeoutMs(): number {
const raw = Number(process.env.LOCODE_WORKFLOW_TIMEOUT_MS);
if (Number.isFinite(raw) && raw >= 1000) return Math.floor(raw);
return 15 * 60 * 1000;
}
/** A JSON-Schema-ish spec for forcing structured output from a sub-agent. Minimal validation only
* (required keys, property/item types, enums, nested objects) — locode can't force a native tool
* call for structured output the way Claude Code can, so this is best-effort: the sub-agent is told
* to return JSON, we parse it, and retry once if it doesn't validate. */
interface JsonSchemaProperty {
type?: string;
enum?: unknown[];
items?: JsonSchemaProperty;
properties?: Record<string, JsonSchemaProperty>;
required?: string[];
additionalProperties?: boolean;
}
interface JsonSchema extends JsonSchemaProperty {}
const schema = z.object({
script: z
.string()
.describe(
"A self-contained JavaScript workflow script. It may begin with `const meta = { name, description, phases }` " +
"(metadata only). Use the provided globals to orchestrate: agent(prompt, opts?) runs a sub-agent and returns its " +
"answer (or a validated object when opts.schema is given); parallel([() => ..., () => ...]) runs thunks concurrently " +
"(failures become null); pipeline(items, stage1, stage2, ...) runs each item through all stages with no barrier between " +
"stages; phase(title) marks a progress group; log(message) surfaces a progress line to the user. Return a value to make " +
"it the workflow's result. No filesystem or Node APIs; no Date.now/Math.random. A thrown error fails the workflow.",
),
args: z
.any()
.optional()
.describe("Free-form value exposed to the script as the global `args` (pass arrays/objects, not a stringified string)."),
});
/** Strips ES-module `export` keywords so a script written in Claude-Code's `export const meta` style
* runs as plain script source inside the vm. Handles `export const/function/default` at line starts. */
function stripExports(src: string): string {
return src.replace(/^export\s+(default\s+)?/gm, "");
}
/** Extracts the first balanced JSON value (object or array) from text that may have surrounding
* prose/code fences — local models often wrap JSON in ```json … ``` (or ```javascript, ```ts, …)
* and add commentary. Tolerates trailing commas, which local models emit frequently even though
* they're invalid JSON. */
function extractJson(text: string): unknown {
const trimmed = text.trim();
// Strip a surrounding ```lang ... ``` fence with any language tag (json, javascript, ts, …).
// Local models label fences with whatever language they think the content is, so accept any tag.
const fenced = trimmed.match(/```[a-zA-Z0-9+#]*\s*([\s\S]*?)```/);
const candidate = (fenced ? fenced[1]! : trimmed).trim();
// Local models frequently emit trailing commas before } or ] (invalid JSON). Strip them.
const sanitized = stripTrailingCommas(candidate);
try {
return JSON.parse(sanitized);
} catch {
// Fall back to the first {...} or [...] span, also comma-sanitized.
const span = candidate.match(/(\{[\s\S]*\}|\[[\s\S]*\])/);
if (span) {
try {
return JSON.parse(stripTrailingCommas(span[1]!));
} catch {
/* fall through */
}
}
}
throw new Error("Sub-agent did not return valid JSON for structured output.");
}
/** Removes trailing commas that immediately precede a closing } or ] — a common local-model
* artifact that makes otherwise-valid JSON unparseable. */
function stripTrailingCommas(s: string): string {
return s.replace(/,\s*([}\]])/g, "$1");
}
/** Minimal JSON-Schema validation: checks the top-level shape (object/array), required keys,
* property types, array item types, enum membership, and nested object properties. Intentionally
* not a full validator — just enough to catch an obviously-wrong shape and trigger the one retry
* in agentFn. Error messages are kept stable (e.g. "missing required field: x") so the retry nudge
* can quote them back to the sub-agent. */
function validateAgainstSchema(value: unknown, schema: JsonSchema): void {
// Top-level shape checks keep the stable, friendly messages callers (and the retry nudge) rely on.
if (schema.type === "object") {
if (typeof value !== "object" || value === null || Array.isArray(value)) {
throw new Error("expected a JSON object");
}
} else if (schema.type === "array") {
if (!Array.isArray(value)) throw new Error("expected a JSON array");
}
// Deeper checks (required keys, property types, items, enum, nested objects) share validateProperty.
validateProperty(value, schema, "");
}
/** Validates a value against a single property/schema descriptor. `path` is the dotted location used
* in error messages ("" at the top level, "field a" / "field a.b" / "item[2]" beneath it). */
function validateProperty(value: unknown, prop: JsonSchemaProperty, path: string): void {
if (prop.type) checkType(value, prop.type, path);
if (prop.enum && !prop.enum.includes(value)) {
throw new Error(`${path || "value"}: value not in allowed enum`);
}
if (prop.type === "object" && typeof value === "object" && value !== null && !Array.isArray(value)) {
const obj = value as Record<string, unknown>;
for (const key of prop.required ?? []) {
if (!(key in obj)) throw new Error(`${path ? path + ": " : ""}missing required field: ${key}`);
}
if (prop.properties) {
for (const [key, sub] of Object.entries(prop.properties)) {
if (key in obj) validateProperty(obj[key], sub, path ? `${path}.${key}` : `field ${key}`);
}
}
if (prop.additionalProperties === false && prop.properties) {
for (const key of Object.keys(obj)) {
if (!(key in prop.properties)) throw new Error(`${path ? path + ": " : ""}unexpected additional property: ${key}`);
}
}
}
if (prop.type === "array" && Array.isArray(value) && prop.items) {
value.forEach((item, i) => validateProperty(item, prop.items!, path ? `${path}[${i}]` : `item[${i}]`));
}
}
function checkType(v: unknown, type: string, field: string): void {
const jsType = Array.isArray(v) ? "array" : v === null ? "null" : typeof v;
// integer: must be a number AND a whole number. number: any number (incl. floats).
if (type === "integer") {
if (typeof v !== "number" || !Number.isInteger(v)) throw new Error(`field ${field}: expected integer`);
return;
}
if (type === "number") {
if (typeof v !== "number") throw new Error(`field ${field}: expected number`);
return;
}
if (jsType !== type) throw new Error(`field ${field}: expected ${type}, got ${jsType}`);
}
interface AgentOpts {
label?: string;
schema?: JsonSchema;
toolNames?: string[];
systemPromptAddendum?: string;
cwd?: string;
agentType?: string;
/** Opt the agent into running inside a throwaway git worktree (detached at HEAD), so its file
* writes can't collide with the main repo or with other concurrently-running agents. Use this for
* writable agents (general-purpose, debugger, test-writer, or plugin agents) that you fan out in
* parallel — their edits are discarded and only the returned answer matters. If the parent cwd
* isn't a git repo, isolation is best-effort and the agent runs in the shared cwd. Don't use it
* for read-only agents (explore, code-reviewer, planner) — they should see current uncommitted
* state, so run them in the shared cwd. */
isolation?: "worktree";
}
/** Builds the SubAgentOverrides for an agent() call, resolving agentType to its toolset+addendum
* unless the script explicitly supplies either. */
function buildOverrides(opts: AgentOpts): SubAgentOverrides | undefined {
if (opts.toolNames || opts.systemPromptAddendum) {
return { toolNames: opts.toolNames, systemPromptAddendum: opts.systemPromptAddendum, cwd: opts.cwd };
}
if (opts.agentType && opts.agentType !== "general-purpose") {
const spec = getAgentType(opts.agentType);
return { toolNames: spec.toolNames, systemPromptAddendum: spec.systemPromptAddendum, cwd: opts.cwd };
}
return opts.cwd ? { cwd: opts.cwd } : undefined;
}
/** Runs a workflow script with the orchestration globals in scope. Returns whatever the script
* returns (or throws). agent/parallel/pipeline run sub-agents via ctx.runSubAgent, bounded by the
* concurrency cap and the overall timeout. */
async function runWorkflow(
script: string,
args: unknown,
ctx: {
cwd?: string;
runSubAgent?: (task: { description: string; prompt: string }, overrides?: SubAgentOverrides) => Promise<SubAgentResult>;
emitNotice?: (text: string, isError?: boolean) => void;
},
): Promise<unknown> {
if (!ctx.runSubAgent) throw new Error("Sub-agents are not available in this context.");
const concurrency = workflowConcurrency();
// A bounded concurrency runner: at most `limit` thunks in flight at once. Used by parallel() and
// pipeline() so a fan-out over many items doesn't dump a huge burst onto a local model server.
async function runBounded<I, O>(limit: number, items: I[], fn: (item: I, index: number) => Promise<O>): Promise<O[]> {
const results: O[] = new Array(items.length);
let next = 0;
async function worker(): Promise<void> {
while (true) {
const i = next++;
if (i >= items.length) return;
results[i] = await fn(items[i]!, i);
}
}
const workers = Array.from({ length: Math.min(limit, items.length) }, () => worker());
await Promise.all(workers);
return results;
}
async function agentFn(prompt: string, opts: AgentOpts = {}): Promise<unknown> {
const description = opts.label ?? "workflow-agent";
// Opt-in worktree isolation for writable parallel agents (see AgentOpts.isolation). Created here
// and cleaned up in the finally below — runSubAgentTurn treats a cwd != parent.cwd as "isolated"
// (edits discarded, not resumable) and appends the discard-notice itself. Best-effort: if the
// parent cwd isn't a git repo, createWorktree returns cwd:undefined and we run in the shared cwd.
let isolation: WorktreeIsolation | undefined;
if (opts.isolation === "worktree") isolation = await createWorktree(ctx.cwd ?? ".");
const worktreeCwd = isolation?.cwd;
try {
const overrides = buildOverrides({ ...opts, cwd: opts.cwd ?? worktreeCwd });
const structuredAddendum = opts.schema
? `\n\nReturn ONLY a JSON ${opts.schema.type ?? "object"} matching this schema (no prose, no code fences):\n${JSON.stringify(opts.schema)}`
: undefined;
const callOverrides = structuredAddendum
? {
...overrides,
systemPromptAddendum: overrides?.systemPromptAddendum
? `${overrides.systemPromptAddendum}\n${structuredAddendum}`
: structuredAddendum,
}
: overrides;
const r1 = await ctx.runSubAgent!({ description, prompt }, callOverrides);
if (!opts.schema) return r1.result;
// Structured output: parse + validate, retry once with a targeted nudge if it fails. The nudge
// includes the actual validation error so the sub-agent can correct the specific defect.
let parsed: unknown;
try {
parsed = extractJson(r1.result);
validateAgainstSchema(parsed, opts.schema);
return parsed;
} catch (err) {
const why = (err as Error).message;
const required = opts.schema.required ?? Object.keys(opts.schema.properties ?? {});
const nudge = `${prompt}\n\nYour previous response was not valid JSON matching the schema (${why}). Return ONLY a JSON object with these fields: ${JSON.stringify(required)}. No prose, no code fences, no trailing commas.`;
const r2 = await ctx.runSubAgent!({ description, prompt: nudge }, callOverrides);
parsed = extractJson(r2.result);
validateAgainstSchema(parsed, opts.schema);
return parsed;
}
} finally {
await isolation?.cleanup();
}
}
function parallelFn<T>(thunks: Array<() => Promise<T>>): Promise<Array<T | null>> {
if (!Array.isArray(thunks)) throw new Error("parallel() expects an array of thunks");
return runBounded<() => Promise<T>, T | null>(concurrency, thunks, async (thunk) => {
try {
return await thunk();
} catch (err) {
// A thunk that throws (or whose agent errors) resolves to null — the call itself never
// rejects, so one failing branch doesn't abort the whole parallel batch.
ctx.emitNotice?.(`parallel branch failed: ${(err as Error).message ?? String(err)}`, true);
return null;
}
});
}
function pipelineFn<T, R>(items: T[], ...stages: Array<(prev: unknown, original: T, index: number) => Promise<unknown>>): Promise<Array<R | null>> {
if (!Array.isArray(items)) throw new Error("pipeline() expects an array of items");
if (stages.length === 0) throw new Error("pipeline() needs at least one stage");
return runBounded<T, R | null>(concurrency, items, async (item, index) => {
try {
let val: unknown = item;
for (const stage of stages) {
val = await stage(val, item, index);
}
return val as R;
} catch (err) {
// A stage that throws drops just this item to null (skipping its remaining stages).
ctx.emitNotice?.(`pipeline item ${index} failed: ${(err as Error).message ?? String(err)}`, true);
return null;
}
});
}
function phaseFn(title: string): void {
ctx.emitNotice?.(`▶ ${title}`);
}
function logFn(message: string): void {
ctx.emitNotice?.(String(message));
}
const sandbox = {
agent: agentFn,
parallel: parallelFn,
pipeline: pipelineFn,
phase: phaseFn,
log: logFn,
args,
// Nested workflows are one level only (matches Claude Code); there's no child context to run in.
workflow: () => {
throw new Error("workflow() cannot be nested.");
},
};
const wrapped = `(async () => {\n${stripExports(script)}\n})()`;
const context = vm.createContext(sandbox);
let timeoutId: ReturnType<typeof setTimeout>;
const timeout = new Promise<never>((_, reject) => {
timeoutId = setTimeout(
() => reject(new Error(`Workflow timed out after ${Math.round(workflowTimeoutMs() / 1000)}s.`)),
workflowTimeoutMs(),
);
});
try {
const promise = vm.runInContext(wrapped, context, { filename: "workflow.js" }) as Promise<unknown>;
return await Promise.race([promise, timeout]);
} finally {
clearTimeout(timeoutId!);
}
}
export const workflowTool: ToolDef<z.infer<typeof schema>> = {
name: "workflow",
description:
"Run a multi-agent workflow from a self-contained JavaScript script that deterministically orchestrates sub-agents — " +
"for being comprehensive (decompose and cover in parallel), confident (independent perspectives + adversarial checks " +
"before committing), or scaling work one context can't hold (migrations, audits). Use it only when the user asks for " +
"multi-agent orchestration (e.g. 'use a workflow', 'fan out agents'); a single agent() call inside is just a delegation. " +
"The script globals are agent, parallel, pipeline, phase, log, args. agent(prompt, opts) accepts opts.isolation: 'worktree' " +
"to run a writable agent in a throwaway git worktree so parallel writers don't collide (edits discarded; only the answer " +
"matters) — use it for parallel general-purpose/debugger/test-writer agents, not for read-only explore/code-reviewer ones. " +
"Runs headless; the return value is the result.",
schema,
mutating: false,
handler: async (args, ctx) => {
const result = await runWorkflow(args.script, args.args, ctx);
return { result };
},
};
// Exported for tests.
export { runWorkflow, extractJson, validateAgainstSchema, stripExports };
+174
View File
@@ -0,0 +1,174 @@
import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
import { execa } from "execa";
import { existsSync, mkdtempSync, rmSync, writeFileSync } from "node:fs";
import path from "node:path";
import os from "node:os";
import { enterWorktreeTool, exitWorktreeTool } from "./worktreeSession.js";
import { removeInteractiveWorktree } from "../utils/worktree.js";
import type { ToolContext } from "./types.js";
// These tests shell out to real git for the create/remove paths, and use a mocked ctx to observe
// the setCwd/setWorktree/getWorktree wiring without a full Session.
async function gitInit(repo: string): Promise<void> {
await execa("git", ["init", "-q"], { cwd: repo });
await execa("git", ["config", "user.email", "t@t"], { cwd: repo });
await execa("git", ["config", "user.name", "t"], { cwd: repo });
writeFileSync(path.join(repo, "README.md"), "hello\n");
await execa("git", ["add", "."], { cwd: repo });
await execa("git", ["commit", "-q", "-m", "init"], { cwd: repo });
}
// A mock ctx that records setCwd calls and tracks the worktree over a shared object, mirroring how
// gateAndRun wires these over the Session. `cwd` is the current working dir (the repo, or the
// worktree once switched).
function ctxFor(cwd: string): { ctx: ToolContext; state: { cwd: string; worktree?: { dir: string; branch: string; originalCwd: string } } } {
const state: { cwd: string; worktree?: { dir: string; branch: string; originalCwd: string } } = { cwd };
const ctx: ToolContext = {
cwd,
setCwd: (newCwd) => {
state.cwd = newCwd;
},
getWorktree: () => state.worktree,
setWorktree: (wt) => {
state.worktree = wt;
},
};
return { ctx, state };
}
describe("enter_worktree / exit_worktree tool wiring", () => {
let repo: string;
beforeEach(() => {
repo = mkdtempSync(path.join(os.tmpdir(), "locode-wt-tool-test-"));
});
afterEach(async () => {
if (existsSync(repo)) {
await execa("git", ["worktree", "prune"], { cwd: repo, reject: false }).catch(() => {});
try {
rmSync(repo, { recursive: true, force: true });
} catch {
/* leave for the OS temp sweep */
}
}
});
it("enter_worktree throws when the ctx doesn't support worktree sessions", async () => {
const ctx: ToolContext = { cwd: repo };
await expect(enterWorktreeTool.handler({}, ctx)).rejects.toThrow(/not available/i);
});
it("enter_worktree refuses when already in a worktree session", async () => {
await gitInit(repo);
const { ctx, state } = ctxFor(repo);
state.worktree = { dir: "/tmp/prev", branch: "locode-wt-prev", originalCwd: repo };
const result = (await enterWorktreeTool.handler({ name: "second" }, ctx)) as { error: string };
expect(result.error).toMatch(/already in a worktree/i);
});
it("enter_worktree returns an error (without switching) for a non-git cwd", async () => {
// repo dir exists but no git init.
const { ctx, state } = ctxFor(repo);
const result = (await enterWorktreeTool.handler({ name: "x" }, ctx)) as { error: string };
expect(result.error).toMatch(/Not a git repository/i);
expect(state.worktree).toBeUndefined();
expect(state.cwd).toBe(repo);
});
it("enter_worktree creates a worktree, switches cwd, and records the worktree", async () => {
await gitInit(repo);
const { ctx, state } = ctxFor(repo);
const result = (await enterWorktreeTool.handler({ name: "feature" }, ctx)) as {
dir: string;
branch: string;
message: string;
};
expect(result.branch).toBe("locode-wt-feature");
expect(existsSync(result.dir)).toBe(true);
// The tool switched the cwd into the worktree and recorded the tracking (with the original cwd).
expect(state.cwd).toBe(result.dir);
expect(state.worktree).toEqual({ dir: result.dir, branch: "locode-wt-feature", originalCwd: repo });
// Clean up the created worktree+branch so afterEach's repo removal is clean.
await removeInteractiveWorktree(repo, result.dir, result.branch);
});
it("exit_worktree returns an error when not in a worktree session", async () => {
const { ctx } = ctxFor(repo);
const result = (await exitWorktreeTool.handler({ action: "keep" }, ctx)) as { error: string };
expect(result.error).toMatch(/not in a worktree/i);
});
it("exit_worktree(keep) restores the original cwd and clears the worktree tracking", async () => {
await gitInit(repo);
const { ctx, state } = ctxFor(repo);
const enter = (await enterWorktreeTool.handler({ name: "keepme" }, ctx)) as { dir: string; branch: string };
expect(state.cwd).toBe(enter.dir);
const result = (await exitWorktreeTool.handler({ action: "keep" }, ctx)) as {
action: string;
restoredCwd: string;
};
expect(result.action).toBe("keep");
expect(result.restoredCwd).toBe(repo);
expect(state.cwd).toBe(repo);
expect(state.worktree).toBeUndefined();
// keep leaves the worktree dir + branch in place.
expect(existsSync(enter.dir)).toBe(true);
// Clean up the kept worktree so afterEach can remove the repo.
await removeInteractiveWorktree(repo, enter.dir, enter.branch);
});
it("exit_worktree(remove) refuses a dirty worktree without discardChanges (no removal, no restore)", async () => {
await gitInit(repo);
const { ctx, state } = ctxFor(repo);
const enter = (await enterWorktreeTool.handler({ name: "dirty" }, ctx)) as { dir: string; branch: string };
writeFileSync(path.join(enter.dir, "README.md"), "changed\n"); // uncommitted change
const result = (await exitWorktreeTool.handler({ action: "remove" }, ctx)) as { error: string };
expect(result.error).toMatch(/uncommitted changes/i);
// Refusal leaves everything in place: still in the worktree, worktree still tracked, dir still exists.
expect(state.cwd).toBe(enter.dir);
expect(state.worktree).toBeDefined();
expect(existsSync(enter.dir)).toBe(true);
// Now opt in to discard → removes + restores.
const ok = (await exitWorktreeTool.handler({ action: "remove", discardChanges: true }, ctx)) as {
action: string;
restoredCwd: string;
};
expect(ok.action).toBe("remove");
expect(ok.restoredCwd).toBe(repo);
expect(state.cwd).toBe(repo);
expect(state.worktree).toBeUndefined();
expect(existsSync(enter.dir)).toBe(false);
});
it("exit_worktree(remove) on a clean worktree removes + restores without discardChanges", async () => {
await gitInit(repo);
const { ctx, state } = ctxFor(repo);
const enter = (await enterWorktreeTool.handler({ name: "clean" }, ctx)) as { dir: string; branch: string };
const result = (await exitWorktreeTool.handler({ action: "remove" }, ctx)) as {
action: string;
restoredCwd: string;
};
expect(result.action).toBe("remove");
expect(state.cwd).toBe(repo);
expect(state.worktree).toBeUndefined();
expect(existsSync(enter.dir)).toBe(false);
});
});
describe("worktree tool flags + schema", () => {
it("enter_worktree is non-mutating; exit_worktree is mutating (destructive on remove)", () => {
expect(enterWorktreeTool.mutating).toBe(false);
expect(exitWorktreeTool.mutating).toBe(true);
});
it("exit_worktree action is required and must be keep|remove", () => {
expect(() => exitWorktreeTool.schema.parse({})).toThrow();
expect(() => exitWorktreeTool.schema.parse({ action: "other" })).toThrow();
expect(() => exitWorktreeTool.schema.parse({ action: "keep" })).not.toThrow();
});
it("enter_worktree name is optional", () => {
expect(() => enterWorktreeTool.schema.parse({})).not.toThrow();
expect(() => enterWorktreeTool.schema.parse({ name: "feature" })).not.toThrow();
});
});
+135
View File
@@ -0,0 +1,135 @@
import { z } from "zod";
import type { ToolDef } from "./types.js";
import {
createInteractiveWorktree,
hasUncommittedChanges,
removeInteractiveWorktree,
} from "../utils/worktree.js";
const enterSchema = z.object({
name: z
.string()
.optional()
.describe(
"Optional name for the worktree and its branch (locode-wt-<name>). Letters, digits, dot, " +
"underscore, dash, starting alphanumeric, max 64 chars. Omit to auto-generate. Must be unique — " +
"re-using an existing branch name returns an error.",
),
});
const exitSchema = z.object({
action: z
.enum(["keep", "remove"])
.describe(
"'keep' leaves the worktree directory and branch in place (restored to the original working " +
"directory; the branch is preserved in git). 'remove' deletes the worktree directory AND the " +
"branch (irreversible).",
),
discardChanges: z
.boolean()
.optional()
.describe(
"Only matters for action 'remove'. If the worktree has uncommitted changes, removal is refused " +
"unless this is true (the changes are then discarded along with the worktree and branch).",
),
});
/** Enters an interactive git worktree session: creates a new branch `locode-wt-<name>` at HEAD in a
* throwaway directory, switches the session's working directory into it, and remembers the original
* cwd so `exit_worktree` can restore it. While in the worktree, all file tools operate there (the
* workspace root becomes the worktree), so experiments can't touch the user's uncommitted work in
* the main repo. Non-mutating: it creates an isolated copy, not an edit to user files. Refuses if
* already in a worktree (exit first) or the cwd isn't a git repo. */
export const enterWorktreeTool: ToolDef<z.infer<typeof enterSchema>> = {
name: "enter_worktree",
description:
"Create an isolated git worktree on a new branch (locode-wt-<name>) at the current HEAD and switch the " +
"session's working directory into it. Use it to try changes without touching the main working tree — " +
"the worktree starts from the last commit, so uncommitted changes in the main repo don't carry over. " +
"While inside, every file tool (read_file, edit_file, grep, bash, etc.) operates in the worktree. Leave " +
"with exit_worktree (keep preserves the branch; remove discards it). Refuses if already in a worktree " +
"(exit first) or the cwd isn't a git repo. Only the main session should use this (not sub-agents).",
schema: enterSchema,
mutating: false,
handler: async (args, ctx) => {
if (!ctx.setCwd || !ctx.setWorktree || !ctx.getWorktree) {
throw new Error("Worktree sessions are not available in this context.");
}
if (ctx.getWorktree()) {
return { error: "Already in a worktree session. Use exit_worktree (keep or remove) before entering another." };
}
const parentCwd = ctx.cwd;
try {
const wt = await createInteractiveWorktree(parentCwd, args.name);
ctx.setWorktree({ dir: wt.dir, branch: wt.branch, originalCwd: parentCwd });
ctx.setCwd(wt.dir);
return {
dir: wt.dir,
branch: wt.branch,
message:
`Switched into worktree at ${wt.dir} on branch ${wt.branch}. File tools now operate there. ` +
`Use exit_worktree to leave (keep the branch, or remove it).`,
};
} catch (err) {
return { error: (err as Error).message };
}
},
};
/** Leaves the active worktree session, restoring the session's working directory to the original cwd.
* `action: "keep"` leaves the worktree dir + branch in place (the branch persists in git; the temp dir
* remains until the process/OS reclaims it). `action: "remove"` deletes the worktree dir AND the
* branch — refused if the worktree has uncommitted changes unless `discardChanges: true`. Mutating
* (a remove is destructive), so the user is asked to confirm. Refuses if not in a worktree session. */
export const exitWorktreeTool: ToolDef<z.infer<typeof exitSchema>> = {
name: "exit_worktree",
description:
"Leave the active worktree session, restoring the working directory to where it was before enter_worktree. " +
"action 'keep' preserves the worktree directory and branch (the branch stays in git — recover the work via " +
"git checkout/worktree add later); action 'remove' deletes the worktree directory AND the branch. A remove " +
"is refused when the worktree has uncommitted changes unless discardChanges is true. Use this only after " +
"enter_worktree; it returns an error if you're not in a worktree session.",
schema: exitSchema,
mutating: true,
handler: async (args, ctx) => {
if (!ctx.setCwd || !ctx.setWorktree || !ctx.getWorktree) {
throw new Error("Worktree sessions are not available in this context.");
}
const wt = ctx.getWorktree();
if (!wt) {
return { error: "Not in a worktree session — nothing to exit." };
}
if (args.action === "remove") {
// Refuse to silently destroy uncommitted work; the model must opt in via discardChanges.
let dirty = false;
try {
dirty = await hasUncommittedChanges(wt.dir);
} catch {
dirty = false;
}
if (dirty && !args.discardChanges) {
return {
error:
`Worktree "${wt.branch}" has uncommitted changes. Re-run with discardChanges: true to discard them ` +
`along with the worktree and branch, or use action: "keep" to preserve them.`,
};
}
try {
await removeInteractiveWorktree(wt.originalCwd, wt.dir, wt.branch);
} catch (err) {
return { error: `Failed to remove worktree: ${(err as Error).message}` };
}
}
// Restore the session cwd for both keep and remove, and clear the tracking.
ctx.setCwd(wt.originalCwd);
ctx.setWorktree(undefined);
return {
action: args.action,
restoredCwd: wt.originalCwd,
message:
args.action === "keep"
? `Left worktree ${wt.dir} in place (branch ${wt.branch} is preserved in git). Restored working directory to ${wt.originalCwd}.`
: `Removed worktree ${wt.dir} and deleted branch ${wt.branch}. Restored working directory to ${wt.originalCwd}.`,
};
},
};
+29 -38
View File
@@ -1,45 +1,36 @@
import { mkdirSync, mkdtempSync, readFileSync, rmSync } from "node:fs";
import os from "node:os";
import { mkdtemp, readFile, rm } from "node:fs/promises";
import { tmpdir } from "node:os";
import path from "node:path";
import { afterEach, beforeEach, describe, expect, it } from "vitest";
import { describe, expect, it } from "vitest";
import { writeFileTool } from "./writeFile.js";
import type { ToolContext } from "./types.js";
describe("writeFile tool — path containment", () => {
let cwd: string;
let ctx: ToolContext;
beforeEach(() => {
cwd = mkdtempSync(path.join(os.tmpdir(), "locode-writefile-"));
ctx = { cwd };
describe("writeFileTool", () => {
it("rejects paths that escape the working directory", async () => {
const cwd = await mkdtemp(path.join(tmpdir(), "locode-writefile-"));
try {
await expect(
writeFileTool.handler({ path: "../outside.txt", content: "x" }, { cwd }),
).rejects.toThrow("Path resolves outside the working directory");
await expect(
writeFileTool.handler({ path: "sub/../../outside.txt", content: "x" }, { cwd }),
).rejects.toThrow("Path resolves outside the working directory");
await expect(
writeFileTool.handler({ path: "/etc/passwd", content: "x" }, { cwd }),
).rejects.toThrow("Path resolves outside the working directory");
} finally {
await rm(cwd, { recursive: true, force: true });
}
});
afterEach(() => {
rmSync(cwd, { recursive: true, force: true });
});
it("writes a file inside the working directory", async () => {
const result = (await writeFileTool.handler({ path: "note.txt", content: "hi" }, ctx)) as { path: string };
expect(readFileSync(result.path, "utf-8")).toBe("hi");
});
it("refuses to write outside the working directory via ../ traversal", async () => {
await expect(writeFileTool.handler({ path: "../escape.txt", content: "oops" }, ctx)).rejects.toThrow(/outside the working directory/);
});
it("refuses to write to an absolute path outside the working directory", async () => {
const outside = path.join(os.tmpdir(), "locode-outside-target.txt");
await expect(writeFileTool.handler({ path: outside, content: "oops" }, ctx)).rejects.toThrow(/outside the working directory/);
});
it("preview reports the block instead of showing a diff", async () => {
const preview = await writeFileTool.preview!({ path: "../escape.txt", content: "oops" }, ctx);
expect(preview).toMatch(/outside the working directory/);
});
it("still applies even when the escaping subdirectory already exists", async () => {
// Sanity check that the guard runs before mkdir/write, not after.
mkdirSync(path.join(cwd, "sub"), { recursive: true });
await expect(writeFileTool.handler({ path: "sub/../../escape.txt", content: "oops" }, ctx)).rejects.toThrow(/outside the working directory/);
it("writes files inside the working directory", async () => {
const cwd = await mkdtemp(path.join(tmpdir(), "locode-writefile-"));
try {
const result = (await writeFileTool.handler({ path: "nested/file.txt", content: "hello" }, { cwd })) as { path: string };
expect(result.path).toBe(path.join(cwd, "nested/file.txt"));
const content = await readFile(path.join(cwd, "nested/file.txt"), "utf-8");
expect(content).toBe("hello");
} finally {
await rm(cwd, { recursive: true, force: true });
}
});
});
+16 -22
View File
@@ -1,9 +1,8 @@
import { createPatch } from "diff";
import { randomUUID } from "node:crypto";
import { mkdir, readFile as fsReadFile, rename, unlink, writeFile as fsWriteFile } from "node:fs/promises";
import { mkdir, readFile as fsReadFile, writeFile as fsWriteFile } from "node:fs/promises";
import path from "node:path";
import { z } from "zod";
import { resolveWithinCwd } from "./pathGuard.js";
import { assertWithinWorkspace } from "../utils/path.js";
import type { ToolDef } from "./types.js";
const schema = z.object({
@@ -14,24 +13,21 @@ const schema = z.object({
async function readExisting(resolved: string): Promise<string | null> {
try {
return await fsReadFile(resolved, "utf-8");
} catch (err: any) {
if (err?.code === "ENOENT") return null;
throw err;
} catch {
return null;
}
}
export const writeFileTool: ToolDef<z.infer<typeof schema>> = {
name: "write_file",
description: "Create or overwrite a file with the given content. Use for new files or full rewrites. For small changes to an existing file, prefer edit_file instead.",
description:
"Create or overwrite a file with the given content. Use this for new files or when rewriting " +
"most of a file; prefer edit_file for small, targeted changes.",
schema,
mutating: true,
preview: async ({ path: filePath, content }, ctx) => {
let resolved: string;
try {
resolved = resolveWithinCwd(ctx.cwd, filePath);
} catch (err) {
return (err as Error).message;
}
const resolved = path.resolve(ctx.cwd, filePath);
assertWithinWorkspace(resolved, ctx.cwd, filePath);
const existing = await readExisting(resolved);
if (existing === null) {
return `Create new file ${resolved} (${content.length} chars)`;
@@ -39,16 +35,14 @@ export const writeFileTool: ToolDef<z.infer<typeof schema>> = {
return createPatch(resolved, existing, content, "", "");
},
handler: async ({ path: filePath, content }, ctx) => {
const resolved = resolveWithinCwd(ctx.cwd, filePath);
const resolved = path.resolve(ctx.cwd, filePath);
assertWithinWorkspace(resolved, ctx.cwd, filePath);
const previousContent = await readExisting(resolved);
await mkdir(path.dirname(resolved), { recursive: true });
const tmpPath = resolved + ".tmp-" + randomUUID();
try {
await fsWriteFile(tmpPath, content, "utf-8");
await rename(tmpPath, resolved);
} catch (err) {
try { await unlink(tmpPath); } catch {}
throw err;
await fsWriteFile(resolved, content, "utf-8");
if (ctx.setLastEdit) {
ctx.setLastEdit({ path: filePath, previousContent: previousContent ?? "" });
}
return { path: resolved, bytesWritten: Buffer.byteLength(content, "utf-8") };
},
};
};
+565 -465
View File
File diff suppressed because it is too large Load Diff
+54
View File
@@ -0,0 +1,54 @@
import { describe, expect, it } from "vitest";
import { getGraphemeBoundaries, nextGraphemeBoundary } from "./ChatInput.js";
describe("getGraphemeBoundaries", () => {
it("returns [0, length] for an empty string", () => {
expect(getGraphemeBoundaries("")).toEqual([0]);
});
it("returns boundaries for plain ASCII", () => {
expect(getGraphemeBoundaries("abc")).toEqual([0, 1, 2, 3]);
});
it("treats surrogate pairs (emoji/Hangul) as single graphemes", () => {
// "a👍b" — thumbs up is a surrogate pair (2 JS indices, 1 displayed cell).
const boundaries = getGraphemeBoundaries("a👍b");
expect(boundaries).toEqual([0, 1, 3, 4]);
});
it("treats ZWJ emoji sequences as single graphemes", () => {
// "👨‍👩‍👧‍👦" is a family emoji made of multiple code points joined with ZWJs.
const str = "👨‍👩‍👧‍👦";
const boundaries = getGraphemeBoundaries(str);
expect(boundaries).toHaveLength(2);
expect(boundaries).toContain(0);
expect(boundaries).toContain(str.length);
});
});
describe("nextGraphemeBoundary", () => {
it("moves right past a surrogate-pair emoji", () => {
// "a👍b", cursor after "a" (index 1) should jump to index 3 (after emoji).
expect(nextGraphemeBoundary("a👍b", 1, 1)).toBe(3);
});
it("moves left past a surrogate-pair emoji", () => {
// cursor at index 3 (after emoji) should jump back to index 1 (before emoji).
expect(nextGraphemeBoundary("a👍b", 3, -1)).toBe(1);
});
it("does not move past the start or end", () => {
expect(nextGraphemeBoundary("ab", 0, -1)).toBe(0);
expect(nextGraphemeBoundary("ab", 2, 1)).toBe(2);
});
it("snaps an invalid offset to the next boundary when moving right", () => {
// index 2 is inside the emoji surrogate pair.
expect(nextGraphemeBoundary("a👍b", 2, 1)).toBe(3);
});
it("snaps an invalid offset to the previous boundary when moving left", () => {
// index 2 is inside the emoji surrogate pair.
expect(nextGraphemeBoundary("a👍b", 2, -1)).toBe(1);
});
});

Some files were not shown because too many files have changed in this diff Show More