perf: 검색 결과 읽기 지침을 시스템 프롬프트에서 빼 결과가 생긴 뒤에만 붙인다

효율 감사 실측(사용자 턴 1,011건 / LLM 호출 3,613건 / prompt 누적 29.1M tok):
  · 고정비 2,494tok/턴 = 전체 prompt의 31%
  · 출력 60토큰 미만 호출이 47% — 도구 라운드·가드 재프롬프트라 고정비가 라운드 수만큼 곱해진다
  · **59%의 턴이 grounding 도구를 한 번도 부르지 않는다**

searchReadingRule(481tok)에는 성격이 다른 두 가지가 한 덩어리로 섞여 있었다:
  ① 검색어를 어떻게 쓸 것인가(세대 추측 금지, 154tok) — 검색 **전에** 필요
  ② 돌아온 결과를 어떻게 읽을 것인가(스테일 타임스탬프·표 행 정렬, 333tok) — 결과가 **생긴 뒤**
     에만 쓸모가 있다
②를 groundingResultReadingRule()로 분리해, handle-chat이 첫 grounding 결과가 들어온 시점에
대화 뒤에 한 번 덧붙인다. 읽을 결과가 없는 59%의 턴은 이제 이 값을 아예 내지 않고, 부르는
턴도 검색어를 쓰는 1라운드에는 내지 않는다. 인사 턴 시스템 프롬프트 1,273 → 904tok(-29%).

**시스템 메시지를 갈아끼우지 않고 뒤에 덧붙인 이유는 KV 캐시다.** 시스템 프롬프트는 라운드 루프
앞에서 한 번 만들어지므로, 라운드마다 다시 만들면 프리픽스가 달라져 캐시가 통째로 무효화된다 —
라운드마다 7K 토큰을 다시 계산하게 되어 아낀 것보다 잃는 게 크다. 뒤에 붙이면 프리픽스는 그대로다.

부수 효과로 지침이 대상(검색 결과) 바로 옆에 놓인다. 제약이 겹치면 먼 것부터 버리는 모델에게는
이쪽이 유리하다([[feedback_local_model_needs_code_backstop]]).

2026-09-02 감사가 만든 hasGroundingTools 게이트는 사실상 죽어 있었다는 점도 이번에 드러났다 —
tool-scope에 web_search를 거르는 분기가 없어 그 조건은 항상 참이다. 이번 분리는 그 게이트에
의존하지 않는다.

기존 회귀 테스트가 이동을 정확히 잡아냈다(2026-08-10 캐시된 기상표 오독 사고의 예시 문자열).
테스트를 새 위치로 옮겨 내용이 사라지지 않았음을 계속 고정한다. 546개 통과.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rfme1WVEPkpNwnf5oVNXXc
This commit is contained in:
kim
2026-09-04 20:59:35 +09:00
co-authored by Claude Opus 5
parent 83e9bd54f9
commit 46a5c3f59f
3 changed files with 80 additions and 9 deletions
+17 -1
View File
@@ -78,7 +78,7 @@ import { createPreflightAdvisor, type PreflightAdvisorCtx } from './preflight-ad
import { createBrowserDesktopAdvisors, type BrowserDesktopAdvisorCtx } from './browser-desktop-advisor';
import type { SkillsManager } from '../skills-manager';
import { selectToolsForTurn, bootAllowedTools, codeAiBlockedTools } from './tool-scope';
import { buildChatSystemPrompt } from './system-prompt';
import { buildChatSystemPrompt, groundingResultReadingRule } from './system-prompt';
import { replyLooksEmpty, appendDroppedSearchImages } from './reply-content';
import { evaluateSearchBudget } from './search-budget';
import { decideAutoRecover, shouldForceMessagingRetry as decideMessagingRetry, shouldForceEmptyGroundingRetry as decideEmptyGroundingRetry, GROUNDING_TOOL_PATTERN } from './retry-decisions';
@@ -321,6 +321,8 @@ async function handleChat(
// 이미 나와 있는데 라운드만 더 쓴다.
let deadEndFallbackNudges = 0;
const MAX_DEAD_END_FALLBACK_NUDGES = 1;
// 검색 결과 읽기 지침은 턴당 한 번만 붙인다.
let groundingReadingRuleInjected = false;
let sycophancyNudges = 0;
const MAX_SYCOPHANCY_NUDGES = 1;
let staleVersionQueryNudges = 0;
@@ -3105,6 +3107,20 @@ async function handleChat(
// any call) is deliberately spreading across topics for diversity per the "call news_search
// several times across categories" instruction elsewhere; collapsing THAT to top-N by
// relevance-to-nothing would undo the exact reason it exists.
// 검색 결과 읽는 법을, 읽을 결과가 실제로 생긴 뒤에 한 번만 붙인다(2026-09-04 효율 감사).
//
// 예전에는 이 333토큰이 시스템 프롬프트에 상주했다. 실측해 보니 사용자 턴 1,011건 중 **59%가
// grounding 도구를 한 번도 부르지 않는다** — 그 턴들은 읽을 결과가 영영 없는데도 매 라운드
// 이 값을 냈고, 부르는 턴에서도 검색어를 쓰는 1라운드에는 아직 읽을 게 없었다.
//
// 시스템 메시지를 갈아끼우지 않고 **뒤에 덧붙이는** 이유는 KV 캐시다. 프리픽스가 바뀌면
// 라운드마다 프롬프트 전체를 다시 계산하게 되어, 아낀 토큰보다 잃는 게 크다.
if (!groundingReadingRuleInjected
&& allToolResults.some(r => GROUNDING_TOOL_PATTERN.test(String(r?.name || '')) && !r.error)) {
groundingReadingRuleInjected = true;
messages.push({ role: 'user', content: groundingResultReadingRule(dateStr) });
}
{
const roundResults = allToolResults.slice(roundStartResultsIdx);
const newsCallsThisRound = roundResults.filter(r => r.name === 'news_search' && !r.error);
+37 -1
View File
@@ -48,6 +48,27 @@ export interface SystemPromptInput {
message?: string;
}
/**
* 검색 결과를 어떻게 읽을 것인가. 시스템 프롬프트가 아니라, **첫 grounding 결과가 들어온 뒤**
* 대화 뒤에 한 번 덧붙인다(handle-chat).
*
* 시스템 프롬프트에 있던 것을 2026-09-04 효율 감사로 옮겼다. 근거는 두 가지다.
* · 실측: 사용자 턴 1,011건 중 59%가 grounding 도구를 한 번도 부르지 않는다. 그 턴들은 읽을
* 결과가 영영 없는데도 이 333토큰을 매 라운드 냈다.
* · 위치: 이 지침이 말하는 대상(검색 결과)이 바로 앞에 있을 때 읽힌다. 제약이 겹치면 먼 것부터
* 버리는 모델에게는 붙여 두는 편이 낫다([[feedback_local_model_needs_code_backstop]]).
*
* 시스템 메시지를 라운드마다 다시 만드는 대신 뒤에 덧붙이는 이유는 KV 캐시다 — 프리픽스가
* 바뀌면 라운드마다 프롬프트 전체를 다시 계산하게 된다.
*/
export function groundingResultReadingRule(dateStr: string): string {
return `You now have search results in front of you. Read them with these rules:
- A result reporting a "closing price" or "today's headline" is not proof today is a trading/business day — cross-check against the current date (${dateStr}). If they conflict, trust the current date and say so rather than presenting the result as live.
- Before treating any result as current, look for its own timestamp: a date in the URL (e.g. "tm=2024.12.13.20:00", "?date=..."), a dateline, or "as of ..." phrasing. If it is more than a day or two old, it is NOT "지금"/"현재" — say plainly that you could not find live data and name the stale date you found, instead of presenting old numbers as current. This applies to "is it raining right now" as much as to prices.
- When a result is a data TABLE (observation/measurement tables are the common case), read every row by its own row label, not by position. Misaligning one row's figures onto the next is worse than finding nothing, because it looks authoritative while being wrong. If you are not confident you are reading the right column for the right label, say so rather than guessing.
- If the results describe an older generation of a product than the user implies, say so and search again rather than reporting the old model as new.`;
}
export function buildChatSystemPrompt(input: SystemPromptInput): string {
const executionMode = String(input.executionMode || '');
const dateStr = input.dateStr;
@@ -127,8 +148,23 @@ export function buildChatSystemPrompt(input: SystemPromptInput): string {
// What survives without any grounding tool: the date itself is ground truth, and "is X open
// right now" is answerable from that date alone. Everything else needs a result to judge.
//
// 2026-09-04 효율 감사로 이 블록을 둘로 쪼갰다. 원래는 481토큰이 한 덩어리였고, 그 안에
// 성격이 다른 두 가지가 섞여 있었다:
// ① 검색어를 어떻게 쓸 것인가(세대 추측 금지) — 검색 **전에** 필요하다. 여기 남는다.
// ② 돌아온 결과를 어떻게 읽을 것인가(스테일 타임스탬프·표 행 정렬) — 결과가 **생긴 뒤**에만
// 쓸모가 있다. groundingResultReadingRule()로 빼내 handle-chat이 첫 grounding 결과가
// 들어온 시점에 주입한다.
//
// 왜 시스템 프롬프트에서 빼는가. 실측(사용자 턴 1,011건): **59%의 턴이 grounding 도구를 한
// 번도 부르지 않는다.** 그 턴들은 읽을 결과가 영영 없는데도 매 라운드 이 333토큰을 냈다.
// 부르는 턴에서도 1라운드(검색어를 쓰는 시점)에는 아직 읽을 게 없다.
//
// 왜 라운드별로 시스템 프롬프트를 다시 만들지 않는가. 시스템 메시지는 라운드 루프 앞에서 한 번
// 만들어지고, 중간에 바꾸면 프리픽스가 달라져 KV 캐시가 통째로 무효화된다 — 라운드마다 7K
// 토큰을 다시 계산하게 된다. 뒤에 **덧붙이면** 프리픽스는 그대로다.
const searchReadingRule = hasGroundingTools
? ` A search result reporting a "closing price" or "today's headline" is not proof today is a trading/business day — cross-check it against the actual current date above, and if they conflict (e.g. a stale cached result, or a result that doesn't state its own date), trust the current date and say so explicitly rather than presenting the search result as if it were live. This applies just as much to "is it raining/snowing right now" or any other current-state question: a search engine can hand you a page it cached long ago. Before treating a search result as live, check whether it carries its own timestamp — a date in the URL (e.g. "tm=2024.12.13.20:00", "?date=..."), a dateline, or "as of ..." phrasing — and compare it to the current date above. If that timestamp is more than a day or two old, it is NOT "지금"/"현재": say plainly that you couldn't find live data and name the stale date you found instead of presenting old numbers as current. When a result is a data TABLE (observation/measurement tables are the common case), read every row by its own row label, not by position — misaligning one city's or one row's figures onto the next is worse than finding nothing, because it looks authoritative while being wrong. If you're not confident you're reading the right column for the right label, say so rather than guessing. When the user asks about the "newest"/"latest"/"신형"/"최신" version of a product WITHOUT naming a specific generation, do NOT put a generation you remember (e.g. "M4", "RTX 4090", "iPhone 15") into your search query — whatever was newest when you were trained is probably not newest now. Search with the current year and neutral terms ("Mac mini ${new Date().getFullYear()}", "latest Mac mini"), let the results tell you which generation is current, and answer about that one. If the results still look like an older generation than the user implies, say so and search again rather than reporting the old model as new.`
? ` When the user asks about the "newest"/"latest"/"신형"/"최신" version of a product WITHOUT naming a specific generation, do NOT put a generation you remember (e.g. "M4", "RTX 4090", "iPhone 15") into your search query — whatever was newest when you were trained is probably not newest now. Search with the current year and neutral terms ("Mac mini ${new Date().getFullYear()}", "latest Mac mini") and let the results tell you which generation is current.`
: '';
const newsFormatRule = hasNewsTool
+26 -7
View File
@@ -12,7 +12,7 @@
import { test, describe } from 'node:test';
import assert from 'node:assert/strict';
import { buildChatSystemPrompt, type SystemPromptInput } from '../src/gateway/chat/system-prompt';
import { buildChatSystemPrompt, groundingResultReadingRule, type SystemPromptInput } from '../src/gateway/chat/system-prompt';
const build = (over: Partial<SystemPromptInput> = {}) => buildChatSystemPrompt({
toolNames: [],
@@ -78,14 +78,33 @@ describe('무조건 들어가는 블록', () => {
// 타임스탬프를 갖는지 보라", "결과가 표면 행 라벨로 읽어라"). 그런데 무조건 붙어 있어서
// "안녕하세요"에도 ~700토큰이 실렸다. 그라운딩 도구 보유 여부로 게이팅한다 — 도구가 없으면
// 판단할 결과 자체가 없다. 날짜 자체가 ground truth라는 핵심 문장은 계속 무조건이다.
test('검색결과 해석 지침은 그라운딩 도구가 있을 때만 붙는다', () => {
// [2026-09-04 효율 감사] 이 블록을 다시 둘로 쪼갰다. 실측해 보니 사용자 턴 1,011건 중 **59%가
// grounding 도구를 한 번도 부르지 않는다** — 그 턴들은 읽을 결과가 영영 없는데도 매 라운드
// 333토큰을 냈고, 부르는 턴에서도 검색어를 쓰는 1라운드에는 아직 읽을 게 없었다.
// · 검색어 작성 규칙(세대 추측 금지)은 검색 **전에** 필요하므로 시스템 프롬프트에 남는다.
// · 결과 읽기 규칙은 groundingResultReadingRule()로 빠져, handle-chat이 첫 grounding
// 결과가 들어온 뒤 대화 **뒤에 덧붙인다**(시스템 프리픽스를 안 건드려야 KV 캐시가 산다).
test('검색어 작성 규칙은 그라운딩 도구가 있을 때 시스템 프롬프트에 남는다', () => {
const withSearch = build({ toolNames: ['web_search'] });
assert.ok(/tm=2024\.12\.13\.20:00/.test(withSearch), '실제 사고 사례의 URL 패턴이 예시로 남아있어야 함');
assert.ok(/misaligning one city|misaligning one row/i.test(withSearch), '표 오독(열 뒤바뀜) 경고가 있어야 함');
assert.ok(/newest.*latest.*신형/is.test(withSearch), '세대 추측 금지 규칙이 있어야 함');
assert.ok(!/newest/i.test(build()), '도구 없는 턴엔 빠져야 함');
assert.ok(/ground truth/i.test(build()), '날짜=ground truth 핵심 문장은 남아야 함');
});
const noTools = build();
assert.ok(!/tm=2024\.12\.13\.20:00/.test(noTools), '도구 없는 턴엔 검색결과 해석 지침이 빠져야 함');
assert.ok(/ground truth/i.test(noTools), '날짜=ground truth 핵심 문장은 남아야 함');
test('결과 읽기 지침은 시스템 프롬프트를 떠났다 — 매 턴 내던 고정비였다', () => {
const withSearch = build({ toolNames: ['web_search'] });
assert.ok(!/tm=2024\.12\.13\.20:00/.test(withSearch), '결과 읽기 예시가 시스템 프롬프트에 남아 있으면 안 됨');
assert.ok(!/misaligning one row/i.test(withSearch), '표 오독 경고도 옮겨갔어야 함');
});
test('결과 읽기 지침의 내용은 groundingResultReadingRule에 그대로 살아 있다', () => {
// 2026-08-10 실제 사고(세션 papa@2d4f3f6e): "지금 비오는 곳이 있나?"에 web_search를 부르긴
// 했으나 결과가 URL에 "tm=2024.12.13.20:00"이 박힌 2024년 캐시 페이지였는데도 "현재"라며
// 소개했고, 표를 옮기며 동두천의 강수량을 파주 것으로 뒤바꿨다. 그 지침이 사라지면 안 된다.
const r = groundingResultReadingRule('Friday, September 4, 2026');
assert.ok(/tm=2024\.12\.13\.20:00/.test(r), '실제 사고 사례의 URL 패턴이 예시로 남아있어야 함');
assert.ok(/misaligning one row/i.test(r), '표 오독(행 뒤바뀜) 경고가 있어야 함');
assert.ok(/Friday, September 4, 2026/.test(r), '현재 날짜가 주입돼야 함');
});
test('뉴스·날씨 출력형식은 해당 도구가 있을 때만 붙는다', () => {