fix: 뉴스 검색이 2분 넘게 걸리던 원인 두 가지

사용자 신고: "지금 뉴스검색을 2분 넘게 하고 있는데?"
"오늘 미국 주요 뉴스" 한 턴에서 도구를 11번 불렀다. 라운드마다 muse-glimmer가
통째로 한 번씩 생성해야 하는데 22 tok/s dense라 라운드 수가 그대로 시간이 된다.

1) 1면 RSS가 제목과 링크만 넘겨서, 모델이 요약을 쓰려고 기사를 하나씩
   web_fetch 했다(6번, 그중 NYT는 403). 정작 피드엔 description이 이미 실려
   있었다 — NYT 138자, NPR 293자 — 파서가 지나치고 버리고 있었을 뿐이다.
   Headline에 summary를 넣고 formatHeadlines가 함께 내보낸다. 실측: us 피드
   12건 전부 요약 확보.

   NPR은 <em> 식으로 이중 인코딩해 보내므로 디코드→태그제거→디코드를
   한 번 더 돈다. 한 번만 돌면 <em>이 그대로 남는다.

2) 가드 두 개가 서로 물려 헛바퀴 3회를 돌았다. 모델이 category를 붙여 다시
   부르면 → breadth 가드가 "요청 안 한 category"라며 떼어냄 → 인자가 앞 호출과
   같아짐 → 중복으로 스킵 → 모델은 데이터를 못 받았으니 다른 category로 또 시도.

   중복 스킵에는 원래 replay 경로가 있는데 대상이 coder_list_files와
   coder_read_file 둘뿐이라, 뉴스·검색은 데이터 대신 "Already ran this exact
   call"이라는 문구만 받았다. 결과가 없으니 변형해서 또 부르는 게 당연하다.
   news_search/web_search/web_fetch를 replay 대상에 넣어 이전 결과를 그대로
   돌려준다 — 모델이 원하던 걸 받으면 반복이 끝난다.

테스트 8개 추가(요약 추출, Atom summary, 이중 인코딩, CDATA, 요약 없는 항목,
길이 제한, 출력 포함, 빈 줄 미발생) — 354개 통과.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
kim
2026-08-15 08:39:45 +09:00
co-authored by Claude Opus 5
parent 554b66d299
commit faee483cf1
3 changed files with 100 additions and 4 deletions
+12 -1
View File
@@ -405,8 +405,19 @@ async function handleChat(
const browserPacketMaxItems = Math.max(12, Math.min(60, Math.min(browserMaxCollectedItems, 40)));
const seenToolCalls = new Set<string>();
const cachedReadOnlyToolResults = new Map<string, ToolResult>();
// Read-only lookups whose answer is stable for the length of one turn, so a repeat can be served
// from the cache instead of being turned away.
//
// The distinction matters more than it looks. A turned-away call gets "Already ran this exact
// call" — a scolding with no data in it — and a model that still needs the data simply asks
// again with something varied. Observed 2026-08-15 on "오늘 미국 주요 뉴스": the model kept
// adding a category, the breadth guard stripped the unrequested category back off, that made the
// args identical to the previous call, and it was skipped — three wasted rounds of two guards
// feeding each other, each round a full generation on a 22 tok/s model. Handing back the
// previous result ends it, because the model gets what it was actually after.
const canReplayReadOnlyCall = (toolName: string): boolean =>
toolName === 'coder_list_files' || toolName === 'coder_read_file';
toolName === 'coder_list_files' || toolName === 'coder_read_file'
|| toolName === 'news_search' || toolName === 'web_search' || toolName === 'web_fetch';
const loopDetectionEnabled = orchRuntimeCfg?.triggers?.loop_detection !== false;
const loopWarningThreshold = 3;
const loopCriticalThreshold = 5;
+23 -3
View File
@@ -38,12 +38,20 @@ const FEEDS: Record<string, string[]> = {
const FETCH_TIMEOUT_MS = 10_000;
const PER_FEED = 6;
export interface Headline { title: string; link: string; source: string }
export interface Headline { title: string; link: string; source: string; summary: string }
const decodeEntities = (s: string) => s
.replace(/&lt;/g, '<').replace(/&gt;/g, '>').replace(/&quot;/g, '"')
.replace(/&#0?39;|&apos;/g, "'").replace(/&nbsp;/g, ' ').replace(/&amp;/g, '&');
/**
* Feed summaries arrive with entities encoded at two depths — `&lt;em&gt;` unescapes into real
* markup that then has to be stripped, which one decode-then-strip pass leaves behind as visible
* tags. Decode, strip, and decode again so the model gets prose instead of `<em>` litter.
*/
const cleanSummary = (raw: string): string =>
decodeEntities(decodeEntities(raw).replace(/<[^>]+>/g, ' ')).replace(/\s+/g, ' ').trim();
const tagValue = (block: string, tag: string): string => {
const m = block.match(new RegExp(`<${tag}[^>]*>(?:<!\\[CDATA\\[)?([\\s\\S]*?)(?:\\]\\]>)?</${tag}>`));
return m ? decodeEntities(m[1].replace(/<[^>]+>/g, '').trim()) : '';
@@ -55,6 +63,12 @@ const tagValue = (block: string, tag: string): string => {
* Exported and pure so it can be tested against saved feed text: these are third-party documents
* whose markup changes without notice, and the failure that matters is the silent one — a parser
* that quietly returns nothing looks exactly like a quiet news day.
*
* The summary is not decoration. Handing the model bare headlines made it go fetch each article
* to write about it — measured 2026-08-15 on one "오늘 미국 주요 뉴스" turn: six web_fetch calls,
* each its own generation round on a 22 tok/s model, which is most of why the answer took over two
* minutes. Every one of these feeds already ships a description (NYT ~138 chars, NPR ~293), so the
* summary was being parsed past and thrown away.
*/
export function parseFeedItems(xml: string, source: string, limit = PER_FEED): Headline[] {
const doc = String(xml || '');
@@ -64,7 +78,8 @@ export function parseFeedItems(xml: string, source: string, limit = PER_FEED): H
const title = tagValue(b, 'title');
if (!title) continue;
const link = tagValue(b, 'link') || (b.match(/<link[^>]*href="([^"]+)"/)?.[1] ?? '');
out.push({ title, link, source });
const rawSummary = (b.match(/<(description|summary|content:encoded)[^>]*>(?:<!\[CDATA\[)?([\s\S]*?)(?:\]\]>)?<\/\1>/) || [])[2] || '';
out.push({ title, link, source, summary: cleanSummary(rawSummary).slice(0, 400) });
if (out.length >= limit) break;
}
return out;
@@ -116,6 +131,11 @@ export async function fetchTopHeadlines(country: string): Promise<Headline[]> {
export function formatHeadlines(items: readonly Headline[], limit = 12): string {
return items.slice(0, limit)
.map((h, i) => `[${i + 1}] ${h.title}\n ${h.link}\n Source: ${h.source}`)
.map((h, i) => {
const lines = [`[${i + 1}] ${h.title}`];
if (h.summary) lines.push(` ${h.summary}`);
lines.push(` ${h.link}`, ` Source: ${h.source}`);
return lines.join('\n');
})
.join('\n');
}
+65
View File
@@ -78,3 +78,68 @@ describe('formatHeadlines', () => {
assert.match(out, /Source: nytimes\.com/);
});
});
/**
* 요약 추출 (2026-08-15 추가)
*
* 제목만 넘기면 모델이 기사를 하나씩 web_fetch 하러 간다 — 실측: "오늘 미국 주요 뉴스" 한 턴에
* web_fetch 6번, 매번 22 tok/s 모델의 생성 라운드 하나씩. 정작 피드엔 description이 이미
* 실려 있었는데(NYT 138자, NPR 293자) 파서가 지나치고 버리고 있었다.
*/
describe('요약 추출', () => {
test('description을 뽑는다', () => {
const xml = `<rss><channel><item>
<title>제목</title><link>https://e.com/1</link>
<description>디에고 가르시아가 물류 허브가 되었다.</description>
</item></channel></rss>`;
assert.equal(parseFeedItems(xml, 'e.com')[0].summary, '디에고 가르시아가 물류 허브가 되었다.');
});
test('Atom의 summary도 뽑는다', () => {
const xml = `<feed><entry>
<title>T</title><link href="https://e.com/2"/><summary>요약본이다.</summary>
</entry></feed>`;
assert.equal(parseFeedItems(xml, 'e.com')[0].summary, '요약본이다.');
});
test('이중 인코딩된 마크업을 걷어낸다 — NPR이 &lt;em&gt;로 보내온다', () => {
const xml = `<rss><channel><item><title>T</title><link>u</link>
<description>Clint Smith&apos;s &lt;em&gt;How the Word Is Passed&lt;/em&gt; explores it.</description>
</item></channel></rss>`;
const s = parseFeedItems(xml, 'x')[0].summary;
assert.ok(!s.includes('<em>') && !s.includes('&lt;'), `태그가 남았다: ${s}`);
assert.ok(s.includes("Clint Smith's") && s.includes('How the Word Is Passed'));
});
test('CDATA 안의 요약도 뽑는다', () => {
const xml = `<rss><channel><item><title>T</title><link>u</link>
<description><![CDATA[<p>본문 요약</p>]]></description></item></channel></rss>`;
assert.equal(parseFeedItems(xml, 'x')[0].summary, '본문 요약');
});
test('요약이 없어도 기사 자체는 살린다', () => {
const xml = `<rss><channel><item><title>제목만</title><link>u</link></item></channel></rss>`;
const items = parseFeedItems(xml, 'x');
assert.equal(items.length, 1);
assert.equal(items[0].summary, '');
});
test('지나치게 긴 요약은 자른다', () => {
const xml = `<rss><channel><item><title>T</title><link>u</link>
<description>${'가'.repeat(2000)}</description></item></channel></rss>`;
assert.ok(parseFeedItems(xml, 'x')[0].summary.length <= 400);
});
test('formatHeadlines가 요약을 실어 보낸다', () => {
const xml = `<rss><channel><item><title>제목</title><link>https://e.com/1</link>
<description>핵심 요약</description></item></channel></rss>`;
const out = formatHeadlines(parseFeedItems(xml, 'e.com'));
assert.ok(out.includes('핵심 요약'), out);
assert.ok(out.includes('https://e.com/1'));
});
test('요약이 없으면 빈 줄을 넣지 않는다', () => {
const xml = `<rss><channel><item><title>제목만</title><link>u</link></item></channel></rss>`;
assert.ok(!/\n\s*\n/.test(formatHeadlines(parseFeedItems(xml, 'x'))));
});
});