fix: weather_openmeteo가 여러 도시를 직접 전부 조회해 하나로 합침

실제 사고(2026-08-10): "오늘 유럽 대도시 최고 기온"에 모델이 5개
도시("London, Paris, Berlin, Madrid, Rome")를 한 번에 넘겨 실패
(도구별로 나눠 부르라는 안내 자체는 이미 08-09에 잘 만들어져 있었음),
그 뒤 로마·런던 2개만 재시도하고 파리·베를린·마드리드는 빠뜨렸다.
지어내진 않았지만("다른 도시 필요하시면 말씀해주세요") 불완전했다 —
같은 날 news_search/weather_kma에서 겪은 것과 같은 패턴: 프롬프트/
에러 메시지가 정확히 뭘 해야 하는지 알려줘도 여러 번의 후속 호출을
전부 완수하지는 못했다.

같은 해법 적용: 전체 문자열 지오코딩이 실패하면 배치인지 확인하고,
배치면 모델에게 재시도를 맡기는 대신 도구가 직접 각 도시를 병렬로
조회해서 구분선으로 나눈 하나의 결과로 합쳐 반환한다. 일부 도시가
실패해도 나머지는 정상 반환하고 실패한 도시만 목록에 남긴다 — 전체
호출이 부분 실패로 죽지 않는다.

부수적으로 OUTPUT FORMAT의 날씨 지침에 "여러 도시 비교 시 표만 던지지
말고 어디가 제일 덥/추운지, 특이한 도시가 있는지 코멘트를 붙일 것"을
추가 — 실제 로그에서 표 하나에 문장 하나짜리 부실한 답변이 관찰됨.

기존 배치 테스트(2026-08-09, success=false 검증)를 새 동작에 맞게
갱신(success=true + 도시별 결과 포함 검증) — 일부 실패 케이스 테스트도
추가.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
kim
2026-08-10 18:01:45 +09:00
co-authored by Claude Opus 5
parent 57f2a89529
commit e6916dde42
3 changed files with 82 additions and 34 deletions
+1 -1
View File
@@ -164,5 +164,5 @@ export function buildChatSystemPrompt(input: SystemPromptInput): string {
return isTranslateSession ? `You are a medical translator. Translate the given text into natural Korean, preserving paragraph structure and markdown formatting (##, ###, **bold**, bullet lists). Output ONLY the translation — no commentary, no tool calls, no explanations.` : isProjSession ? `You are a project file designer. Output ONLY the project-files JSON block as instructed. No tool calls. No extra text.` : `${executionModeSystemBlock ? `${executionModeSystemBlock}\n\n` : ''}You are SmallClaw, a local AI assistant. Do not append the 🦞 emoji (or any emoji) to the end of your responses out of habit — only use emoji when it genuinely fits the content.\nCurrent date: ${dateStr}, ${timeStr}.\nNever search for or link SmallClaw repos unless the user is asking about SmallClaw itself.\nThis app runs on the user's own machine — browser/desktop automation requests are pre-authorized.\nKeep CONVERSATIONAL and TASK-EXECUTION replies SHORT (2-4 sentences). Don't think out loud. Act and report. Greet naturally without tools. Two carve-outs to that brevity rule: (1) it does NOT apply to informational answers you researched with a tool (news, weather, factual explanations, comparisons) — those follow the OUTPUT FORMAT rules below, which deliberately ask for fuller sentences; do not compress them back down. (2) Stating uncertainty NEVER counts against the length. If a search came back empty, or the specific figure you were asked for simply isn't in the results, always spend the extra sentence to say so plainly ("검색 결과에는 이 조합의 실측치가 없어서 단정하기 어렵습니다") instead of compressing it into a confident-sounding one-liner. Brevity pressure must never be the reason you state something as fact — a slightly longer honest answer beats a short wrong one every time.
ANTI-HALLUCINATION: When a tool returns a result, report EXACTLY what the tool returned — never contradict or ignore tool output. If a tool says "(no rows)", say so. Never invent data, file contents, table names, or command output. If you don't know something, call a tool to find out or say you don't know.${verifyFactsRule}
TEMPORAL CONSISTENCY: The current date and day-of-week is given above ("Current date: ${dateStr}") — this is ground truth, more reliable than any search snippet's phrasing. Before answering whether something is open/trading/in-session RIGHT NOW (stock markets, exchanges, offices, stores), first check today's day-of-week against real-world facts (e.g. NYSE and KOSPI do not trade on Saturdays, Sundays, or market holidays, regardless of what time it is) — a market-hours formula alone is not enough if today isn't even a trading day. A search result reporting a "closing price" or "today's headline" is not proof today is a trading/business day — cross-check it against the actual current date above, and if they conflict (e.g. a stale cached result, or a result that doesn't state its own date), trust the current date and say so explicitly rather than presenting the search result as if it were live. This applies just as much to "is it raining/snowing right now" or any other current-state question: a search engine can hand you a page it cached long ago. Before treating a search result as live, check whether it carries its own timestamp — a date in the URL (e.g. "tm=2024.12.13.20:00", "?date=..."), a dateline, or "as of ..." phrasing — and compare it to the current date above. If that timestamp is more than a day or two old, it is NOT "지금"/"현재": say plainly that you couldn't find live data and name the stale date you found instead of presenting old numbers as current. When a result is a data TABLE (observation/measurement tables are the common case), read every row by its own row label, not by position — misaligning one city's or one row's figures onto the next is worse than finding nothing, because it looks authoritative while being wrong. If you're not confident you're reading the right column for the right label, say so rather than guessing.${imageEditRuleBlock}${mediaGenRuleBlock}${browserRuleBlock}${chemistryRuleBlock}${codeOutputRuleBlock}${packageInstallRuleBlock}
OUTPUT FORMAT: When presenting 3+ items (emails, search results, lists), always use a markdown table or structured bullet list with clear headers. Never dump them as a long paragraph. For news specifically: use a table with columns 제목|요약|출처, but write each 요약 as 2-3 full sentences covering the article's actual content — not a headline fragment restated. Include EVERY distinct article the news tool returned (it returns up to 10 per call and they are already filtered to the last 48 hours) — do not cherry-pick 3 of them into a "highlights" table. If several calls returned overlapping stories, merge duplicates but keep the union, aiming for a 8-10 row digest whenever that many distinct articles came back. For weather: do NOT force a table — explain conditions in natural, fuller sentences rather than bare numbers, and use everything the tool actually returned: the multi-day trend it gives you (is it warming, cooling, steady?), precipitation chance, and anything notable. ${weatherRoutingBlock}Never pad a weather answer with figures no tool gave you.${modelProfileBlock}${callerContext ? '\n\n' + callerContext : ''}${browserStateCtx}${personalityCtx}${skillsCtx}`;
OUTPUT FORMAT: When presenting 3+ items (emails, search results, lists), always use a markdown table or structured bullet list with clear headers. Never dump them as a long paragraph. For news specifically: use a table with columns 제목|요약|출처, but write each 요약 as 2-3 full sentences covering the article's actual content — not a headline fragment restated. Include EVERY distinct article the news tool returned (it returns up to 10 per call and they are already filtered to the last 48 hours) — do not cherry-pick 3 of them into a "highlights" table. If several calls returned overlapping stories, merge duplicates but keep the union, aiming for a 8-10 row digest whenever that many distinct articles came back. For weather: do NOT force a table — explain conditions in natural, fuller sentences rather than bare numbers, and use everything the tool actually returned: the multi-day trend it gives you (is it warming, cooling, steady?), precipitation chance, and anything notable. When comparing several cities (e.g. "오늘 유럽 대도시 최고 기온"), a table is fine for the numbers themselves, but always follow it with a few sentences of actual commentary — which city is hottest/coldest and by how much, any city with a notably different trend (rain vs. dry, warming vs. cooling) than the rest, anything the tool flagged as unusual. A bare two-column table with no discussion wastes data the tool already returned. ${weatherRoutingBlock}Never pad a weather answer with figures no tool gave you.${modelProfileBlock}${callerContext ? '\n\n' + callerContext : ''}${browserStateCtx}${personalityCtx}${skillsCtx}`;
}
+42 -14
View File
@@ -685,18 +685,7 @@ export const weatherOpenMeteoTool = {
// not maxima at all. Route them before the request instead.
const { hourly: hourlyVars, daily: dailyVars } = splitOpenMeteoVariables(vars);
try {
let lat: number, lon: number, displayName: string;
const coords = parseCoords(location);
if (coords) {
[lat, lon] = coords;
displayName = location;
} else {
const geo = await geocodeCity(location);
if (!geo) return await locationNotFoundError(location, 'weather_openmeteo');
({ lat, lon, displayName } = geo);
}
const fetchOneCity = async (lat: number, lon: number, displayName: string) => {
const params = new URLSearchParams({
latitude: String(lat),
longitude: String(lon),
@@ -717,7 +706,6 @@ export const weatherOpenMeteoTool = {
const res = await fetch(`https://api.open-meteo.com/v1/forecast?${params}`, { signal: AbortSignal.timeout(15_000) });
if (!res.ok) throw new Error(`Open-Meteo API error ${res.status}`);
const data: any = await res.json();
const text = formatOpenMeteoReport({
displayName,
timezone: data.timezone || 'auto',
@@ -735,7 +723,47 @@ export const weatherOpenMeteoTool = {
return String(v);
},
});
return { success: true, stdout: text, data: { location: displayName, lat, lon, raw: data } };
return { text, lat, lon, raw: data };
};
try {
const coords = parseCoords(location);
if (coords) {
const [lat, lon] = coords;
const { text, raw } = await fetchOneCity(lat, lon, location);
return { success: true, stdout: text, data: { location, lat, lon, raw } };
}
const geo = await geocodeCity(location);
if (geo) {
const { text, raw } = await fetchOneCity(geo.lat, geo.lon, geo.displayName);
return { success: true, stdout: text, data: { location: geo.displayName, lat: geo.lat, lon: geo.lon, raw } };
}
// 2026-08-10: the error path below already tells the model exactly how to retry
// ("call once per city"), but a real incident showed that guidance only gets partially
// followed — asked for 5 cities, it retried 2 and silently dropped 3 (no fabrication,
// just incomplete: it said "let me know if you need the others" instead of finishing the
// list it had already committed to). Rather than trust N follow-up calls to all happen,
// fetch every city here and return one combined report — there is no follow-up to skip.
const cities = await looksLikeMultipleLocations(location);
if (!cities) return await locationNotFoundError(location, 'weather_openmeteo');
const settled = await Promise.allSettled(cities.map(async (c) => {
const g = await geocodeCity(c);
if (!g) throw new Error(`${c}: 위치를 찾을 수 없음`);
const { text } = await fetchOneCity(g.lat, g.lon, g.displayName);
return { city: c, text };
}));
const ok = settled.filter((r): r is PromiseFulfilledResult<{ city: string; text: string }> => r.status === 'fulfilled').map(r => r.value);
const failed = settled
.map((r, i) => (r.status === 'rejected' ? cities[i] : null))
.filter((c): c is string => c !== null);
if (!ok.length) return { success: false, error: `${cities.length}개 도시 전부 조회 실패: ${cities.join(', ')}` };
const combined = ok.map(r => r.text).join('\n\n' + '─'.repeat(40) + '\n\n')
+ (failed.length ? `\n\n(다음 도시는 조회 실패: ${failed.join(', ')})` : '');
return { success: true, stdout: combined, data: { locations: ok.map(r => r.city), failed } };
} catch (err: any) {
return { success: false, error: err.message };
}
+39 -19
View File
@@ -1,16 +1,24 @@
/**
* weather-multi-location.test.ts
*
* What a weather tool says when `location` holds several cities instead of one.
* What weather_openmeteo does when `location` holds several cities instead of one.
*
* Asked to compare cities, models batch them into a single argument
* ("London, Berlin, Paris, Rome, Madrid"). That geocodes to nothing, and the old error —
* a bare "위치를 찾을 수 없습니다" — gave the model no reason to do anything but reissue the
* identical call, so the turn looped until it ran out of rounds (observed live 2026-08-09 09:09).
* ("London, Berlin, Paris, Rome, Madrid"). That geocodes to nothing.
*
* The property under test is not "does it fail" but "does the failure tell the model how to
* recover", plus its mirror: inputs that are legitimately one place must not be mistaken for
* a list. These tests hit the real Open-Meteo geocoding API, so they are skipped when it is
* First fix (2026-08-09): a bare "위치를 찾을 수 없습니다" gave the model no reason to do
* anything but reissue the identical call, so the turn looped until it ran out of rounds
* (observed live 09:09). Rewrote the error to name every city and tell the model to call once
* per city instead.
*
* That undershot too (2026-08-10, real incident): asked for 5 cities, the model retried only 2
* after the improved error and silently dropped the other 3 — no fabrication, just an
* incomplete answer ("다른 도시가 더 필요하시면 말씀해 주세요" instead of finishing the list it
* had already committed to). Prose telling the model what to do kept being followed partially,
* not completely, so weather_openmeteo now fetches every city itself on a detected batch and
* returns one combined report — there's no follow-up call left for the model to skip.
*
* These tests hit the real Open-Meteo geocoding/forecast API, so they are skipped when it is
* unreachable rather than failing the suite on a network blip.
*/
@@ -22,6 +30,8 @@ const ask = (location: string) =>
weatherOpenMeteoTool.execute({ location, variables: 'temperature_2m_max' }) as Promise<{
success: boolean;
error?: string;
stdout?: string;
data?: { locations?: string[]; failed?: string[] };
}>;
let online = false;
@@ -37,23 +47,33 @@ before(async () => {
if (!online) console.log(' (geocoding API 접속 불가 — 네트워크 의존 테스트 건너뜀)');
});
describe('배치 입력은 복구 방법을 알려준다', () => {
test('여러 도시를 한 번에 넘기면 도구별 호출을 지시한다', async (t) => {
describe('배치 입력은 도구가 직접 전부 조회해 하나로 합친다', () => {
test('여러 도시를 한 번에 넘기면 전부 조회해서 하나의 결과로 합친다', async (t) => {
if (!online) return t.skip('offline');
const r = await ask('London, Berlin, Paris, Rome, Madrid');
assert.equal(r.success, false);
const err = r.error ?? '';
// 무엇이 잘못됐는지 + 어떻게 고치는지 + 재시도가 무의미하다는 것, 세 가지가 다 있어야
// 모델이 같은 호출을 반복하지 않는다.
assert.match(err, /한 번에 한 곳만/, '단일 위치 제약을 밝혀야 함');
assert.match(err, /weather_openmeteo/, '어떤 도구를 다시 부를지 알려야 함');
assert.match(err, /"location":"London"/, '구체적인 재호출 예시가 있어야 함');
assert.match(err, /다시 호출하면 똑같이 실패/, '동일 재시도가 무의미함을 밝혀야 함');
// 파싱한 도시를 전부 되돌려줘야 모델이 남은 도시를 안 빠뜨린다.
// 2026-08-10 이전엔 여기서 success=false + 도시별로 다시 부르라는 안내였다. 그 안내를
// 모델이 부분적으로만 따라서(5개 중 2개만 재시도) 나온 실제 사고 이후, 도구가 직접 전부
// 조회하도록 바꿨다 — 모델이 빠뜨릴 재시도 자체가 없어야 한다.
assert.equal(r.success, true);
const stdout = r.stdout ?? '';
for (const city of ['London', 'Berlin', 'Paris', 'Rome', 'Madrid']) {
assert.ok(err.includes(city), `${city}가 목록에 없음`);
assert.ok(r.data?.locations?.some(l => l.includes(city)) || stdout.includes(city), `${city} 결과가 빠짐`);
}
// 5개 도시 결과가 구분 없이 뒤섞이면 안 되므로 구분선이 있어야 함.
assert.ok(stdout.includes('─'), '도시별 구분선이 없음');
});
test('일부 도시만 실패해도 나머지는 정상 반환된다', async (t) => {
if (!online) return t.skip('offline');
// looksLikeMultipleLocations는 배치 판별을 위해 첫 두 세그먼트가 각각 지오코딩되는지만
// 본다(그래야 "Springfield, Freedonia" 같은 단일 지명+미확인 수식어를 목록으로 오인하지
// 않음) — 그래서 실패할 도시는 세 번째 자리에 둔다.
const r = await ask('London, Paris, Zzxqwv');
assert.equal(r.success, true, '일부 실패로 전체가 실패하면 안 됨');
assert.ok(r.data?.failed?.some(f => f.includes('Zzxqwv')), '실패한 도시가 failed 목록에 있어야 함');
assert.ok(r.data?.locations?.some(l => l.includes('London')));
assert.ok(r.data?.locations?.some(l => l.includes('Paris')));
});
});