feat: GPU 스펙 구조화 DB를 검색 결과에 자동 주입

"시어엔진에 이런 데이터 전용 API가 있나" → 253개 엔진 중 하드웨어 스펙
엔진은 0개였다(geizhals는 403 봇차단, ddg definitions는 GPU 모델명에 빈
응답). 대신 GitHub에서 RightNow-AI/RightNow-GPU-Database(Apache 2.0,
TechPowerUp 기반, 2,824개 GPU, 55필드)를 찾았고, README엔 언급이 없지만
실물엔 RTX 50 시리즈가 들어있는 것을 데이터를 받아 확인했다.

이게 지금까지의 우회로보다 나은 이유:
- nvidia.com 공식 페이지는 스펙 키워드 0개(JS 렌더링), techpowerup은
  봇 차단. HTML 스크래핑으로 살아있는 출처가 nanoreview 하나뿐이었다
- 봇차단·JS·파싱실패·죽은링크가 전부 없고, 캐시 워밍 후 64ms
- 무엇보다 모든 레코드에 releaseDate가 있다. 5060은 2025-05-19다.
  이 사태의 출발점이던 "5060은 미출시" 환각은 데이터와 만나는 순간
  성립하지 않는다 — 가드는 쓰인 거짓말을 잡지만, 이건 쓸 이유를 없앤다

신선도가 진짜 리스크라 거기에 설계를 집중했다. 호스팅 API가 아니라 GitHub
저장소라서, 저쪽이 멈추면 우리도 멈추고 그러면 지금 고치는 실패가 데이터
계층에서 재현된다. 그래서 주간 갱신 + 결과에 캐시 나이를 항상 같이 실어
낡은 답이 조용히 틀리는 대신 눈에 띄게 낡도록 했다. 갱신은 백그라운드라
질문을 느리게 만들지 않고, 실패해도 옛 캐시로 계속 답한다.

도구가 아니라 주입으로 넣은 이유: 모델이 도구를 안 부른다. 로그 3,300줄
에서 web_fetch 호출 1번, GPU 질문엔 0번이었다(오늘 같은 결론 네 번째).

- extractGpuMentions: 맥락이 있을 때만 맨숫자를 모델명으로 본다.
  "4080이 5060보다"는 잡고 "2024년 매출 3800억"은 안 잡는다 — 무관한
  답변에 스펙이 끼어들면 그 자체가 오염이다
- findGpu: "RTX 4080"은 SUPER/Mobile 이름에도 부분일치하므로 변형에
  가중치를 줘 기본 카드를 고른다. 데스크탑 질문에 노트북 칩 수치를
  물려주면 조용히 틀린 답이 된다
- 캐시 쓰기는 write-then-rename (config.json 3회 손상과 같은 실패 방지)

실측: RTX 4080 FP32 48.74 TFLOPS / 716.8 GB/s vs RTX 5060 19.18 TFLOPS /
448 GB/s — 2.54배 차이를 근거를 갖고 말할 수 있게 됐다.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
kim
2026-08-11 14:34:34 +09:00
co-authored by Claude Opus 5
parent de5407b701
commit 231ecbfce6
3 changed files with 402 additions and 3 deletions
+258
View File
@@ -0,0 +1,258 @@
/**
* gpu-specs.ts
*
* Structured GPU specifications, looked up from a local cache instead of scraped from the web.
*
* WHY THIS EXISTS. Measured 2026-08-11 while chasing a bad "RTX 4080 vs RTX 5060" answer: the
* sources a web search surfaces for a spec question are mostly unusable to us.
* nvidia.com (the manufacturer!) 6016 chars fetched, ZERO spec keywords — the spec table is
* JS-rendered, so the body text is marketing copy
* techpowerup.com 275 chars — bot wall
* nanoreview.net works, and is what the comparison-split path fetches today
* One working source reached by HTML scraping is a thin margin. This path has none of those
* failure modes: no bot wall, no JS, no parse, no dead link, and no 9-second fetch.
*
* AND IT KILLS THE ORIGINAL BUG AT THE ROOT. The failure that started all of this was the model
* asserting "RTX 5060은 아직 출시되지 않은 차세대 모델" — a knowledge-cutoff artifact. Every record
* here carries a releaseDate (the 5060's is 2025-05-19), so the claim cannot survive contact with
* the data. That is a stronger fix than any guard: the guard catches the lie after it is written,
* this makes it not worth writing.
*
* STALENESS IS THE REAL RISK, so it gets the design attention. The source is a GitHub repo, not a
* hosted API — if it stops being updated, our data freezes, and a frozen spec table reproduces the
* exact "that card doesn't exist" failure one layer down. Hence: refresh weekly, and always report
* the cache's age alongside the data so a stale answer is visibly stale rather than silently wrong.
*
* Source: https://github.com/RightNow-AI/RightNow-GPU-Database (Apache 2.0), 2,824 GPUs across
* nvidia/amd/intel (plus some historical 3dfx/matrox/xgi), ~55 fields each, derived from
* TechPowerUp's database.
*/
import fs from 'fs';
import path from 'path';
import os from 'os';
const SOURCE_URL = 'https://raw.githubusercontent.com/RightNow-AI/RightNow-GPU-Database/main/data/all-gpus.json';
const CACHE_PATH = path.join(os.homedir(), '.smallclaw', 'gpu-specs-cache.json');
const MAX_AGE_MS = 7 * 24 * 60 * 60 * 1000;
const FETCH_TIMEOUT_MS = 30_000;
export interface GpuRecord {
name?: string;
vendor?: string;
architecture?: string;
releaseDate?: string;
processSize?: number;
shaders?: number;
tensorCores?: number;
baseClock?: number;
boostClock?: number;
memorySize?: number;
memoryType?: string;
memoryBus?: number;
memoryBandwidth?: number;
fp32?: number;
tdp?: number;
url?: string;
[k: string]: any;
}
interface CacheFile { fetchedAt: number; gpus: GpuRecord[] }
let memoryCache: CacheFile | null = null;
let refreshInFlight: Promise<void> | null = null;
function readCacheFile(): CacheFile | null {
try {
const raw = JSON.parse(fs.readFileSync(CACHE_PATH, 'utf8'));
if (Array.isArray(raw?.gpus) && raw.gpus.length) return raw as CacheFile;
} catch { /* missing or corrupt cache is the same as no cache */ }
return null;
}
async function downloadToCache(): Promise<CacheFile | null> {
try {
const res = await fetch(SOURCE_URL, { signal: AbortSignal.timeout(FETCH_TIMEOUT_MS) });
if (!res.ok) return null;
const gpus = await res.json() as GpuRecord[];
if (!Array.isArray(gpus) || gpus.length < 100) return null; // truncated/garbage guard
const payload: CacheFile = { fetchedAt: Date.now(), gpus };
fs.mkdirSync(path.dirname(CACHE_PATH), { recursive: true });
// Write-then-rename: a half-written cache read by the next request would look like corruption
// and silently disable the whole lookup ([[project_config_json_corruption]]).
const tmp = `${CACHE_PATH}.tmp-${process.pid}`;
fs.writeFileSync(tmp, JSON.stringify(payload));
fs.renameSync(tmp, CACHE_PATH);
console.log(`[gpu-specs] 데이터셋 갱신 완료 — ${gpus.length}개 GPU`);
return payload;
} catch (err: any) {
console.log(`[gpu-specs] 갱신 실패 (${err?.message || err}) — 기존 캐시를 계속 사용합니다`);
return null;
}
}
/**
* Serves whatever is on disk immediately and refreshes in the background when it is over a week
* old. A weekly refresh must never make a user's question slower, and a network failure must never
* make it fail — stale data beats no data, as long as its age is disclosed (see formatGpuSpecs).
*/
async function loadGpus(): Promise<CacheFile | null> {
if (!memoryCache) memoryCache = readCacheFile();
if (!memoryCache) {
const fresh = await downloadToCache();
if (fresh) memoryCache = fresh;
return memoryCache;
}
if (Date.now() - memoryCache.fetchedAt > MAX_AGE_MS && !refreshInFlight) {
refreshInFlight = downloadToCache()
.then(fresh => { if (fresh) memoryCache = fresh; })
.finally(() => { refreshInFlight = null; });
}
return memoryCache;
}
// ── Model-name matching ───────────────────────────────────────────────────────
/** "GeForce RTX 4080 SUPER" → "rtx 4080 super" — vendor prefixes carry no distinguishing info. */
function normName(name: string): string {
return String(name || '')
.toLowerCase()
.replace(/\b(geforce|nvidia|radeon|amd|intel)\b/g, ' ')
.replace(/\s+/g, ' ')
.trim();
}
export interface GpuMention { series: string; number: string; suffix: string; raw: string }
const SERIES_MODEL = /\b(rtx|gtx|rx|arc)\s*([a-z]?\d{3,4})\s*((?:ti|super|xtx|xt)\b(?:\s*super\b)?)?/gi;
/** Words that make a bare "4080" readable as a graphics card rather than a year or a price. */
const GPU_CONTEXT = /\b(gpu|graphics\s*card|vga|rtx|gtx|radeon|geforce|nvidia)\b|그래픽\s*카드|비디오\s*카드|그래픽카드/i;
const BARE_MODEL = /\b([1-9]\d{3})\b/g;
/**
* Finds the GPU models named in a query. Bare numbers only count when the text also establishes
* we are talking about graphics cards at all — otherwise "2024년 실적" and "1080p 영상" would both
* read as model numbers, and the injected specs would be pure noise in an unrelated answer.
*/
export function extractGpuMentions(text: string): GpuMention[] {
const s = String(text || '');
const out: GpuMention[] = [];
const seen = new Set<string>();
const add = (series: string, number: string, suffix: string, raw: string) => {
const key = `${series} ${number} ${suffix}`.toLowerCase().replace(/\s+/g, ' ').trim();
if (seen.has(key)) return;
seen.add(key);
out.push({ series: series.toLowerCase(), number: number.toLowerCase(), suffix: suffix.trim().toLowerCase(), raw });
};
for (const m of s.matchAll(SERIES_MODEL)) add(m[1], m[2], m[3] || '', m[0].trim());
if (GPU_CONTEXT.test(s)) {
// "4080이 5060보다 빠른가?" — the series word appears once (or not at all) while the numbers
// are written bare, which is how people actually type it.
const series = /\brx\b|radeon/i.test(s) ? 'rx' : /\barc\b/i.test(s) ? 'arc' : 'rtx';
for (const m of s.matchAll(BARE_MODEL)) {
if (out.some(o => o.number === m[1])) continue;
// A 4-digit number that is plausibly a year is more likely a year than a card.
if (/^(19|20)\d{2}$/.test(m[1])) continue;
add(series, m[1], '', m[0]);
}
}
return out;
}
/**
* Picks the one card a mention means. "RTX 4080" substring-matches the SUPER and the Mobile
* variants too, and handing the model three near-identical spec blocks invites it to quote the
* laptop chip's numbers for a desktop question — so variants lose to the plain card unless the
* query asked for them.
*/
export function findGpu(mention: GpuMention, gpus: readonly GpuRecord[]): GpuRecord | null {
const stem = `${mention.series} ${mention.number}`;
const want = `${stem}${mention.suffix ? ' ' + mention.suffix : ''}`;
const cands = gpus.filter(g => normName(g.name || '').includes(stem));
if (!cands.length) return null;
const variant = /\b(mobile|max-q|laptop|oem|d v2|\bd\b)\b/;
const scored = cands.map(g => {
const n = normName(g.name || '');
let score = n.length; // tie-break toward the plainest name
if (n === want) score -= 1000;
if (variant.test(n) !== variant.test(want)) score += 500;
if (/\bti\b/.test(n) !== /\bti\b/.test(want)) score += 200;
if (/\bsuper\b/.test(n) !== /\bsuper\b/.test(want)) score += 200;
if (/\bxtx?\b/.test(n) !== /\bxtx?\b/.test(want)) score += 200;
return { g, score };
}).sort((a, b) => a.score - b.score);
return scored[0].g;
}
// ── Formatting ────────────────────────────────────────────────────────────────
const num = (v: any) => (typeof v === 'number' && Number.isFinite(v) ? String(v) : null);
/** One compact block per card. Only decision-relevant fields — all 55 would drown the answer. */
export function formatGpuRecord(g: GpuRecord): string {
const lines: string[] = [];
const head = [g.name, g.vendor ? `(${g.vendor}` : null].filter(Boolean).join(' ');
lines.push(`${head}${g.releaseDate ? `, 출시 ${g.releaseDate})` : g.vendor ? ')' : ''}`);
const arch = [g.architecture, num(g.processSize) ? `${g.processSize}nm` : null].filter(Boolean).join(' · ');
if (arch) lines.push(` 아키텍처: ${arch}`);
const cores = [
num(g.shaders) ? `셰이딩 유닛 ${g.shaders}` : null,
num(g.tensorCores) ? `텐서 코어 ${g.tensorCores}` : null,
].filter(Boolean).join(' · ');
if (cores) lines.push(` ${cores}`);
const clock = [num(g.baseClock) ? `베이스 ${g.baseClock}MHz` : null, num(g.boostClock) ? `부스트 ${g.boostClock}MHz` : null]
.filter(Boolean).join(' → ');
if (clock) lines.push(` 클럭: ${clock}`);
const mem = [
num(g.memorySize) ? `${g.memorySize}GB` : null,
g.memoryType || null,
num(g.memoryBus) ? `${g.memoryBus}bit` : null,
num(g.memoryBandwidth) ? `대역폭 ${g.memoryBandwidth} GB/s` : null,
].filter(Boolean).join(' ');
if (mem) lines.push(` 메모리: ${mem}`);
const perf = [num(g.fp32) ? `FP32 ${g.fp32} TFLOPS` : null, num(g.tdp) ? `TDP ${g.tdp}W` : null]
.filter(Boolean).join(' · ');
if (perf) lines.push(` ${perf}`);
if (g.url) lines.push(` 출처: ${g.url}`);
return lines.join('\n');
}
const ageLabel = (fetchedAt: number): string => {
const days = Math.floor((Date.now() - fetchedAt) / (24 * 60 * 60 * 1000));
return days <= 0 ? '오늘 받음' : `${days}일 전 받음`;
};
/**
* The block appended to a search result, or null when the query named no GPU we have data for.
* The age line is not decoration: it is how a reader can tell a missing new card from a wrong one.
*/
export async function lookupGpuSpecs(query: string): Promise<string | null> {
const mentions = extractGpuMentions(query);
if (!mentions.length) return null;
const cache = await loadGpus();
if (!cache) return null;
const found = mentions
.map(m => findGpu(m, cache.gpus))
.filter((g): g is GpuRecord => !!g);
if (!found.length) return null;
// Dedupe: "RTX 4080" and a bare "4080" in the same query resolve to the same card.
const unique = Array.from(new Map(found.map(g => [g.name, g])).values()).slice(0, 4);
return `[GPU 스펙 DB — TechPowerUp 기반 구조화 데이터, ${ageLabel(cache.fetchedAt)}]\n`
+ `아래 수치는 검색 결과가 아니라 스펙 데이터베이스에서 온 확정값입니다. 비교 수치를 낼 땐 이걸 쓰세요.\n\n`
+ unique.map(formatGpuRecord).join('\n\n');
}
+38 -3
View File
@@ -1,5 +1,6 @@
import { ToolResult } from '../types.js';
import { getConfig } from '../config/config.js';
import { lookupGpuSpecs } from './gpu-specs.js';
type SearchResultItem = { title: string; url: string; snippet: string };
@@ -902,11 +903,45 @@ export function splitComparisonQuery(query: string): ComparisonSplit | null {
}
// ── Main web_search tool ──────────────────────────────────────────────────────
//
// The outer call is where the query is still the user-facing question: it decides whether to split
// a comparison and whether GPU specs apply, then delegates. Inner calls carry _noSplit so those
// decisions are made exactly once per turn rather than once per provider round-trip.
export async function executeWebSearch(args: { query: string; max_results?: number; _noSplit?: boolean }): Promise<ToolResult> {
if (!args._noSplit) {
const split = splitComparisonQuery(args.query || '');
if (split) return runComparisonSearch(args, split);
if (args._noSplit) return runSearchProviders(args);
const split = splitComparisonQuery(args.query || '');
const base = split ? await runComparisonSearch(args, split) : await runSearchProviders(args);
return withGpuSpecs(args.query, base);
}
/**
* Prepends structured GPU specifications when the query named a card we have data for.
*
* Injected rather than exposed as a tool the model must decide to call: across a 3,300-line
* production log (2026-08-11) the model called web_fetch exactly once and never reached for it on
* any of the GPU questions in that same log. A tool it does not call is worth nothing, and this is
* the fourth time today the same conclusion was reached about the same model
* ([[feedback_local_model_needs_code_backstop]]).
*/
async function withGpuSpecs(query: string, res: ToolResult): Promise<ToolResult> {
if (!res.success) return res;
try {
const specs = await lookupGpuSpecs(query || '');
if (!specs) return res;
console.log('[v2] web_search: GPU 스펙 DB 주입');
return {
...res,
data: { ...(res.data as any || {}), gpu_specs_injected: true },
stdout: `${specs}\n\n${String(res.stdout || '').trim()}`,
};
} catch {
// Spec injection is an add-on; any failure must leave the search result untouched.
return res;
}
}
async function runSearchProviders(args: { query: string; max_results?: number }): Promise<ToolResult> {
if (!args.query?.trim()) return { success: false, error: 'query is required' };
let limit = Math.min(args.max_results ?? 5, 10);
if (isPriceQuery(args.query)) limit = Math.max(limit, 5);
+106
View File
@@ -0,0 +1,106 @@
/**
* gpu-specs — 모델명 인식과 카드 선택
*
* 순수 함수만 테스트한다. 다운로드/캐시는 네트워크와 디스크에 의존하므로 여기서 다루지 않는다.
*
* 이 테스트가 지키는 두 가지:
* 1) 맨숫자를 언제 모델명으로 볼 것인가 — "4080이 5060보다 빠른가"는 잡아야 하지만
* "2024년 실적"과 "1080p 영상"은 잡으면 안 된다. 무관한 답변에 GPU 스펙이 끼어들면
* 그 자체가 오염이다.
* 2) 변형 모델 선택 — "RTX 4080"은 SUPER/Mobile 이름에도 부분일치하는데, 데스크탑 질문에
* 노트북 칩 수치를 물려주면 조용히 틀린 답이 된다.
*/
import { test, describe } from 'node:test';
import assert from 'node:assert/strict';
import { extractGpuMentions, findGpu, formatGpuRecord, type GpuRecord } from '../src/tools/gpu-specs';
const GPUS: GpuRecord[] = [
{ name: 'GeForce RTX 4080', vendor: 'nvidia', releaseDate: '2022-09-20', architecture: 'Ada Lovelace', processSize: 5, shaders: 9728, tensorCores: 304, baseClock: 2205, boostClock: 2505, memorySize: 16, memoryType: 'GDDR6X', memoryBus: 256, memoryBandwidth: 716.8, fp32: 48.74, tdp: 320, url: 'https://example.invalid/4080' },
{ name: 'GeForce RTX 4080 SUPER', vendor: 'nvidia', releaseDate: '2024-01-31' },
{ name: 'GeForce RTX 4080 Mobile', vendor: 'nvidia', releaseDate: '2023-02-22' },
{ name: 'GeForce RTX 5060', vendor: 'nvidia', releaseDate: '2025-05-19', memoryBandwidth: 448, fp32: 19.18 },
{ name: 'GeForce RTX 5060 Ti 16 GB', vendor: 'nvidia', releaseDate: '2025-04-16' },
{ name: 'GeForce RTX 5060 Mobile', vendor: 'nvidia', releaseDate: '2025-05-01' },
{ name: 'Radeon RX 7900 XTX', vendor: 'amd', releaseDate: '2022-12-13' },
{ name: 'Radeon RX 7900 XT', vendor: 'amd', releaseDate: '2022-12-13' },
];
describe('extractGpuMentions — 모델명 인식', () => {
test('시리즈 표기가 붙은 형태', () => {
const m = extractGpuMentions('RTX 4080 vs RTX 5060 performance comparison specs');
assert.deepEqual(m.map(x => `${x.series} ${x.number}`), ['rtx 4080', 'rtx 5060']);
});
test('접미사(Ti/SUPER/XTX)를 보존한다', () => {
assert.equal(extractGpuMentions('RTX 5060 Ti 리뷰')[0].suffix, 'ti');
assert.equal(extractGpuMentions('RX 7900 XTX 성능')[0].suffix, 'xtx');
});
test('GPU 맥락이 있으면 맨숫자도 모델명으로 본다', () => {
// 사용자가 실제로 친 형태. "RTX"가 한 번만 나오고 나머지는 숫자만 쓴다.
const m = extractGpuMentions('RTX 4080이 5060 보다 얼마나 빠르지?');
assert.deepEqual(m.map(x => x.number).sort(), ['4080', '5060']);
});
test('GPU 맥락이 없으면 맨숫자를 건드리지 않는다', () => {
assert.deepEqual(extractGpuMentions('2024년 매출이 3800억이었다'), []);
assert.deepEqual(extractGpuMentions('아파트 3800세대 분양'), []);
});
test('연도 꼴은 맥락이 있어도 모델명으로 보지 않는다', () => {
const m = extractGpuMentions('이 GPU는 2022년에 나왔다');
assert.deepEqual(m, []);
});
test('중복은 한 번만', () => {
const m = extractGpuMentions('RTX 4080 성능, RTX 4080 가격');
assert.equal(m.length, 1);
});
});
describe('findGpu — 변형 모델에 밀리지 않고 기본 카드를 고른다', () => {
const pick = (q: string) => findGpu(extractGpuMentions(q)[0], GPUS)?.name;
test('"RTX 4080"은 SUPER나 Mobile이 아니라 기본 카드', () => {
assert.equal(pick('RTX 4080 specs'), 'GeForce RTX 4080');
});
test('"RTX 5060"은 Ti나 Mobile이 아니라 기본 카드', () => {
assert.equal(pick('RTX 5060 specs'), 'GeForce RTX 5060');
});
test('접미사를 명시하면 그 변형을 고른다', () => {
assert.equal(pick('RTX 4080 SUPER 스펙'), 'GeForce RTX 4080 SUPER');
assert.equal(pick('RTX 5060 Ti 스펙'), 'GeForce RTX 5060 Ti 16 GB');
});
test('XT와 XTX를 구분한다', () => {
assert.equal(pick('RX 7900 XTX 성능'), 'Radeon RX 7900 XTX');
assert.equal(pick('RX 7900 XT 성능'), 'Radeon RX 7900 XT');
});
test('데이터에 없는 모델은 null', () => {
assert.equal(findGpu({ series: 'rtx', number: '9090', suffix: '', raw: 'RTX 9090' }, GPUS), null);
});
});
describe('formatGpuRecord — 판단에 필요한 필드만', () => {
const out = formatGpuRecord(GPUS[0]);
test('출시일이 들어간다 — "미출시" 환각을 원천 차단하는 필드다', () => {
assert.match(out, /2022-09-20/);
});
test('비교에 쓰이는 수치가 들어간다', () => {
assert.match(out, /716\.8 GB\/s/);
assert.match(out, /FP32 48\.74 TFLOPS/);
assert.match(out, /9728/);
});
test('없는 필드는 빈 줄로 남기지 않는다', () => {
const sparse = formatGpuRecord({ name: 'GeForce RTX 5060', releaseDate: '2025-05-19' });
assert.equal(sparse.split('\n').some(l => l.trim() === ''), false);
assert.match(sparse, /2025-05-19/);
});
});