v4.3.19-26: 환각 방어 전역화 + 답변 길이 압박 완화 + 토큰 로깅 구멍 수정

실사용 로그 감사(07-29 GPU/PCIe 대화)에서 발견한 문제들 일괄 수정.

환각 방어:
- looksLikeUnverifiedSpecClaim이 놓치던 패턴 2개 추가 — 단위 없는
  "15토큰"(strict tier), 퍼센트 벤치마크 "약 5%"(ambiguous tier,
  SPEC_CONTEXT_KEYWORD 게이트로 오탐 방지)
- forceToolChoiceOnLiveData를 모델별 opt-in → 전역 기본 ON + opt-out.
  모델을 수시로 바꾸는데 profiles에 등재된 2개만 방어되던 구멍이 원인
  (프로필 없던 gemini-3.5-flash-lite가 tok/s 날조 + 없는 제품명 창작)

답변 길이:
- "Keep responses SHORT" 1-2 → 2-4문장, 대화/실행 답변으로 범위 한정.
  정보성 답변(뉴스·날씨·설명·비교)은 OUTPUT FORMAT을 따르도록 명시
- 불확실성 표현은 길이 제한 면제. 1-2문장 압박이 "모른다"고 말할 자리를
  없애 환각을 확신형 한 줄로 포장하던 것이 실측으로 확인됨
- 뉴스는 도구가 반환한 기사 전부 표기(3건 체리픽 금지, 8-10행 목표)

토큰 로깅:
- logUsage가 ollama-adapter 내부 private이라 Ollama 경유만 기록되고,
  openai-compat은 usage 필드를 추출조차 안 하고 버리던 문제. 비용 비교
  대상인 Gemini가 통째로 누락돼 공짜처럼 보이는 상태였음
- providers/usage-log.ts로 공용화 → openai-compat(chat/generate)과
  openai-codex(SSE response.completed)에서도 캡처·로깅·호출자 반환

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
kim
2026-07-29 14:44:21 +09:00
co-authored by Claude Opus 5
parent 3bff465c08
commit 8359901cac
7 changed files with 105 additions and 30 deletions
+14 -6
View File
@@ -1808,18 +1808,26 @@ function getModelProfileExtraPrompt(modelName: string): string {
// alone doesn't fix (verified 2026-07-25: mistral-large-3 still hallucinated a full news answer
// — see extraSystemPrompt above — despite the prompt-level warning). tool_choice: 'required'
// makes round 0 mechanically unable to answer in plain text when a live-data/factual question
// is detected, instead of just asking nicely. Scoped per-model via config rather than applied
// globally — forcing an unwanted tool call on a gate false-positive is worse than today's no-op
// for models that don't actually have this problem (verified 07-18: qwen3.5 1 hallucination vs
// mistral's 7 in the same 20-question comparison).
// is detected, instead of just asking nicely.
//
// 2026-07-29: INVERTED from per-model opt-in to global default with per-model opt-out. The
// original opt-in reasoning ("forcing an unwanted tool call on a gate false-positive is worse
// than a no-op for models that don't have this problem") assumed a stable model choice, but
// the user switches the active model frequently — so in practice only the two profiled models
// (mistral-large-3, kimi-k2.6) ever had this defense, and every other model ran with the
// round-0 gap wide open. Log audit of 2026-07-29 confirmed the cost: gemini-3.5-flash-lite,
// which has no profile entry at all, fabricated GPU tok/s figures and an entirely nonexistent
// product name across a long hardware chat. A spurious search on a gate false-positive is a
// few wasted seconds; an unguarded fabrication is a wrong answer the user acts on. Set
// models.profiles[<model>].forceToolChoiceOnLiveData = false to opt a specific model out.
function getModelProfileForceToolChoice(modelName: string): boolean {
const m = String(modelName || '').trim();
if (!m) return false;
try {
const raw = getConfig().getConfig() as any;
return raw?.models?.profiles?.[m]?.forceToolChoiceOnLiveData === true;
return raw?.models?.profiles?.[m]?.forceToolChoiceOnLiveData !== false;
} catch {
return false;
return true;
}
}