v4.3.19-26: 환각 방어 전역화 + 답변 길이 압박 완화 + 토큰 로깅 구멍 수정
실사용 로그 감사(07-29 GPU/PCIe 대화)에서 발견한 문제들 일괄 수정. 환각 방어: - looksLikeUnverifiedSpecClaim이 놓치던 패턴 2개 추가 — 단위 없는 "15토큰"(strict tier), 퍼센트 벤치마크 "약 5%"(ambiguous tier, SPEC_CONTEXT_KEYWORD 게이트로 오탐 방지) - forceToolChoiceOnLiveData를 모델별 opt-in → 전역 기본 ON + opt-out. 모델을 수시로 바꾸는데 profiles에 등재된 2개만 방어되던 구멍이 원인 (프로필 없던 gemini-3.5-flash-lite가 tok/s 날조 + 없는 제품명 창작) 답변 길이: - "Keep responses SHORT" 1-2 → 2-4문장, 대화/실행 답변으로 범위 한정. 정보성 답변(뉴스·날씨·설명·비교)은 OUTPUT FORMAT을 따르도록 명시 - 불확실성 표현은 길이 제한 면제. 1-2문장 압박이 "모른다"고 말할 자리를 없애 환각을 확신형 한 줄로 포장하던 것이 실측으로 확인됨 - 뉴스는 도구가 반환한 기사 전부 표기(3건 체리픽 금지, 8-10행 목표) 토큰 로깅: - logUsage가 ollama-adapter 내부 private이라 Ollama 경유만 기록되고, openai-compat은 usage 필드를 추출조차 안 하고 버리던 문제. 비용 비교 대상인 Gemini가 통째로 누락돼 공짜처럼 보이는 상태였음 - providers/usage-log.ts로 공용화 → openai-compat(chat/generate)과 openai-codex(SSE response.completed)에서도 캡처·로깅·호출자 반환 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
+14
-6
@@ -1808,18 +1808,26 @@ function getModelProfileExtraPrompt(modelName: string): string {
|
||||
// alone doesn't fix (verified 2026-07-25: mistral-large-3 still hallucinated a full news answer
|
||||
// — see extraSystemPrompt above — despite the prompt-level warning). tool_choice: 'required'
|
||||
// makes round 0 mechanically unable to answer in plain text when a live-data/factual question
|
||||
// is detected, instead of just asking nicely. Scoped per-model via config rather than applied
|
||||
// globally — forcing an unwanted tool call on a gate false-positive is worse than today's no-op
|
||||
// for models that don't actually have this problem (verified 07-18: qwen3.5 1 hallucination vs
|
||||
// mistral's 7 in the same 20-question comparison).
|
||||
// is detected, instead of just asking nicely.
|
||||
//
|
||||
// 2026-07-29: INVERTED from per-model opt-in to global default with per-model opt-out. The
|
||||
// original opt-in reasoning ("forcing an unwanted tool call on a gate false-positive is worse
|
||||
// than a no-op for models that don't have this problem") assumed a stable model choice, but
|
||||
// the user switches the active model frequently — so in practice only the two profiled models
|
||||
// (mistral-large-3, kimi-k2.6) ever had this defense, and every other model ran with the
|
||||
// round-0 gap wide open. Log audit of 2026-07-29 confirmed the cost: gemini-3.5-flash-lite,
|
||||
// which has no profile entry at all, fabricated GPU tok/s figures and an entirely nonexistent
|
||||
// product name across a long hardware chat. A spurious search on a gate false-positive is a
|
||||
// few wasted seconds; an unguarded fabrication is a wrong answer the user acts on. Set
|
||||
// models.profiles[<model>].forceToolChoiceOnLiveData = false to opt a specific model out.
|
||||
function getModelProfileForceToolChoice(modelName: string): boolean {
|
||||
const m = String(modelName || '').trim();
|
||||
if (!m) return false;
|
||||
try {
|
||||
const raw = getConfig().getConfig() as any;
|
||||
return raw?.models?.profiles?.[m]?.forceToolChoiceOnLiveData === true;
|
||||
return raw?.models?.profiles?.[m]?.forceToolChoiceOnLiveData !== false;
|
||||
} catch {
|
||||
return false;
|
||||
return true;
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user