fix: 종료된 deepseek-v4-flash:cloud → gemma4:31b-cloud (기억 재랭킹·Google 폴백)

Ollama Cloud가 2026-09-25에 deepseek-v4-flash를 종료해 재랭킹이 HTTP 410으로 실패하고
거리 임계값 폴백만 쓰이던 문제.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
kim
2026-09-26 15:47:43 +09:00
co-authored by Claude Sonnet 5
parent 03a462d72e
commit 6626132eaf
2 changed files with 2 additions and 2 deletions
+1 -1
View File
@@ -14,7 +14,7 @@ import { getConfig } from '../../config/config.js';
// "flash"-tier model chosen specifically for latency, not quality — this call is a small
// classification task (pick relevant indices from a short list), not a task worth spending
// primary-model-grade reasoning time on. Revisit if this model is ever unpulled/renamed.
const RERANK_MODEL = 'deepseek-v4-flash:cloud';
const RERANK_MODEL = 'gemma4:31b-cloud';
// Only fires on genuinely ambiguous retrieval (see the gating in personality-context.ts) —
// most turns never reach this call at all, so a generous timeout here doesn't cost anything
// on the common path. 8s was measured to occasionally clip real (successful, just slow)
+1 -1
View File
@@ -28,7 +28,7 @@ import { isGoogleNearLimit } from './google-usage';
// reactive to a 429 (see google-usage.ts). Hardcoded rather than reusing
// models.fallback: that field is a separate, more general "provider errored
// mid-turn" mechanism (see handle-chat.ts) and may be tuned independently.
const GOOGLE_FALLBACK_MODEL = 'deepseek-v4-flash:cloud';
const GOOGLE_FALLBACK_MODEL = 'gemma4:31b-cloud';
const LEGACY_BLOCKED_MODELS = new Set(['codex-davinci-002']);
const DEFAULT_OPENAI_MODEL = 'gpt-4o';