fix: 임베딩 모델 콜드로드로 벡터메모리 회상이 타임아웃되던 문제

Ollama 재시작(예: 버전 업그레이드) 후 embeddinggemma가 언로드되면
다음 채팅의 buildPersonalityContext에서 콜드 로드(~30s)가
15초 타임아웃을 넘겨 "vector memory query failed ... aborted due to
timeout" 경고가 났다. 답변엔 영향 없지만 회상이 빠지고 첫 응답이
15초 지연됐다.

- EMBED_KEEP_ALIVE '24h' → -1 (영구 상주). embeddinggemma가 유일한
  로컬 모델이고 ~0.6GB라 안전. ollama.service의 OLLAMA_KEEP_ALIVE=-1
  drop-in과 이중 방어.
- warmupEmbedding() 추가 — 게이트웨이 부팅 시 백그라운드 예열해서
  첫 채팅이 콜드 로드를 안 기다리게.
- embedText에 timeoutMs 파라미터 추가(기본 15s 유지, 워밍업만 60s).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gh45CPB94UFQFiHov2CQe7
This commit is contained in:
kim
2026-08-27 13:11:44 +09:00
co-authored by Claude Sonnet 5
parent 38320a9162
commit 63bbc36c6e
2 changed files with 26 additions and 3 deletions
+5
View File
@@ -21,6 +21,7 @@ import {
queryMemoryVectors,
embedQuery,
queryVectorsWithEmbedding,
warmupEmbedding,
USER_FACTS_COLLECTION,
DAILY_EXTRACTS_COLLECTION,
} from './memory/memory-vector';
@@ -4805,6 +4806,10 @@ server.listen(PORT, HOST, async () => {
// Auto-connect enabled MCP servers
getMCPManager().startEnabledServers().catch(err => console.warn('[MCP] Startup error:', err?.message));
// Preload the embedding model (background, non-blocking) so the first chat that needs
// vector-memory recall isn't stuck behind a cold model load.
warmupEmbedding();
// GPU voice engine (faster-whisper + OmniVoice) — only if configured as the active provider,
// starts lazily in the background so a slow/failed GPU load never blocks gateway boot.
const voiceCfg = (liveConfig as any).voice;