feat: 모델별 num_ctx 고정 옵션 추가 — 지서버 재로드 문제 대응

사용자 관찰("크기가 바뀔 때마다 올라마가 모델을 다시 로드하네")을
직접 재현·측정: 같은 num_ctx로 두 번 호출하면 load_duration 1.3초,
num_ctx만 바꿔 호출하면 18.8초. ollama-adapter.ts의 _resolveCtx()가
매 턴 프롬프트 크기에 맞춰 2의 거듭제곱으로 동적으로 재계산하는데,
그 값이 바뀔 때마다 Ollama가 모델을 통째로 재로드하는 것 — 대화
도중 설명 안 되는 수 초짜리 멈춤으로 나타났다.

동적 사이징 자체는 VRAM이 빠듯한 배포를 위해 존재하는데, 지서버는
그 경우가 아니다: gemma4:26b를 native 최대 컨텍스트(262144)로
로드해도 32GB 중 18.3GB만 쓰고 13.7GB가 남는다(실측, size_vram).
아낄 필요가 없는데 아끼다가 재로드 비용만 치르고 있었다.

models.profiles[<model>].fixedNumCtx로 모델별로 동적 사이징을
끄고 항상 고정값을 요청하게 하는 옵션을 추가했다(getModelProfileForceToolChoice와
같은 패턴). gemma4:26b는 262144로 설정 — 한 번 로드되면 대화가
아무리 길어져도 재로드가 없다. 다른 모델/제공자는 기존 동적
사이징 그대로 유지된다(opt-in이라 이 모델 외엔 영향 없음).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
kim
2026-08-10 17:30:43 +09:00
co-authored by Claude Opus 5
parent 6c02050db6
commit 1a77efd106
2 changed files with 37 additions and 2 deletions
+24
View File
@@ -1861,6 +1861,29 @@ function getModelProfileForceToolChoice(modelName: string): boolean {
}
}
// Same profile mechanism, for opting a specific model OUT of dynamic num_ctx sizing
// (ollama-adapter.ts's _resolveCtx(), which rounds to the next power of 2 based on prompt size).
// Dynamic sizing exists to avoid reserving VRAM a smaller GPU doesn't have — but 2026-08-10,
// measured directly against 지서버's live Ollama instance: gemma4:26b at its native max context
// (262144) loads in 18.3GB, leaving 13.7GB free on its 32GB (2×5060 Ti). It has no VRAM pressure
// to economize for. The cost of dynamic sizing turned out to be real and measured too — every
// time the resolved num_ctx crosses a power-of-2 boundary between turns, Ollama does a full
// model reload before answering (verified: same num_ctx back-to-back = 1.3s load_duration,
// num_ctx changed = 18.8s), which shows up mid-conversation as unexplained multi-second stalls.
// Set models.profiles[<model>].fixedNumCtx to a number to always request exactly that num_ctx
// for that model — loads once, never reloads for a size change again.
function getModelProfileFixedNumCtx(modelName: string): number | undefined {
const m = String(modelName || '').trim();
if (!m) return undefined;
try {
const raw = getConfig().getConfig() as any;
const v = Number(raw?.models?.profiles?.[m]?.fixedNumCtx);
return Number.isFinite(v) && v > 0 ? v : undefined;
} catch {
return undefined;
}
}
function resolveWorkspaceFilePath(workspacePath: string, filename: string): string {
if (!filename) return '';
@@ -2121,6 +2144,7 @@ const { handleChat, handleCodeChat } = createHandleChat({
normalizeToolArgs,
goalLikelyNeedsTextInput,
getModelProfileForceToolChoice,
getModelProfileFixedNumCtx,
stripExplicitThinkTags,
separateThinkingFromContent,
isOrchestrationSkillEnabled,