v4.3.19-5: 컨텍스트 게이지 provider-aware하게 수정 + 캐시 버그 제거

- resolveNumCtx()가 Ollama 전용이라 primary가 Google로 바뀐 뒤 세션
  컨텍스트 게이지가 419%까지 잘못 표시되던 버그 수정 — google provider면
  Gemini API에서 inputTokenLimit을 실시간 조회(하드코딩 아님, API 실패시만
  1048576 폴백)
- 최종 폴백(8192)을 캐시하던 버그 제거 — 재시작 직후 첫 조회가 일시적으로
  실패하면 프로세스 수명 내내 8192에 고정되던 문제, 이제 재시도 가능
- /api/model-context: model 파라미터 없을 때(메인 채팅 사이드바 게이지)
  무조건 200000(클로드 가정) 반환하던 걸 resolveNumCtx() 재사용으로 교체
  — 이제 실제 활성 provider와 항상 동기화됨

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
kim
2026-07-26 17:17:14 +09:00
co-authored by Claude Sonnet 5
parent 5f83803ed4
commit 814f32a63e
2 changed files with 52 additions and 6 deletions
+10 -2
View File
@@ -3494,8 +3494,16 @@ function _readOllamaManifestNumCtx(modelName: string): number | null {
app.get('/api/model-context', async (req, res) => {
const model = String(req.query.model || '').trim().toLowerCase();
// Claude models: 200K (no local metadata)
if (!model || model.includes('claude')) {
// No model specified: caller wants the *current active* model's window (main
// chat sidebar gauge) — reuse resolveNumCtx() so this stays in sync with
// whatever provider is actually primary (ollama/google/etc) instead of the
// old hardcoded 200000 that silently applied no matter what was configured.
if (!model) {
res.json({ contextWindow: resolveNumCtx() });
return;
}
// Explicit Claude model name (e.g. from code.js's per-model lookup): 200K, no local metadata.
if (model.includes('claude')) {
res.json({ contextWindow: 200000 });
return;
}