fix: 해석 실패한 컨텍스트 창 추측값을 브라우저가 캐시해 세션창이 8K에 박히던 문제
사용자 제보: "지서버 모델 연결됐는데 컨텍스트가 세션창에서만 8킬로", 그리고
"웹을 리로드하면 256킬로가 된다".
서버 쪽 마지막 폴백(8192)은 이미 캐시하지 않게 돼 있어서 매번 다시 조회한다
(그 사고를 겪고 고친 주석이 남아 있다). 문제는 브라우저였다 — app.js가
if (_appCtxMax > 0 && !force) return; // 한 번 받으면 끝
로 한 번 받은 값을 페이지 수명 내내 들고 있어서, 지서버가 자고 있을 때 열면
8192가 그대로 박혔다. 리로드가 고친 게 아니라 리로드로 캐시가 비워진 것.
메인 채팅 게이지는 SSE usage로 라이브 값을 받아서 멀쩡했고 세션창만 틀렸던
이유가 이것 — 두 화면이 서로 다른 경로로 같은 숫자를 구하고 있었다.
어댑터 쪽(bfbf04e)과 같은 원칙으로 고친다: **추측을 측정처럼 다루지 않는다.**
- resolveNumCtxInfo()가 value와 known(진짜 해석 결과인지)을 함께 반환
- /api/model-context가 provisional 플래그를 실어 보낸다. /api/show 실패나
마지막 폴백이면 true
- app.js: provisional이면 화면에는 보여주되 _appCtxMax에 저장하지 않아 다음
호출에서 다시 묻는다
- code.js: provisional이면 모델별 1회 래치(_codeAiCtxMaxFetchedFor)를 풀어
다시 조회하게 한다
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -109,7 +109,7 @@ import { getVault, SecretValue } from '../security/vault';
|
||||
import { getOllamaClient } from '../agents/ollama-client';
|
||||
import { startOllamaLocalKeepAlive } from './ollama-local-keepalive';
|
||||
import { spawnAgent } from '../agents/spawner';
|
||||
import { getSession, addMessage, getHistory, getHistoryForApiCall, getWorkspace, setWorkspace, clearHistory, setContextTokens, cleanupSessions, cleanupEmptySessions, migrateGlobalSessionsToUser, listAllSessions, resolveNumCtx, activeOllamaEndpoint } from './session';
|
||||
import { getSession, addMessage, getHistory, getHistoryForApiCall, getWorkspace, setWorkspace, clearHistory, setContextTokens, cleanupSessions, cleanupEmptySessions, migrateGlobalSessionsToUser, listAllSessions, resolveNumCtx, resolveNumCtxInfo, activeOllamaEndpoint } from './session';
|
||||
import { hookBus } from './hooks/hooks';
|
||||
import { loadWorkspaceHooks } from './hooks/hook-loader';
|
||||
import { validateLinksInText, filterImageMarkdown } from './guards/link-validator';
|
||||
@@ -3583,7 +3583,11 @@ app.get('/api/model-context', async (req, res) => {
|
||||
// whatever provider is actually primary (ollama/google/etc) instead of the
|
||||
// old hardcoded 200000 that silently applied no matter what was configured.
|
||||
if (!model) {
|
||||
res.json({ contextWindow: resolveNumCtx() });
|
||||
// `provisional` = every real lookup failed and this is the 8192 last resort. The client must
|
||||
// not cache it, or one request made while 지서버 was asleep pins the panel at 8K for the life
|
||||
// of the page (2026-08-12 report: "세션창에서만 8킬로", fixed by reloading).
|
||||
const info = resolveNumCtxInfo();
|
||||
res.json({ contextWindow: info.value, provisional: !info.known });
|
||||
return;
|
||||
}
|
||||
// Explicit Claude model name (e.g. from code.js's per-model lookup): 200K, no local metadata.
|
||||
@@ -3604,7 +3608,7 @@ app.get('/api/model-context', async (req, res) => {
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({ name: model }),
|
||||
});
|
||||
if (!response.ok) { res.json({ contextWindow: null }); return; }
|
||||
if (!response.ok) { res.json({ contextWindow: null, provisional: true }); return; }
|
||||
const data = await response.json() as any;
|
||||
let ctxWindow: number | null = null;
|
||||
const modelInfo = data.model_info || {};
|
||||
@@ -3618,9 +3622,9 @@ app.get('/api/model-context', async (req, res) => {
|
||||
const m = data.parameters.match(/\bnum_ctx\s+(\d+)/);
|
||||
if (m) ctxWindow = parseInt(m[1], 10);
|
||||
}
|
||||
res.json({ contextWindow: ctxWindow });
|
||||
res.json({ contextWindow: ctxWindow, provisional: !ctxWindow });
|
||||
} catch (err: any) {
|
||||
res.json({ contextWindow: null, error: err.message });
|
||||
res.json({ contextWindow: null, provisional: true, error: err.message });
|
||||
}
|
||||
});
|
||||
|
||||
|
||||
Reference in New Issue
Block a user