feat: OpenAI 모델 목록 조회 라우트 추가, LM Studio thinking 비활성화 스위치

프론트엔드가 예전부터 호출하던 POST /api/openai/models가 백엔드에 구현이
안 되어 있어 항상 404였던 버그 수정 - OpenAICompatAdapter로 실제 OpenAI
API에 모델 목록을 조회하도록 구현.

지서버(LM Studio)의 gemma-4-26b-a4b가 응답당 완성 토큰을 2.5배 더 많이
쓰는 것을 벤치마크로 확인(322 vs 132 토큰/응답, thinking 오버헤드로 추정) -
chat_template_kwargs.enable_thinking=false를 기본 적용하되, 복잡한 질문에서
품질이 떨어질 수 있어 설정 화면(모델 → LM Studio)에 체크박스로 켜고 끌 수
있게 구현. 기본값은 비활성화(현재 상태 유지).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
kim
2026-08-07 12:57:22 +09:00
co-authored by Claude Sonnet 5
parent ae6e347339
commit b001cd9318
6 changed files with 47 additions and 1 deletions
+4
View File
@@ -1075,6 +1075,10 @@
<input id="settings-lmstudio-endpoint" type="text" placeholder="http://localhost:1234" style="width:100%;border:1px solid var(--line);border-radius:10px;padding:8px;font-size:12px;font-family:'IBM Plex Mono',monospace" />
<label style="display:block;font-size:12px;color:var(--muted);margin:10px 0 6px">모델 이름</label>
<input id="settings-lmstudio-model" type="text" placeholder="e.g. qwen2.5-7b-instruct" style="width:100%;border:1px solid var(--line);border-radius:10px;padding:8px;font-size:12px;font-family:'IBM Plex Mono',monospace" />
<label style="display:flex;align-items:center;gap:6px;font-size:12px;color:var(--muted);margin-top:10px;cursor:pointer">
<input id="settings-lmstudio-disable-thinking" type="checkbox" checked />
생각(thinking) 비활성화 — 답변 전 내부 추론 토큰을 끔 (Gemma 4 등 일부 모델에서 속도 개선, 대신 복잡한 질문 품질은 낮아질 수 있음)
</label>
<div style="display:flex;gap:8px;margin-top:8px">
<button class="btn btn-sm" onclick="refreshProviderModels()" style="background:#fff;border:1px solid var(--line);color:var(--text)">모델 감지</button>
<button class="btn btn-sm" onclick="testProviderConnection()" style="background:#fff;border:1px solid var(--line);color:var(--text)">테스트</button>