feat: OpenAI 모델 목록 조회 라우트 추가, LM Studio thinking 비활성화 스위치
프론트엔드가 예전부터 호출하던 POST /api/openai/models가 백엔드에 구현이 안 되어 있어 항상 404였던 버그 수정 - OpenAICompatAdapter로 실제 OpenAI API에 모델 목록을 조회하도록 구현. 지서버(LM Studio)의 gemma-4-26b-a4b가 응답당 완성 토큰을 2.5배 더 많이 쓰는 것을 벤치마크로 확인(322 vs 132 토큰/응답, thinking 오버헤드로 추정) - chat_template_kwargs.enable_thinking=false를 기본 적용하되, 복잡한 질문에서 품질이 떨어질 수 있어 설정 화면(모델 → LM Studio)에 체크박스로 켜고 끌 수 있게 구현. 기본값은 비활성화(현재 상태 유지). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
@@ -1075,6 +1075,10 @@
|
||||
<input id="settings-lmstudio-endpoint" type="text" placeholder="http://localhost:1234" style="width:100%;border:1px solid var(--line);border-radius:10px;padding:8px;font-size:12px;font-family:'IBM Plex Mono',monospace" />
|
||||
<label style="display:block;font-size:12px;color:var(--muted);margin:10px 0 6px">모델 이름</label>
|
||||
<input id="settings-lmstudio-model" type="text" placeholder="e.g. qwen2.5-7b-instruct" style="width:100%;border:1px solid var(--line);border-radius:10px;padding:8px;font-size:12px;font-family:'IBM Plex Mono',monospace" />
|
||||
<label style="display:flex;align-items:center;gap:6px;font-size:12px;color:var(--muted);margin-top:10px;cursor:pointer">
|
||||
<input id="settings-lmstudio-disable-thinking" type="checkbox" checked />
|
||||
생각(thinking) 비활성화 — 답변 전 내부 추론 토큰을 끔 (Gemma 4 등 일부 모델에서 속도 개선, 대신 복잡한 질문 품질은 낮아질 수 있음)
|
||||
</label>
|
||||
<div style="display:flex;gap:8px;margin-top:8px">
|
||||
<button class="btn btn-sm" onclick="refreshProviderModels()" style="background:#fff;border:1px solid var(--line);color:var(--text)">모델 감지</button>
|
||||
<button class="btn btn-sm" onclick="testProviderConnection()" style="background:#fff;border:1px solid var(--line);color:var(--text)">테스트</button>
|
||||
|
||||
Reference in New Issue
Block a user