fix: TTS 중첩클릭 에코 버그 + 음성엔진 한/영 발음 개선 + 생성속도 2배

- web-ui: 🔊 버튼 연타/중첩 클릭 시 이전 요청이 완료되면서 새 요청과 오디오가
  겹쳐 재생되던 버그 수정 (playTTSText/playTTSMsg를 _ttsPlay 공용 코어로 통합,
  AbortController + 활성 버튼 가드로 stale response의 재생을 원천 차단)
- voice.tts.provider를 edge_tts(ko-KR-SunHiNeural)에서 xtts_gpu로 전환 —
  실제로는 실시간 통화와 동일한 GPU 엔진(OmniVoice)을 타서 영어 섞인 문장의
  발음이 훨씬 자연스러워짐
- voice_engine.py: OmniVoice num_step 기본값을 32→16으로 낮춤. STT 왕복
  검증으로 실측: ~1.9배 빠른 생성(2.94s→1.55s), 품질 저하 없음(8은 문장이
  깨져서 폐기). 텍스트 버튼/실시간 통화 양쪽에 공통 적용됨

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
kim
2026-07-10 14:42:44 +09:00
co-authored by Claude Sonnet 5
parent f014fec7e2
commit adfa136597
3 changed files with 52 additions and 59 deletions
+12 -2
View File
@@ -130,11 +130,19 @@ class Engine:
except OSError:
pass
def synthesize_full(self, text: str):
# num_step default lowered from the library default (32) to 16: measured
# ~1.9x faster generation with no detectable quality loss via STT
# round-trip (2026-07-10, see project_tts_edge memory). 8 was tested and
# rejected — clearly audible/transcribable garbling on a full clause.
def synthesize_full(self, text: str, speed=None, num_step=None):
if num_step is None:
num_step = 16
audio = self.tts_model.generate(
text=text[:4000],
language='Korean',
voice_clone_prompt=self.voice_clone_prompt,
speed=speed,
num_step=num_step,
)
pcm = pcm16_from_float(np.asarray(audio[0], dtype=np.float32))
return pcm, self._tts_sample_rate
@@ -257,8 +265,10 @@ async def handle_connection(ws, engine: Engine):
elif mtype == 'tts_synthesize_full':
text = msg.get('text', '')
speed = msg.get('speed')
num_step = msg.get('num_step')
loop = asyncio.get_event_loop()
pcm, sample_rate = await loop.run_in_executor(None, engine.synthesize_full, text)
pcm, sample_rate = await loop.run_in_executor(None, engine.synthesize_full, text, speed, num_step)
await ws.send(json.dumps({
'type': 'tts_result',
'audioBase64': base64.b64encode(pcm).decode('ascii'),