- Replace piper TTS with edge-tts (Microsoft neural Korean voice)
- piper ko_KR-kss-medium uses pygoruut phoneme type unsupported by C++ binary,
causing Chinese-sounding output; edge-tts solves this via online neural TTS
- Added src/tools/edge_tts_synth.py Python helper script
- Server TTS endpoint now dispatches by provider (edge_tts vs piper)
- Config updated: voice.tts.provider = "edge_tts", voice = "ko-KR-SunHiNeural"
- Add TTS toggle switch to web UI (chat input bar)
- Toggles visibility of all 🔊 buttons and stops active playback
- State persisted in localStorage (default: on)
- Mic input now auto-submits after speech recognition completes
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- STT: Replace MediaRecorder+Whisper with browser Web Speech API (ko-KR)
Whisper base model hallucinated English for Korean speech; Chrome's
built-in Google STT is far more accurate for Korean
- TTS: Skip ffmpeg OGG conversion, return WAV directly from Piper
Avoids OGG/Opus codec compatibility issues in browsers
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- POST /api/voice/stt: Whisper-cpp transcription of webm/ogg audio
- POST /api/voice/tts: Piper Korean TTS returning ogg audio
- Web UI: 🎤 mic button in chat input; click to start/stop recording
MediaRecorder → /api/voice/stt → auto-fills chat input
- Web UI: 🔊 TTS button on each AI message; click to play/stop
- config.json: add absolute model/binary paths for whisper-cli and piper
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Luckysheet modal: 8-direction resize handles (edge + corner), title bar drag-to-move
- First drag/resize converts from flex-centered to absolute positioning
- Minimize/close reset to centered state on reopen
- New xlsx-viewer.html: standalone Luckysheet page for opening xlsx in separate window
- Added "↗ 새 창" button on xlsx cards to open viewer in a resizable popup window
- Viewer shares session cookies so auth works transparently
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Skills:
- Add accountant (CPA) and investor (CFA) skills with site: hints
- Add site: hints to lawyer, psychiatrist, musician skills
- Allow multiple skills active simultaneously (remove exclusive mode)
Tools:
- Add excel_read / excel_write tools (openpyxl-based)
- Fix python_eval packages: retry with --break-system-packages on PEP 668 failure
- Fix email_read: spaces→underscores in filenames, return /api/files/ links,
remove absolute savedPaths from data, send SSE 'files' event for attachments
Server:
- FILE_OP v2: fall through to primary when secondary model unavailable
- Add PUT /api/files/{*filePath} endpoint for xlsx editor save
UI:
- xlsx viewer: Luckysheet modal (lazy CDN load, ~3MB on first use)
with edit mode, 💾 save button, Nanum Gothic font
- Email attachment preview: SSE 'files' event appends file links to final reply
- ✏️ 편집 button (blue) for xlsx cards
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
스캔 PDF에서 텍스트 추출 실패 시 자동으로 OCR 실행.
PyMuPDF로 300DPI 이미지 렌더링 후 tesseract kor+eng 적용.
ocr: true 파라미터로 강제 OCR 가능. method 필드로 추출 방법 표시.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
사용자가 업로드한 PDF는 이미 갖고 있으므로 저장 버튼 불필요.
renderMarkdown/renderLinkButton에 hideDownload 옵션 추가,
유저 메시지 렌더링 시 적용. 미리보기 버튼은 유지.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
스캔 악보 PDF 처리 절차 추가:
pdf_read 시도 → 실패 시 pdf_extract_images로 페이지 추출 → image_read로 시각 분석.
도구 목록에 pdf_extract_images, image_read 추가.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Settings that were silently lost on reinstall:
- Track .smallclaw/heartbeat/config.json (active hours 8-22, interval 30m)
- Track .smallclaw/skills_state.json (multi-agent-orchestrator enabled)
- Fix gitignore: heartbeat/ → heartbeat/runs/ (only ignore run logs, not config)
- Fix gitignore: remove skills_state.json from ignored list
Vault (contains channel tokens, passwords) cannot go in git — instead:
- Add scripts/backup-vault.sh: backs up vault+credentials to ~/.smallclaw-backup/
keeping last 5 snapshots. Run manually or via cron (0 3 * * *).
- Run initial backup now → ~/.smallclaw-backup/vault_20260518_232441
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
When no explicit output path is given and source is JPEG/WebP, output
defaults to PNG (lossless). This avoids cumulative quality loss on
repeated edits (brightness → contrast → filter → ...).
Exception: 'convert' operation keeps the original format so explicit
format-change requests still work. Explicit output path always honored.
Also fix quality default mismatch: TS params was 85, Python was 92.
Now both are 92 for when JPEG/WebP output is explicitly requested.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Add per-user session/workspace paths to .gitignore so runtime data
(sessions, uploads, memory, attachments, pubmed) is no longer tracked
- Narrow ollama_web_search description to prevent it being chosen over
web_search for news/info queries — it's now explicitly a fallback only
- Add scripts/install-image-deps.sh to install Pillow/torch/rembg deps
- Remove stray root files (1, pdf-pptx.txt, eng.traineddata)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Add Korean search keywords to WEB category (뉴스, 검색해, etc.) so
Korean news queries trigger web_search instead of ollama_web_search
- Show ✓ SearXNG in startup banner when searxng_url is configured
- Update dental dict databases scripts and DB
- Clean up stale sessions and workspace runtime files
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
image_edit:
- speech_bubble operation with Korean font, rounded rect, tail
- sketch quality: CLAHE pre-processing for better line contrast
- output filename with timestamp to prevent overwrites
- anime/painting stylize timeout 30s→600s (AnimeGANv2 model download)
- remove_bg PNG output fix
server-v2.ts:
- resolveToolImageContent: embed edited image as base64 in tool results
so vision models can see the output (3 execution paths covered)
- _activeModelName: fix vision support detection for Ollama-hosted models
(kimi/gemini run via Ollama — provider='ollama' was masking vision capability)
- Synthetic tool call ID generation to prevent Gemini function_response empty name error
- image_edit path hint format changed to English to prevent Gemini hallucination
- userRequestedImageEdit: expanded keyword list (풍선, 달아, 붙여 등)
- image_read OCR failure returns success:true with visual description hint
- TOOL_BLOCKS photo: image_edit workflow documentation updated
web-ui/index.html:
- renderFileDownloads: skip download buttons for image file extensions
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Reorganize .smallclaw/databases/ around a single entry point with
subcommands (add-terms / add-images / verify / reassign / stats),
absorb pmc_reassign.py into reassign, and route every image source
through inline vision-model verification before commit.
Layout
config.py paths, endpoints, API-key locations (no more
hard-coded absolute paths in source files)
manage.py argparse dispatcher
workflow_terms.py seed-JSON import (JSONC supported, dedupes on
korean/english)
workflow_images.py renamed from dental_image_workflow.py; main()
converted to run(verify=True, ...)
workflow_verify.py validate_image() / verify_db() + verify_updates()
gate used by every add-images source
seeds/ recovered seed_all.json, seed_periodontics.json
scratch/ ad-hoc work area replacing the /tmp habit
(only README.md is tracked)
docs/howto.md moved + expanded
docs/archive_image_rounds/ one-off round scripts + their results
docs/archive_validation/ model-comparison + validation history
Verification
- VERIFY_MODEL defaults to qwen3.5:397b-cloud (50-image test on
procedure-heavy sample: 9.8s/img avg, zero timeouts, Kimi-level
rigor — see archive_validation/).
- Prompt strengthened: in-image text/captions are not valid grounds
for "적합"; technique terms require visible procedure steps, not
generic device/anatomy photos.
- All 9 add-images sources gated through verify_updates() before
commit; --no-verify escape hatch for bulk runs.
.gitignore: API key files, dental_images/, __pycache__/, scratch/*
(README.md kept).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Python stdout already includes the correct markdown download link.
Appending extra bare URL and absolute path caused the model to generate
two download links in its response. Keep only Python's stdout.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The v2.2.3 hint "ALWAYS use /api/files/uploads/filename" was too broad,
causing the model to format PPTX download links as uploads/file.pptx
instead of the correct project_folder/file.pptx path returned by the tool.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Playwright's bundled browser install fails on Ubuntu 26.04 with
"Playwright does not support chromium on ubuntu26.04-x64".
Fall back to system Chromium (/usr/bin/chromium-browser etc.) via
executablePath when launching headless, bypassing the platform check.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>