Release 2.9.5: PPTX wizard improvements + PDF extraction fixes

- PDF: skip references/bibliography section from text extraction
- PDF: add minimal system prompt for translate sessions (translation-only, no tool calls)
- Wizard: image download button per card + bulk download all
- Wizard: translation modal download button (saves as _ko.txt / _en.txt)
- Wizard: save papers.json immediately on upload, after text extraction, and before outline step
- Wizard: short image directory names (22 chars + timestamp suffix)
- Wizard: project-images scans workspace root dir alongside pptx/ folder
- Server: pptx/list scans only pptx/ folder (workspace root scanning removed after file migration)
- create_presentation: save output to workspace/pptx/{slug}/ instead of workspace root
- edit_presentation: fuzzy path search includes pptx/ subfolder

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
kim
2026-06-02 23:10:40 +09:00
co-authored by Claude Sonnet 4.6
parent bba06bab6f
commit b292061c71
4 changed files with 259 additions and 50 deletions
+37 -12
View File
@@ -598,12 +598,14 @@ function setOrchestrationEnabled(enabled: boolean): void {
// can never silently diverge from getOrchestrationConfig() or getOrchestrationConfigForApi().
const clamped = clampOrchestrationConfig(current);
const preempt = clampPreemptConfig(current.preempt || {});
const merged = {
enabled,
secondary: {
const secBase: Record<string, any> = {
provider: String(current.secondary?.provider || '').trim(),
model: String(current.secondary?.model || '').trim(),
},
};
if (current.secondary?.vision !== undefined) secBase.vision = current.secondary.vision;
const merged = {
enabled,
secondary: secBase,
...clamped,
preempt,
};
@@ -4860,22 +4862,29 @@ async function handleChat(
const historyTurns = (getConfig().getConfig() as any)?.session?.historyTurns ?? 8;
const history = getHistoryForApiCall(sessionId, historyTurns, username);
const isCodeAiSession = String(sessionId || '').startsWith('code_ai_');
const isTranslateSession = String(sessionId || '').startsWith('translate-');
// ── Code AI session: block tools that break streaming ───────────────────
// create_file writes to disk in one shot (breaks live streaming).
// start_task / task_control would spawn background tasks (inappropriate in editor).
const codeAiBlockedTools = new Set(['start_task', 'task_control', 'create_file', 'python_eval', 'shell', 'run_command']);
const weatherToolNames = new Set(['weather_search', 'weather_airpollution', 'weather_openmeteo', 'weather_kma', 'weather_airkorea', 'weather_nasa_power', 'weather_era5', 'weather_cds', 'weather_cmip6']);
const legalToolNames = new Set(['korean_law_search', 'korean_law_fetch', 'us_case_search']);
const pptxToolNames = new Set(['create_presentation', 'edit_presentation']);
const meteorologistEnabled = isSkillEnabledForUser('meteorologist', userWorkspace);
const lawyerEnabled = isSkillEnabledForUser('lawyer', userWorkspace);
const presenterEnabled = isSkillEnabledForUser('presenter', userWorkspace);
const hasPptxKeyword = /슬라이드|발표|pptx|ppt\b|프레젠테이션|피피티|덱\b|presentation/i.test(message);
const skillToolFilter = (t: any) => {
const name = String(t?.function?.name || '');
if (weatherToolNames.has(name) && !meteorologistEnabled) return false;
if (legalToolNames.has(name) && !lawyerEnabled) return false;
if (pptxToolNames.has(name) && presenterEnabled && !hasPptxKeyword) return false;
return true;
};
const tools = isBootStartupTurn
? buildTools().filter((t: any) => bootAllowedTools.has(String(t?.function?.name || '')))
: isTranslateSession
? []
: buildTools().filter((t: any) => {
if (isCodeAiSession && codeAiBlockedTools.has(String(t?.function?.name || ''))) return false;
return skillToolFilter(t);
@@ -5301,12 +5310,13 @@ RULES:
const messages: any[] = [
{
role: 'system',
content: isCodeAiSession ? codeAiSystemPrompt : `${executionModeSystemBlock ? `${executionModeSystemBlock}\n\n` : ''}You are SmallClaw 🦞, a local AI assistant.\nCurrent date: ${dateStr}, ${timeStr}.\nNever search for or link SmallClaw repos unless the user is asking about SmallClaw itself.\nThis app runs on the user's own machine — browser/desktop automation requests are pre-authorized.\nKeep responses SHORT (1-2 sentences). Don't think out loud. Act and report. Greet naturally without tools.
content: isCodeAiSession ? codeAiSystemPrompt : isTranslateSession ? `You are a medical translator. Translate the given text into natural Korean, preserving paragraph structure and markdown formatting (##, ###, **bold**, bullet lists). Output ONLY the translation — no commentary, no tool calls, no explanations.` : `${executionModeSystemBlock ? `${executionModeSystemBlock}\n\n` : ''}You are SmallClaw 🦞, a local AI assistant.\nCurrent date: ${dateStr}, ${timeStr}.\nNever search for or link SmallClaw repos unless the user is asking about SmallClaw itself.\nThis app runs on the user's own machine — browser/desktop automation requests are pre-authorized.\nKeep responses SHORT (1-2 sentences). Don't think out loud. Act and report. Greet naturally without tools.
ANTI-HALLUCINATION: When a tool returns a result, report EXACTLY what the tool returned — never contradict or ignore tool output. If a tool says "(no rows)", say so. Never invent data, file contents, table names, or command output. If you don't know something, call a tool to find out or say you don't know.
IMAGE EDITING RULE: NEVER call image_edit (or any editing tool) when a user uploads a photo without explicitly requesting edits. Uploading a photo is NOT a request to edit it. Only call image_edit when the user's message explicitly asks for an edit (e.g. "수채화로 바꿔줘", "회전해줘"). Violating this rule is a critical error.
BROWSER RULE: NEVER call browser_open, browser_snapshot, browser_click, browser_fill, browser_scroll, or any browser_* tool unless the user EXPLICITLY asks to use the browser (e.g. "브라우저로 열어줘", "Chrome으로 봐줘", "사이트 직접 들어가봐"). "열어줘" or "보여줘" alone is NOT a browser request — answer from knowledge, use web_fetch/web_search, or output the relevant URL/embed. Pasting a URL or asking a question is also NOT a browser request.
CODE OUTPUT: When writing code in a fenced code block, start with a filename comment on line 1: \`# filename: snake_game.py\` (Python), \`// filename: app.js\` (JS/C), \`<!-- filename: index.html -->\` (HTML). Never repeat code already written in this conversation. For modifications to existing files, use replace_lines or insert_after (not create_file — it only works for NEW files). All code changes are presented as diffs for the user to review before being applied. Write code directly — do not ask for permission.
PACKAGE INSTALL: NEVER run pip install, npm install, apt-get, or any package installation command. If a package is missing, just write the code and mention the package name in a comment — let the user decide whether to install it. Do NOT attempt to install packages yourself.${callerContext ? '\n\n' + callerContext : ''}${browserStateCtx}${personalityCtx}${skillsManager.buildPromptContextForUser(username ? getUserWorkspace(username) : null, 16000, message)}${workflowCtx ? '\n\n' + workflowCtx : ''}`,
PACKAGE INSTALL: NEVER run pip install, npm install, apt-get, or any package installation command. If a package is missing, just write the code and mention the package name in a comment — let the user decide whether to install it. Do NOT attempt to install packages yourself.
OUTPUT FORMAT: When presenting 3+ items (news articles, emails, search results, lists), always use a markdown table or structured bullet list with clear headers. Never dump them as a long paragraph. Example: news → table with columns 제목|요약|출처.${callerContext ? '\n\n' + callerContext : ''}${browserStateCtx}${personalityCtx}${skillsManager.buildPromptContextForUser(username ? getUserWorkspace(username) : null, 16000, message)}${workflowCtx ? '\n\n' + workflowCtx : ''}`,
},
];
@@ -9319,6 +9329,7 @@ app.get('/api/pptx/list', requireGatewayAuth, async (req: express.Request, res:
const pptxBaseDir = path.join(workspacePath, PPTX_BASE);
fs.mkdirSync(pptxBaseDir, { recursive: true });
const projectMap = new Map<string, { mtime: number; pptxRelPath?: string; pptxFilename?: string }>();
const topEntries = fs.readdirSync(pptxBaseDir, { withFileTypes: true });
for (const entry of topEntries) {
if (!entry.isDirectory()) continue;
@@ -9937,7 +9948,10 @@ app.post('/api/pptx/extract-images', requireGatewayAuth, async (req: express.Req
let outDir = explicitOutDir;
if (!outDir && projectSlugEx && pdfPath) {
const pdfBase = path.basename(pdfPath, path.extname(pdfPath));
outDir = pptxRelPath(projectSlugEx, `${pdfBase}_images`);
const sanitized = pdfBase.replace(/\s+/g, '_').replace(/[^a-zA-Z0-9_-]/g, '');
const tsMatch = sanitized.match(/_(\d{10,})$/);
const shortBase = sanitized.slice(0, 22).replace(/_+$/, '') + (tsMatch ? '_' + tsMatch[1].slice(-6) : '');
outDir = pptxRelPath(projectSlugEx, `${shortBase}_images`);
}
if (!pdfPath) { res.status(400).json({ error: 'pdf_path required' }); return; }
const workspacePath = username ? getUserWorkspace(username) : getConfig().getWorkspacePath();
@@ -10070,20 +10084,31 @@ app.get('/api/pptx/project-images', requireGatewayAuth, async (req: express.Requ
const project = String(req.query.project || '').trim().replace(/\.\./g, '');
if (!project) { res.status(400).json({ error: 'project required' }); return; }
const projectDir = path.join(workspacePath, PPTX_BASE, project);
if (!fs.existsSync(projectDir)) { res.json({ images: [] }); return; }
const IMG_EXT = /\.(jpe?g|png|gif|webp)$/i;
const images: { path: string; name: string }[] = [];
const scan = (dir: string, rel: string) => {
const seenPaths = new Set<string>();
const scan = (dir: string, relBase: string) => {
if (!fs.existsSync(dir)) return;
for (const f of fs.readdirSync(dir)) {
const abs = path.join(dir, f);
const relF = rel ? `${rel}/${f}` : f;
const relF = relBase ? `${relBase}/${f}` : f;
try {
if (fs.statSync(abs).isDirectory()) { scan(abs, relF); continue; }
} catch { continue; }
if (IMG_EXT.test(f)) images.push({ path: pptxRelPath(project, relF), name: f });
if (IMG_EXT.test(f) && !seenPaths.has(abs)) {
seenPaths.add(abs);
images.push({ path: relF, name: f });
}
}
};
scan(projectDir, '');
// Primary: pptx project folder
scan(projectDir, pptxRelPath(project, '').replace(/\/$/, ''));
// Secondary: workspace root folder with same slug (agent-extracted images)
const slugOnly = project.split('/')[0];
const wsRootDir = path.join(workspacePath, slugOnly);
if (wsRootDir !== projectDir && fs.existsSync(wsRootDir) && fs.statSync(wsRootDir).isDirectory()) {
scan(wsRootDir, slugOnly);
}
res.json({ images });
} catch (err) {
res.status(500).json({ error: String(err) });
+129 -12
View File
@@ -17,6 +17,7 @@ function isPathInsideDir(base: string, target: string): boolean {
const COLUMN_EXTRACT_SCRIPT = `
import sys, fitz, re
from collections import Counter
pdf_path = sys.argv[1]
page_from = int(sys.argv[2]) if len(sys.argv) > 2 else 1
@@ -28,6 +29,7 @@ end = min(page_to, total) if page_to > 0 else total
LINE_TOL = 4
MARGIN = 0.07
TABLE_GAP = 60 # min px gap between words -> skip as table row
FURNITURE_RE = re.compile(
r'@[\\w.]+\\.|https?://|doi\\.org|\\u00a9|All rights are reserved'
@@ -46,6 +48,77 @@ def is_furniture(text, y0, y1, ph):
if FURNITURE_RE.search(text): return True
return False
def get_body_size(doc, start, end_p):
sizes = []
for pi in range(start, min(end_p, start+4)):
page = doc[pi]
ph = page.rect.height
for b in page.get_text("dict")["blocks"]:
if b.get("type") != 0: continue
if b["bbox"][1] < ph*0.07 or b["bbox"][3] > ph*0.93: continue
for ln in b.get("lines", []):
for sp in ln.get("spans", []):
if len(sp["text"].strip()) > 3:
sizes.append(round(sp["size"] * 2) / 2)
if not sizes: return 10.0
return Counter(sizes).most_common(1)[0][0]
body_size = get_body_size(doc, page_from-1, end)
def build_font_map(page, ph):
fmap = {}
for b in page.get_text("dict")["blocks"]:
if b.get("type") != 0: continue
if b["bbox"][1] < ph*0.07 or b["bbox"][3] > ph*0.93: continue
for ln in b.get("lines", []):
spans = [s for s in ln.get("spans", []) if s["text"].strip()]
if not spans: continue
y = ln["bbox"][1]
max_size = max(s["size"] for s in spans)
all_bold = all(bool(s["flags"] & 16) for s in spans)
matched = next((k for k in fmap if abs(k-y) <= LINE_TOL), None)
if matched is not None:
prev = fmap[matched]
fmap[matched] = (max(prev[0], max_size), prev[1] and all_bold)
else:
fmap[y] = (max_size, all_bold)
return fmap
def get_font_info(y, fmap):
best, bd = None, float('inf')
for ky, v in fmap.items():
d = abs(ky - y)
if d < bd and d <= LINE_TOL*3:
bd, best = d, v
return best
def format_line(text, y, fmap):
info = get_font_info(y, fmap)
if not info: return text
max_size, all_bold = info
ratio = max_size / body_size if body_size > 0 else 1.0
t = text.strip()
if ratio >= 1.35:
return '## ' + t
if ratio >= 1.12 or (all_bold and 3 < len(t) < 100):
return '### ' + t
if all_bold:
return '**' + t + '**'
return text
def col_baseline(words):
if not words: return None
xs = sorted(w[0] for w in words)
return xs[max(0, len(xs)//10)]
def make_indent(x_start, base_x, col_w):
if base_x is None or col_w <= 0: return ''
off = x_start - base_x
if off < col_w * 0.04: return ''
if off < col_w * 0.12: return ' '
if off < col_w * 0.25: return ' '
return ' '
def words_to_lines(wlist):
wlist.sort(key=lambda w: (w[1], w[0]))
groups, cur = [], []
@@ -57,18 +130,30 @@ def words_to_lines(wlist):
if cur: groups.append(cur)
lines = []
for g in groups:
text = ' '.join(w[4] for w in sorted(g, key=lambda w: w[0]))
g_sorted = sorted(g, key=lambda w: w[0])
text = ' '.join(w[4] for w in g_sorted)
text = re.sub(r'^([A-Z]) ([a-z][a-z])', lambda m: m.group(1)+m.group(2), text)
lines.append((g[0][1], text))
gaps = [g_sorted[k+1][0] - g_sorted[k][2] for k in range(len(g_sorted)-1)]
max_gap = max(gaps) if gaps else 0
lines.append((g[0][1], g_sorted[0][0], text, max_gap))
return lines
BULLET_RE = re.compile(r'^[\\u2022\\u00b7]\\s*|^[\\-\\*]\\s+(?=\\S)')
REF_HEADING_RE = re.compile(
r'^\\s*(References|Bibliography|참고문헌|REFERENCES|BIBLIOGRAPHY|Literature Cited)\\s*$',
re.IGNORECASE)
pages_text = []
stop_extraction = False
for pi in range(page_from-1, end):
if stop_extraction: break
page = doc[pi]
ph, pw = page.rect.height, page.rect.width
mid = pw * 0.52
full_w_thr = pw * 0.55
fmap = build_font_map(page, ph)
skip_rects = []
for b in page.get_text("blocks"):
x0,y0,x1,y1,txt = b[0],b[1],b[2],b[3],b[4]
@@ -97,19 +182,52 @@ for pi in range(page_from-1, end):
elif col=='left': left_w.append((wx0,wy0,wx1,wy1,word))
else: right_w.append((wx0,wy0,wx1,wy1,word))
full_base = col_baseline(full_w)
left_base = col_baseline(left_w)
right_base = col_baseline(right_w)
full_lines = words_to_lines(full_w)
left_lines = words_to_lines(left_w)
right_lines = words_to_lines(right_w)
all_entries = [(y,'F',t) for y,t in full_lines] \
+ [(y,'L',t) for y,t in left_lines] \
+ [(y,'R',t) for y,t in right_lines]
all_entries.sort(key=lambda e: (e[0] if e[1] in ('F','L') else e[0]+10000))
pages_text.append('\\n'.join(t for _,_,t in all_entries))
all_ys = sorted(e[0] for e in full_lines + left_lines)
lh = None
if len(all_ys) >= 4:
diffs = [all_ys[k+1]-all_ys[k] for k in range(len(all_ys)-1) if 0 < all_ys[k+1]-all_ys[k] < 40]
if diffs: lh = sorted(diffs)[len(diffs)//2]
def render(line_list, base_x, col_w, col_type):
out = []
prev_y = None
for y, x_start, text, max_gap in line_list:
if max_gap >= TABLE_GAP:
prev_y = None; continue
if prev_y is not None and lh and (y - prev_y) > lh * 1.8:
out.append((y - 0.5, col_type, ''))
t = text.strip()
indent = make_indent(x_start, base_x, col_w)
if BULLET_RE.match(t):
t = BULLET_RE.sub('- ', t)
out.append((y, col_type, indent + t))
else:
out.append((y, col_type, indent + format_line(t, y, fmap)))
prev_y = y
return out
all_entries = render(full_lines, full_base, pw * 0.85, 'F')
all_entries += render(left_lines, left_base, pw * 0.45, 'L')
all_entries += render(right_lines, right_base, pw * 0.45, 'R')
all_entries.sort(key=lambda e: e[0] if e[1] in ('F','L') else e[0]+10000)
page_lines = [t for _,_,t in all_entries]
# Stop at References/Bibliography heading
for idx, ln in enumerate(page_lines):
if REF_HEADING_RE.match(ln.lstrip('# ').lstrip('*').strip()):
page_lines = page_lines[:idx]
stop_extraction = True
break
pages_text.append('\\n'.join(page_lines))
full = '\\n\\n'.join(pages_text)
# drop cap 후처리
lines = full.split('\\n')
result = []
i = 0
@@ -121,7 +239,6 @@ while i < len(lines):
else:
result.append(lines[i])
i += 1
print('\\n'.join(result))
`;
@@ -173,7 +290,7 @@ async function runOcr(resolved: string, pageFrom: number, pageTo: number): Promi
export const pdfReadTool = {
name: 'pdf_read',
description: 'Extract text from a PDF file. For text-based PDFs uses pdftotext; for scanned/image PDFs automatically falls back to Tesseract OCR (kor+eng). Returns full text, optionally limited to a page range.',
description: 'Extract text from a PDF file. Returns markdown-formatted text: headings (##/###), bold (**), indentation, and bullet lists. Multi-column academic PDFs are handled correctly. Complex tables are skipped (use pdf_extract_images for table figures). For scanned/image PDFs falls back to Tesseract OCR.',
schema: {
path: 'Path to the PDF file (absolute, or relative to workspace)',
page_from: 'First page to extract, 1-indexed (default: 1)',
@@ -219,7 +336,7 @@ export const pdfReadTool = {
let method = 'pymupdf';
if (!forceOcr) {
// Primary: PyMuPDF column-aware extraction (handles 2-column academic PDFs)
// Primary: PyMuPDF markdown-aware extraction (headings, bold, indentation, table skip)
try {
const { stdout } = await execFileAsync(
'python3', ['-c', COLUMN_EXTRACT_SCRIPT, resolved, String(pageFrom), String(pageTo)],
+23 -4
View File
@@ -703,7 +703,7 @@ export const pptxTool: import('./registry.js').Tool = {
.replace(/^_+|_+$/g, '')
.toLowerCase()
.slice(0, 60) || 'presentation';
const projectDir = path.join(workspacePath, projectSlug);
const projectDir = path.join(workspacePath, 'pptx', projectSlug);
if (!fs.existsSync(projectDir)) {
fs.mkdirSync(projectDir, { recursive: true });
}
@@ -751,10 +751,18 @@ export const pptxTool: import('./registry.js').Tool = {
const resolvedSrc = fs.existsSync(srcPath) ? srcPath : fs.existsSync(srcPathAlt) ? srcPathAlt : null;
if (resolvedSrc) {
const ext = path.extname(resolvedSrc) || '.png';
const resolvedReal = fs.realpathSync(resolvedSrc);
const projectReal = fs.realpathSync(projectDir);
if (resolvedReal.startsWith(projectReal + path.sep) || resolvedReal.startsWith(projectReal + '/')) {
// Already inside project folder — use relative path, no copy needed
slide.image_path = path.relative(projectDir, resolvedReal);
console.log(`[pptx] slide ${i + 1} image_url already in project -> ${slide.image_path}`);
} else {
const fname = `slide${i + 1}_image${ext}`;
fs.copyFileSync(resolvedSrc, path.join(projectDir, fname));
slide.image_path = fname;
console.log(`[pptx] slide ${i + 1} image_url local copy -> ${fname}`);
}
} else {
// Last resort: search uploads/ subdirs for a file with the same basename
const basename = path.basename(url);
@@ -812,11 +820,18 @@ export const pptxTool: import('./registry.js').Tool = {
? slide.image_path
: path.join(workspacePath, slide.image_path);
if (fs.existsSync(inWorkspace)) {
const resolvedReal = fs.realpathSync(inWorkspace);
const projectReal = fs.realpathSync(projectDir);
if (resolvedReal.startsWith(projectReal + path.sep) || resolvedReal.startsWith(projectReal + '/')) {
slide.image_path = path.relative(projectDir, resolvedReal);
console.log(`[pptx] slide ${i + 1} image_path already in project -> ${slide.image_path}`);
} else {
const ext = path.extname(slide.image_path) || '.png';
const fname = `slide${i + 1}_image${ext}`;
fs.copyFileSync(inWorkspace, path.join(projectDir, fname));
slide.image_path = fname;
console.log(`[pptx] slide ${i + 1} image_path resolved from workspace -> ${fname}`);
}
} else {
// Last resort: search uploads/ subdirs for a file with the same basename
const basename = path.basename(slide.image_path);
@@ -972,15 +987,19 @@ export const editPptxTool: import('./registry.js').Tool = {
const givenFolder = path.basename(path.dirname(absPath));
let fuzzyMatch: string | null = null;
try {
for (const entry of fs.readdirSync(workspacePath)) {
const searchDirs = [workspacePath, path.join(workspacePath, 'pptx')];
outer: for (const searchDir of searchDirs) {
if (!fs.existsSync(searchDir)) continue;
for (const entry of fs.readdirSync(searchDir)) {
if (!entry.startsWith(givenFolder) || entry === givenFolder) continue;
const candidate = path.join(workspacePath, entry);
const candidate = path.join(searchDir, entry);
if (!fs.statSync(candidate).isDirectory()) continue;
const pptxFiles = fs.readdirSync(candidate).filter(f => f.toLowerCase().endsWith('.pptx'));
if (pptxFiles.length > 0) {
fuzzyMatch = path.join(candidate, pptxFiles[0]);
console.log(`[pptx] edit_presentation: fuzzy match "${existingPath}" → "${fuzzyMatch}"`);
break;
break outer;
}
}
}
} catch {}
+51 -3
View File
@@ -344,6 +344,7 @@ body{background:var(--bg);color:var(--text);font-family:'Segoe UI',system-ui,san
<div style="display:flex;align-items:center;gap:8px;margin-bottom:6px">
<span style="font-size:10px;color:var(--muted);font-weight:600;text-transform:uppercase;letter-spacing:.04em">캡처 이미지 <span id="imgcnt"></span></span>
<button id="btn-analyze-imgs" onclick="analyzeImages()" style="font-size:11px;color:#10b981;background:rgba(16,185,129,.1);border:1px solid #10b981;border-radius:5px;padding:2px 8px;cursor:pointer;flex-shrink:0" title="미분석 이미지를 mistral vision으로 재분석 (준비 실행 시 자동 처리됨)">🔍 재분석</button>
<button onclick="downloadAllImgs()" style="font-size:11px;color:#60a5fa;background:rgba(96,165,250,.1);border:1px solid #60a5fa;border-radius:5px;padding:2px 8px;cursor:pointer;flex-shrink:0" title="전체 이미지 다운로드">⬇ 전체</button>
<label style="display:flex;align-items:center;gap:4px;font-size:11px;color:var(--brand2);background:rgba(99,102,241,.1);border:1px solid var(--brand);border-radius:5px;padding:2px 8px;cursor:pointer;flex-shrink:0" title="이미지 파일 업로드">
<input type="file" id="img-upload-input" accept="image/*" multiple style="display:none" onchange="handleImgUpload(this.files)">
+ 이미지 업로드
@@ -532,6 +533,7 @@ body{background:var(--bg);color:var(--text);font-family:'Segoe UI',system-ui,san
<button class="btn bg" style="padding:3px 10px;font-size:11px" id="txl-tab-en" onclick="switchTxlTab('en')">원문</button>
<button class="btn bg" style="padding:3px 10px;font-size:11px" id="txl-tab-ko" onclick="switchTxlTab('ko')">번역</button>
<button class="btn bg" style="padding:3px 10px;font-size:11px" onclick="copyTxlText()">복사</button>
<button class="btn bg" style="padding:3px 10px;font-size:11px" onclick="downloadTxlText()">⬇ 다운로드</button>
<button class="btn bg" style="padding:3px 10px;font-size:11px" onclick="closeTxlModal()">✕</button>
</div>
<div class="txlmodal-body" id="txl-body"></div>
@@ -878,8 +880,13 @@ async function deleteSelectedProject(){
const r=await fetch('/api/pptx/delete-project',{...jh(),method:'DELETE',body:JSON.stringify({project})});
const d=await r.json();
if(!d.success) throw new Error(d.error||'삭제 실패');
papers=[]; selected=new Set(); paperMeta={};
paperTexts={}; paperTranslations={}; txlStatus={};
prepLog={}; extractedImages=[]; imageCaptions={}; abPaper=null;
outline=null; outlineEdits={};
await loadProjList();
sel.value='';
goStep(1);
alert(`"${project}" 삭제 완료.`);
}catch(e){ alert('삭제 오류: '+e.message); }
}
@@ -1519,7 +1526,7 @@ function updOutlineBtn() {
btn=document.createElement('button');
btn.id='s2-outline-btn'; btn.className='btn bp';
btn.textContent='✨ 아웃라인 생성 →';
btn.onclick=()=>{ goStep(3); generateOutline(); };
btn.onclick=()=>{ savePapersToProject(); goStep(3); generateOutline(); };
document.querySelector('#step-2 .bbar').appendChild(btn);
}
} else {
@@ -1564,6 +1571,19 @@ function copyTxlText(){
const s=document.getElementById('txl-status'); s.textContent='클립보드에 복사됨'; setTimeout(()=>s.textContent='',2000);
});
}
function downloadTxlText(){
const t=document.getElementById('txl-body').textContent;
if(!t.trim()) return;
const label=txlViewKey?paperLabel(txlViewKey):'translation';
const suffix=txlTabMode==='ko'?'_ko':'_en';
const filename=(label+suffix).replace(/[/\\?%*:|"<>]/g,'_').slice(0,80)+'.txt';
const blob=new Blob([t],{type:'text/plain;charset=utf-8'});
const a=document.createElement('a');
a.href=URL.createObjectURL(blob);
a.download=filename;
document.body.appendChild(a); a.click(); document.body.removeChild(a);
setTimeout(()=>URL.revokeObjectURL(a.href),1000);
}
// ── PDF 직접 업로드 ──────────────────────────────────────────────────────────
function handlePdfDrop(e) {
@@ -1597,6 +1617,7 @@ async function uploadPdfs(files) {
}
dz.innerHTML='📄 PDF 파일을 드래그하거나 클릭해서 선택<br><span style="font-size:10px">여러 파일 동시 업로드 가능</span>';
renderUploadList(); updSelUI();
savePapersToProject();
}
function removeUpload(idx){
const p=papers[idx]; if(!p||!p._isUpload)return;
@@ -2060,6 +2081,8 @@ async function startPrep() {
prepLog[key].txt='skip';
}
updProg(key); updOverall();
// 텍스트 추출 완료 즉시 저장 (번역 전에도 경로 보존)
if(prepLog[key].txt==='done') savePapersToProject();
// 텍스트 추출 성공 + 자동 번역 ON이면 백그라운드 번역 시작
if(autoTranslate && prepLog[key].txt==='done' && paperTexts[key]) translatePdfText(key,paperTexts[key]);
@@ -2103,9 +2126,10 @@ function renderImgGrid(){
const cap=imageCaptions[img.url]||'';
const excluded=cap==='제외';
const capHtml=cap?`<div style="font-size:8px;padding:1px 4px;line-height:1.3;color:${excluded?'#f87171':'#34d399'};overflow:hidden;text-overflow:ellipsis;white-space:nowrap" title="${esc(cap)}">${excluded?'⛔ 제외':cap}</div>`:'';
return `<div class="icard" title="${esc(img.label)}" style="position:relative;opacity:${excluded?.5:1}" onmouseenter="this.querySelector('.icard-del').style.opacity=1" onmouseleave="this.querySelector('.icard-del').style.opacity=0">
return `<div class="icard" title="${esc(img.label)}" style="position:relative;opacity:${excluded?.5:1}" onmouseenter="this.querySelectorAll('.icard-hover').forEach(b=>b.style.opacity=1)" onmouseleave="this.querySelectorAll('.icard-hover').forEach(b=>b.style.opacity=0)">
<img src="${esc(img.url)}" loading="lazy" onerror="this.style.display='none'" onclick="showImgPreview('${esc(img.url)}','${esc(img.label)}')" style="cursor:zoom-in;width:100%;height:100%;object-fit:cover">
<button class="icard-del" onclick="event.stopPropagation();deleteImg(${i})" style="position:absolute;top:4px;right:4px;opacity:0;transition:.15s;background:rgba(220,38,38,.85);border:none;color:#fff;border-radius:4px;width:22px;height:22px;font-size:13px;cursor:pointer;display:flex;align-items:center;justify-content:center;line-height:1" title="삭제">✕</button>
<button class="icard-hover icard-del" onclick="event.stopPropagation();deleteImg(${i})" style="position:absolute;top:4px;right:4px;opacity:0;transition:.15s;background:rgba(220,38,38,.85);border:none;color:#fff;border-radius:4px;width:22px;height:22px;font-size:13px;cursor:pointer;display:flex;align-items:center;justify-content:center;line-height:1" title="삭제">✕</button>
<button class="icard-hover" onclick="event.stopPropagation();downloadImg(${i})" style="position:absolute;top:4px;left:4px;opacity:0;transition:.15s;background:rgba(37,99,235,.85);border:none;color:#fff;border-radius:4px;width:22px;height:22px;font-size:13px;cursor:pointer;display:flex;align-items:center;justify-content:center;line-height:1" title="다운로드">⬇</button>
<div class="ilabel" style="overflow:hidden;text-overflow:ellipsis;white-space:nowrap;font-size:9px;padding:2px 5px">${esc(img.label)}</div>
${capHtml}
</div>`;
@@ -2154,6 +2178,30 @@ async function deleteImg(i){
renderImgGrid();
}
async function downloadImg(i) {
const img = extractedImages[i];
const filename = decodeURIComponent(img.url.split('/').pop().split('?')[0]) || 'image.png';
try {
const r = await fetch(img.url, {credentials: 'include'});
const blob = await r.blob();
const a = document.createElement('a');
a.href = URL.createObjectURL(blob);
a.download = filename;
document.body.appendChild(a);
a.click();
document.body.removeChild(a);
setTimeout(() => URL.revokeObjectURL(a.href), 1000);
} catch(e) { alert('다운로드 실패: ' + e.message); }
}
async function downloadAllImgs() {
if (!extractedImages.length) return;
for (let i = 0; i < extractedImages.length; i++) {
await downloadImg(i);
if (i < extractedImages.length - 1) await new Promise(r => setTimeout(r, 400));
}
}
async function doImgSearch() {
const q = document.getElementById('img-search-input').value.trim();
const res = document.getElementById('img-search-results');