fix: 새어나온 바이트 폴백 토큰을 원래 글자로 복원

실사고(2026-08-12): 뉴스 요약 답변이 이렇게 나왔다.

  "정치적 갈등과 사회적 사건들이 눈에 <0xEB><0x9D><0x95>니다."
                                    EB 9D 95 = U+B755 = 띕

GGUF 토크나이저에 통글자 항목이 없는 글자는 바이트 단위로 쪼개지고, 디토크나이저가
다시 붙여야 하는데 실패하면 바이트 토큰의 "표기"가 그대로 출력에 남는다. 모델은
정확한 바이트를 냈다 — `띕`(ㄸ+ㅢ+ㅂ)이 토큰 테이블에 없을 만큼 드문 음절일 뿐이다.
잃은 정보가 없고 정답이 완전히 결정되는 무손실 복원이라, 코드로 갈 자리다.

처음 발견했을 땐 로그 전체에 3건이라 빈도가 낮다고 보고 미뤘는데, 바로 다음
답변에서 또 나왔다. "눈에 띕니다"가 뉴스 요약에서 흔한 표현이라 체감 빈도가
계산보다 높다 — 미룬 판단이 틀렸다.

복원이 스스로 해를 끼치지 않도록 두 가지를 지킨다:
- 유효한 UTF-8로 디코드되는 것만 바꾼다. 홀로 나온 <0xFF>는 잘린 글자가 아니므로
  U+FFFD로 바꾸면 복원이 아니라 정보 파괴다
- 코드블록·인라인코드 안은 건드리지 않는다. 거기 있는 <0xEB>는 내용이다
  (헥스 덤프, 토크나이저 예제, 이 기능 자체의 문서 등)

조립된 전체 텍스트에 적용한다 — 스트리밍 청크에 걸면 여러 청크에 걸쳐 도착한
바이트 열을 놓친다.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
kim
2026-08-12 12:51:50 +09:00
co-authored by Claude Opus 5
parent 2f6872b026
commit bf07a15112
3 changed files with 122 additions and 1 deletions
+50
View File
@@ -0,0 +1,50 @@
/**
* byte-fallback.ts
*
* Reassembles byte-fallback tokens that reached the output as literal text.
*
* A GGUF tokenizer that has no whole-token entry for a character falls back to encoding it one
* byte at a time. Detokenization is supposed to put those bytes back together; when it doesn't,
* the byte token's printed FORM leaks into the reply instead of the character:
*
* "정치적 갈등과 사회적 사건들이 눈에 <0xEB><0x9D><0x95>니다." ← observed 2026-08-12
* EB 9D 95 = U+B755 = 띕
*
* The intended word was "눈에 띕니다". Nothing was lost — the model emitted the correct bytes, and
* `띕` (ㄸ+ㅢ+ㅂ) is simply an uncommon enough syllable to miss the token table. So this is a
* lossless repair with a fully determined answer, which is exactly the kind of thing that belongs
* in code rather than in a prompt.
*
* Two rules keep the repair from doing damage of its own:
* 1. Only rewrite runs that decode to valid UTF-8. A stray `<0xFF>` is not a truncated
* character, and turning it into U+FFFD would destroy information rather than restore it.
* 2. Leave fenced code blocks alone. `<0xEB>` inside a code sample is content — a hex dump, a
* tokenizer example, this file's own docs — and rewriting it silently corrupts the sample.
*/
const BYTE_RUN = /(?:<0x[0-9A-Fa-f]{2}>)+/g;
const CODE_FENCE = /```[\s\S]*?```|`[^`\n]*`/g;
function decodeRun(run: string): string {
const bytes = (run.match(/[0-9A-Fa-f]{2}/g) || []).map(h => parseInt(h, 16));
const buf = Buffer.from(bytes);
const decoded = buf.toString('utf8');
// Node substitutes U+FFFD for anything it can't decode; that means these bytes were not a
// truncated character, so the original text is the more faithful output.
return decoded.includes('�') ? run : decoded;
}
/** Put leaked byte-fallback tokens back together. Returns the text unchanged when there are none. */
export function repairByteFallbackTokens(text: string): string {
const s = String(text || '');
if (!s.includes('<0x')) return s; // cheap guard: the overwhelmingly common case
// Splice around code spans so their contents are preserved verbatim.
let out = '';
let last = 0;
for (const m of s.matchAll(CODE_FENCE)) {
out += s.slice(last, m.index).replace(BYTE_RUN, decodeRun) + m[0];
last = (m.index ?? 0) + m[0].length;
}
return out + s.slice(last).replace(BYTE_RUN, decodeRun);
}
+4 -1
View File
@@ -8,6 +8,7 @@
*/
import { Ollama } from 'ollama';
import { repairByteFallbackTokens } from './byte-fallback.js';
import fs from 'fs';
import path from 'path';
import type { LLMProvider, ChatMessage, ContentPart, ChatOptions, ChatResult, GenerateOptions, GenerateResult, ModelInfo } from './LLMProvider';
@@ -651,6 +652,7 @@ export class OllamaAdapter implements LLMProvider {
continue;
}
message.content = repairByteFallbackTokens(String(message.content || ''));
return { message, thinking: response?.thinking, usage: _usage };
} catch (error: any) {
lastError = error;
@@ -686,7 +688,7 @@ export class OllamaAdapter implements LLMProvider {
if (textMatch) {
const rawText = textMatch[1].replace(/\{[\s\S]*?\}/g, '').trim();
if (rawText.length > 10) {
return { message: { role: 'assistant' as const, content: rawText } };
return { message: { role: 'assistant' as const, content: repairByteFallbackTokens(rawText) } };
}
}
throw new Error(`Ollama chat failed: ${msg}`);
@@ -843,6 +845,7 @@ export class OllamaAdapter implements LLMProvider {
continue;
}
message.content = repairByteFallbackTokens(String(message.content || ''));
return { message, thinking: fullThinking || undefined, usage: { promptTokens, completionTokens, ctxWindow: modelCtx } };
} catch (error: any) {
lastError = error;
+68
View File
@@ -0,0 +1,68 @@
/**
* byte-fallback — 새어나온 바이트 토큰 복원
*
* 2026-08-12 실사고: 뉴스 요약 답변이 "눈에 <0xEB><0x9D><0x95>니다"로 나왔다. EB 9D 95는
* U+B755 = 띕. 모델은 정확한 바이트를 냈고 디토크나이저가 못 붙인 것뿐이라, 정답이 완전히
* 결정되는 무손실 복원이다.
*
* 이 테스트가 지키는 두 가지 안전선:
* 1) 유효한 UTF-8로 디코드되는 것만 바꾼다 — 아니면 원문이 더 충실하다
* 2) 코드블록 안은 건드리지 않는다 — 거기 있는 <0xEB>는 내용이다
*/
import { test, describe } from 'node:test';
import assert from 'node:assert/strict';
import { repairByteFallbackTokens } from '../src/providers/byte-fallback';
describe('복원', () => {
test('실사고 문장', () => {
assert.equal(
repairByteFallbackTokens('정치적 갈등과 사회적 사건들이 눈에 <0xEB><0x9D><0x95>니다.'),
'정치적 갈등과 사회적 사건들이 눈에 띕니다.',
);
});
test('한 문장에 여러 군데', () => {
assert.equal(
repairByteFallbackTokens('<0xEB><0x9D><0x95>고 <0xEB><0x9D><0x95>고'),
'띕고 띕고',
);
});
test('바이트 토큰이 없으면 그대로 반환한다', () => {
const s = '평범한 문장입니다. 0x1F 같은 표기도 건드리지 않습니다.';
assert.equal(repairByteFallbackTokens(s), s);
});
test('빈 입력', () => {
assert.equal(repairByteFallbackTokens(''), '');
assert.equal(repairByteFallbackTokens(null as any), '');
});
});
describe('안전선', () => {
test('유효한 UTF-8이 아니면 원문을 남긴다', () => {
// 0xFF는 UTF-8 어디에도 나올 수 없는 바이트다. 잘린 글자가 아니므로 U+FFFD로 바꾸면
// 복원이 아니라 정보 파괴가 된다.
assert.equal(repairByteFallbackTokens('앞 <0xFF> 뒤'), '앞 <0xFF> 뒤');
});
test('코드블록 안은 건드리지 않는다', () => {
const s = '설명입니다.\n```\nbytes = <0xEB><0x9D><0x95>\n```\n끝.';
assert.equal(repairByteFallbackTokens(s), s);
});
test('인라인 코드도 건드리지 않는다', () => {
assert.equal(
repairByteFallbackTokens('토큰 `<0xEB>` 는 바이트 폴백 표기입니다.'),
'토큰 `<0xEB>` 는 바이트 폴백 표기입니다.',
);
});
test('코드블록 밖은 고치고 안은 남긴다 — 한 문서에 섞여 있어도', () => {
assert.equal(
repairByteFallbackTokens('눈에 <0xEB><0x9D><0x95>니다.\n```\n<0xEB><0x9D><0x95>\n```'),
'눈에 띕니다.\n```\n<0xEB><0x9D><0x95>\n```',
);
});
});