Three independent bugs combined to make Shira fabricate the content of
case-46-Friedman's appeal in production (CTS case rewritten as a knee
injury, wrong dates, wrong court file number, wrong %-of-disability —
written into save_memory/Task/Meeting on prod).
* api/services/ocr.py — ocr_docx_images now reads word/document.xml first
(text content), then OCRs embedded images. Previously it only OCR'd the
word/media/* images, so a text-heavy DOCX with one signature image
returned only "[signature]" to the LLM. Verified against the same
Friedman appeal: 16633 chars / 80 paragraphs / 0 XML noise.
* mcp_server/tools/document_tools.py — _ocr_fallback now refuses to wrap
an empty/signature-only OCR result as success. If <80 useful chars or
just "[signature]", returns an explicit failure that instructs the LLM
not to describe / summarize the document.
* api/services/prompt_builder.py — TOOL_RULES adds the absolute rule
"DOCUMENT FAITHFULNESS": never quote a document not literally in the
tool result for THIS turn; never infer content from a file name; treat
dates and case numbers as especially dangerous to invent.
Refs Task Master #5, #6, #7
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two minimal additions for the upcoming v0.1.11 client work:
- /kb/search hits now include chunk_index, so the client can call
/source/{id}/chunks?around=N&ctx=K to fetch the matched chunk plus K
chunks before/after as context. Required by the new search-preview
layout for text sources, which until now duplicated the matched chunk
in both panes.
- /source/{id}/chunks accepts around (chunk_index) + ctx (window radius
clamped to 0..10). When around is given, returns only the windowed
range; otherwise unchanged (returns all chunks of the source).
Refs Task Master #1