fix(2.9.1): extractFromDocx — drop PHPWord dependency, parse word/document.xml directly
Root cause of the Shira hallucination incident in production case 46 Friedman: DocumentAnalyzer::extractFromDocx relied on PhpOffice\PhpWord, but on the EspoCRM prod image the PHPWord files are present in vendor/ yet are NOT registered in the composer PSR-4 autoload map — class_exists() silently returned false, the method returned null, and Shira's read_document fell back to OCR. OCR then only saw the signature image, the AI got "[signature]" as document content and fabricated the entire CTS appeal as a knee injury. This change rewrites extractFromDocx to use ZipArchive + a small regex parser of word/document.xml. Independent of PHPWord, more robust on tables/footnotes/hyperlinks (which PHPWord's element walker missed at depth >1), and verified on prod against the same Friedman appeal: 80 paragraphs / 16633 clean chars extracted (vs 0 before). The paired Python fix in shira-hermes (commit 7b517e1) makes the OCR fallback also read document.xml, so even if this PHP fix regresses again, the AI will not be fed "[signature]" as document content. Refs Task Master #5 Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user