This repository has been archived on 2026-07-19. You can view files and clone it. You cannot open issues or pull requests or push a commit.
Files
shira-hermes/api/services/kb/classifier.py
T
chaim bb23b8a120 fix(kb): classifier reliability + pending-batches review endpoint
Two production-blocker fixes for the v0.8.0 batch upload UI, plus a
small endpoint to recover review state after a hard refresh.

classifier.py:
  - timeout 45s → 90s and concurrency 5 → 2. ai-gateway proxies a
    single Claude OAuth session so requests effectively serialize at
    the gateway; with concurrency=5, the 4th and 5th files in a batch
    of 5 timed out from queue waits alone (observed 2026-04-25 on
    user upload of 5 circulars: jobs 10 and 13 returned empty).
  - labels schema/prompt: arrays of strings via ai-gateway came back
    as [{}, {}, {}] (proxy mangled items.type). Switched the param
    to a comma-separated string with explicit format guidance, then
    split in _normalize. _normalize is defensive — accepts both
    string and list[str|dict] shapes so older callers don't break.
  - Sharpened the prompt distinction between subject (free text on
    the source) and labels (navigation tags), since identical
    examples were causing the model to emit one and skip the other.

ingest_jobs.list_pending_batches + admin_kb GET /admin/kb/batches:
  the review screen lives in JS RAM (_activeBatch). A hard refresh
  wipes it, leaving uploaded jobs orphaned in awaiting_review with
  no UI path back. The new endpoint groups jobs by batch_id with
  status counters; the EspoCRM 'ממתינים לאישור' panel uses it to
  let users resume any pending review.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-25 20:14:00 +00:00

257 lines
12 KiB
Python
Raw Blame History

This file contains invisible Unicode characters
This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""AI metadata extraction for uploaded KB documents.
Phase 6 (Task #20). When a user uploads a PDF/DOCX/TXT through the
batch UI, the background classifier reads the first ~3 pages of text,
calls Claude (via the existing ai-gateway proxy), and returns a draft
metadata payload — kind, title, identifier, subject, labels, dates,
and a short summary. The user reviews these on the new bulk-upload
review screen and edits before committing.
Design notes (see ~/.claude/plans/hidden-tickling-ullman.md):
• Tool-calling guarantees JSON shape — no prompt-engineering
fragility around "please return only JSON".
• Sonnet model (env CLAUDE_MODEL). Hebrew legal text is hard;
Haiku gets citations wrong too often.
• 15s timeout, semaphore-bounded to 5 concurrent calls so a 30-file
drop doesn't trip ai-gateway rate limits.
• Fail-soft: any error → empty dict. Caller transitions the job to
`awaiting_review` with `ai_suggestions = {}`, and the UI shows
"אין הצעות AI — מלא ידנית".
The classifier is an aid, not a gate. Wrong suggestions are an
annoyance; the user always confirms before chunks land in the index.
"""
from __future__ import annotations
import asyncio
import json
import logging
import os
from typing import Optional
from openai import AsyncOpenAI
logger = logging.getLogger("shira.kb.classifier")
# ai-gateway proxies a single Claude OAuth session, so requests
# effectively serialize at the gateway. Setting concurrency higher
# than 2 only piles requests in the gateway's queue, where they eat
# our timeout budget. Two in flight gives a tiny pipelining win if
# the gateway can overlap socket I/O with Claude latency, otherwise
# it's no worse than 1.
_CONCURRENCY = int(os.environ.get("KB_CLASSIFIER_CONCURRENCY", "2"))
# Empirically the ai-gateway → Claude tool-call round-trip with a ~3KB
# Hebrew system prompt + 5KB user content runs 18-25s. With 20 files
# queued through the gateway one-at-a-time, the tail can hit ~80s.
# Padded to 90s so we don't fail-soft on serialized queue waits.
_TIMEOUT_S = float(os.environ.get("KB_CLASSIFIER_TIMEOUT_S", "90"))
_MAX_INPUT_CHARS = 20_000 # ~3 pages of Hebrew legal text
_semaphore = asyncio.Semaphore(_CONCURRENCY)
_SYSTEM_PROMPT = """אתה מסווג מסמכים משפטיים בעברית עבור מאגר ידע משפטי.
תקבל את הטקסט של 3 העמודים הראשונים של מסמך, ועליך לחלץ ממנו מטא-דאטה.
הסיווג חייב לכבד את החוקים הבאים:
1. **kind** — אחד מהבאים בלבד:
• law — חוק (הצעת חוק, ספר חוקים, תיקון לחוק)
• regulation — תקנות
• circular — חוזר רשות / משרד / מוסד (חוזר נכות, חוזר מנכ"ל, חוזר לשכה רפואית)
• caselaw — פסק דין / החלטה. כותרת מתחילה לרוב ב-"בבית המשפט", "ע"א", "בג"ץ", "ב"ל",
"תיק", או שם תיק עם מספר.
• tool — כלי הערכה / מאמר / מדריך אקדמי. אינו מסמך משפטי פורמלי.
דוגמאות: GMFCS, MACS, CARS, Mini-MACS, מדריך הערכת תפקוד מאוניברסיטה.
2. **title** — כותרת המסמך כפי שהיא מופיעה בעמוד הראשון. אל תוסיף, אל תקצר.
ללא מרכאות סוגרות.
3. **identifier** — זיהוי קצר ומדויק. דוגמאות לפי kind:
• law/regulation: "חוק הביטוח הלאומי, התשנ\"ה–1995", "תקנה 37(א)"
• circular: "חוזר נכות 1990/02", "חוזר מנכ\"ל 14/2018"
• caselaw: "ע\"א 1234/24", "בג\"ץ 5678/22", "ב\"ל 9876-12-23"
• tool: שם הגוף המוציא ("CanChild — McMaster University")
אם אין זיהוי ברור — null. אל תמציא.
4. **subject** — תיאור תת-נושא קצר וחופשי, כטקסט אחד (≤80 תווים). זהו תיאור
אינפורמטיבי שמופיע בכרטיסיית המקור בחיפוש. למשל: "ילד נכה / אוטיזם",
"ילד נכה / שיתוק מוחין", "נפגעי עבודה / מחלת מקצוע".
5. **labels** — **שדה חובה**: 1 עד 3 תוויות לניווט וסינון במאגר, מופרדות בפסיק.
החזר מחרוזת אחת (string) — לא רשימה ולא אובייקט. דוגמאות בהמשך מראות בדיוק את הפורמט.
זהו שדה נפרד מ-subject ואסור להשמיט אותו, גם אם הוא חופף לחלוטין ל-subject.
הכלל: התווית הראשונה היא תמיד התת-קטגוריה הספציפית (זהה ל-subject ברוב המקרים).
התוויות הבאות, אם יש, הן רחבות יותר או הקשרים אחרים — שם כלי הערכה, סוג ועדה,
מספר חוזר וכו' — שיעזרו למצוא את המסמך מזוויות שונות.
כל תווית בפורמט "קטגוריה / תת-קטגוריה" כשרלוונטי.
דוגמאות לערך הסופי של השדה (string מופרד בפסיקים):
"ילד נכה / אוטיזם, ילד נכה, DSM-5"
"ילד נכה / שיתוק מוחין, GMFCS, כלי הערכה"
"ילד נכה / תלות בזולת, ועדה רפואית, ילד נכה"
"ילד נכה / עיכוב התפתחותי, DQ, פעוטות"
"נפגעי עבודה / שמיעה, אודיומטריה"
6. **published_at / effective_at** — בפורמט YYYY-MM-DD בלבד.
אם רק שנה ידועה — YYYY-01-01. אם לא ידוע — null.
effective_at הוא תאריך התחילה לתוקף (אם מצוין במפורש בנפרד מתאריך הפרסום).
7. **summary** — משפט אחד באורך 80–200 תווים שמסכם על מה המסמך. ללא קלישאות
("מסמך חשוב"...). תאר את התוכן הספציפי בקצרה.
‏**אסור: לא להמציא נתונים שלא מופיעים בטקסט. עדיף null על ניחוש.**
"""
_TOOL_SCHEMA = {
"type": "function",
"function": {
"name": "extract_kb_metadata",
"description": "Return extracted metadata for the document.",
"parameters": {
"type": "object",
"required": ["kind", "title", "summary", "labels"],
"properties": {
"kind": {
"type": "string",
"enum": ["law", "regulation", "circular", "caselaw", "tool"],
},
"title": {"type": "string", "maxLength": 500},
"identifier": {"type": ["string", "null"], "maxLength": 200},
"subject": {"type": ["string", "null"], "maxLength": 80},
"labels": {
"type": "string",
"minLength": 1,
"maxLength": 250,
"description": (
"1 to 3 navigation labels separated by commas. "
"Plain string only — NOT an array, NOT an object. "
"Example: 'ילד נכה / אוטיזם, ילד נכה, DSM-5'"
),
},
"published_at": {
"type": ["string", "null"],
"pattern": r"^\d{4}-\d{2}-\d{2}$",
},
"effective_at": {
"type": ["string", "null"],
"pattern": r"^\d{4}-\d{2}-\d{2}$",
},
"summary": {"type": "string", "minLength": 40, "maxLength": 250},
},
"additionalProperties": False,
},
},
}
async def classify(
*,
text: str,
filename: str,
topic_name: Optional[str] = None,
) -> dict:
"""Best-effort extraction. Returns {} on any error.
`text` is the first ~3 pages of the document (caller decides — we
truncate to _MAX_INPUT_CHARS just in case). `topic_name` lets the
LLM bias label suggestions toward the active legal domain
("ביטוח לאומי" vs "דיני עבודה").
"""
if not text or not text.strip():
logger.info("[classifier] empty text for %s — skipping", filename)
return {}
base = (os.environ.get("AI_GATEWAY_URL") or "http://localhost:3000").rstrip("/")
api_key = os.environ.get("AI_GATEWAY_API_KEY", "")
if not api_key:
logger.warning("[classifier] AI_GATEWAY_API_KEY not set — skipping")
return {}
truncated = text[:_MAX_INPUT_CHARS]
user_msg = (
f"topic_hint: {topic_name or 'כללי'}\n"
f"filename: {filename}\n\n"
f"--- מסמך ---\n{truncated}\n--- סוף ---"
)
client = AsyncOpenAI(base_url=f"{base}/v1", api_key=api_key)
async with _semaphore:
try:
resp = await asyncio.wait_for(
client.chat.completions.create(
model=os.environ.get("CLAUDE_MODEL", "sonnet"),
messages=[
{"role": "system", "content": _SYSTEM_PROMPT},
{"role": "user", "content": user_msg},
],
tools=[_TOOL_SCHEMA],
tool_choice={
"type": "function",
"function": {"name": "extract_kb_metadata"},
},
max_tokens=600,
temperature=0.1,
),
timeout=_TIMEOUT_S,
)
except asyncio.TimeoutError:
logger.warning("[classifier] timeout on %s after %ss", filename, _TIMEOUT_S)
return {}
except Exception as e:
logger.warning("[classifier] error on %s: %s", filename, e)
return {}
# Extract the tool call. We forced the model to call exactly one tool;
# if it didn't, we got nothing useful.
try:
choice = resp.choices[0]
tool_calls = choice.message.tool_calls or []
if not tool_calls:
logger.warning("[classifier] no tool call returned for %s", filename)
return {}
args_str = tool_calls[0].function.arguments
parsed = json.loads(args_str)
except (KeyError, IndexError, AttributeError, json.JSONDecodeError) as e:
logger.warning("[classifier] failed to parse tool call for %s: %s", filename, e)
return {}
return _normalize(parsed)
def _normalize(d: dict) -> dict:
"""Coerce types and drop empty values from the LLM output."""
out: dict = {}
if isinstance(d.get("kind"), str) and d["kind"] in (
"law", "regulation", "circular", "caselaw", "tool"
):
out["kind"] = d["kind"]
for key in ("title", "identifier", "subject", "summary"):
v = d.get(key)
if isinstance(v, str) and v.strip():
out[key] = v.strip()
for key in ("published_at", "effective_at"):
v = d.get(key)
if isinstance(v, str) and len(v) == 10 and v[4] == "-" and v[7] == "-":
out[key] = v
# labels comes as a comma-separated string from the model. Older
# callers may still hand us a list (defensive — we used to ask for
# an array and Claude proxies sometimes returned [{}] instead of
# strings, see the 2026-04-25 incident). Accept both shapes.
raw_labels = d.get("labels")
items: list[str] = []
if isinstance(raw_labels, str):
items = [s.strip() for s in raw_labels.split(",")]
elif isinstance(raw_labels, list):
for raw in raw_labels:
if isinstance(raw, str):
items.append(raw.strip())
elif isinstance(raw, dict):
# Empty {} from broken proxy round-trip — drop it.
continue
clean = [s for s in items if s][:3]
if clean:
out["labels"] = clean
return out