Phase 2 of the multi-topic KB refactor. Adds three admin endpoints that
back the new "ניהול" tab in the KnowledgeBase EspoCRM extension:
POST /admin/kb/upload — multipart, queues a kb_ingest_job and
fires asyncio.create_task to process
parse → chunk → embed in the
background. Returns {job_id, status}
immediately so the browser doesn't
block on slow PDFs.
GET /admin/kb/jobs — recent jobs, optionally filtered by
topic_id / status / limit.
GET /admin/kb/jobs/{id} — single-job detail for the polling UI.
Migration 002 adds kb_ingest_job (queued/processing/done/failed) with
indexes on (status, created_at) and (topic_id, created_at).
ingest_source now accepts topic_id (NULL → defaults to 1 for back-compat
with the existing scan-inbox cron path) and writes it to kb_source.
S3 layout for new uploads is topic-aware: inbox/<topic_slug>/<kind>/...
The legacy /scan-inbox path on inbox/<kind>/ is unchanged.
Refs Task Master #13 (espocrm-extensions/KnowledgeBase)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Phase 1 of the multi-domain refactor. Adds a kb_topic table seeded with
the existing 'ביטוח לאומי' domain (id=1) and pins all 6 existing
kb_source rows to it, so search/ask can be scoped per domain without
disturbing the current single-topic UX.
- scripts/migrations/001_topics.sql: idempotent migration. Creates
kb_topic, seeds national-insurance with the system_prompt_addendum
copied verbatim from kb_public.py's hardcoded text, adds nullable
kb_source.topic_id, backfills it to 1, then SET NOT NULL. GRANTs to
shira_kb mirror its existing kb_source privileges. Wrapped in BEGIN
/COMMIT with a sanity-check that raises if the backfill is incomplete.
- api/services/kb/topics.py: list_topics / get_topic / resolve_topic_id.
resolve_topic_id raises ValueError on unknown ids so the route layer
can return 400 instead of silently mixing domains.
- api/services/kb/search.py: search() + _retrieve_rrf accept topic_id;
the SQL adds AND s.topic_id = $N when present. _expand_query is
parametrized by topic name (was hardcoded "הביטוח הלאומי").
- api/routes/kb_public.py: new GET /kb/topics. /kb/search /kb/sources
/kb/ask /kb/ask/stream all accept topic_id and call resolve_topic_id.
_build_ask_runner_context loads system_prompt_addendum from the DB
instead of the hardcoded insurance string.
- mcp_server/tools/legal_kb_tools.py: tool renamed
search_insurance_kb → search_legal_kb, takes topic_id at registration
so each conversation is pinned to its caller-chosen domain. Tool
description interpolates the topic name.
- api/services/agent_runner.py: kb_topic_id + kb_topic_name flow
through to register_legal_kb_tools. Hebrew progress labels updated.
Refs Task Master #12 (espocrm-extensions/KnowledgeBase).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Each file in the inbox can now be accompanied by a `<filename>.meta.json`
sidecar that supplies identifier, published_at, effective_at, and
source_url. The scanner applies filename-derived defaults first and lets
the sidecar override individual fields.
- list_inbox / list_failed skip `.meta.json` so sidecars aren't ingested
as standalone documents.
- mark_processed, mark_failed, and requeue-failed move the sidecar
alongside its parent (best-effort).
- mark_failed now sanitizes the error string to ASCII before writing it
to S3 object metadata (HTTP header encoding).
- search SQL selects published_at/effective_at; the formatter shows the
full heading_path plus publication date so citations are unambiguous.
- scripts/kb_sidecar_template.json documents the expected shape.
Refs Task Master #2
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
One-shot script that fetches the consolidated National Insurance Law from
Hebrew Wikisource and emits plain text with heading lines the chunker can
parse (פרק / סימן / סעיף).
Key quirks of the source markup:
- {{ח:קטע2}} carries the level name inside the title parameter ("פרק א׳:
…"), not in the template name. Emit the title verbatim and let the
chunker's existing prefix regex handle it.
- Named args (e.g. 'אחר=[פרק א׳]') must be filtered — the original
positional parser was treating them as positional titles and producing
gibberish "חלק אחר=[פרק א׳]" headings.
- No חלק level exists in this law — only פרק > סימן > סעיף.
Run as: python3 scripts/fetch_national_insurance_law.py <out_path>
Then upload the output .txt to s3://insurance-kb/inbox/law/ and invoke
POST /admin/kb/scan-inbox.
Refs Task Master #2
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>