Commit Graph

22 Commits

Author SHA1 Message Date
Nitzan Pomerantz 75ee79bbfd Add percentage-based backup outlier filtering for heterogeneous data
Root cause: IQR filtering fails with wide distributions (mixed room counts, property types).
Example: 9,459 NIS/sqm deal (68% below average) passed IQR with k=1.0 due to wide IQR.

Solution: Dual filtering approach - IQR + percentage backup

Changes:
- Add ANALYSIS_USE_PERCENTAGE_BACKUP config (default: true)
- Add ANALYSIS_PERCENTAGE_THRESHOLD config (default: 0.5 = 50%)
- Apply percentage backup after IQR when method=iqr
- Update outlier report with percentage backup parameters
- Update CLAUDE.md docs with dual filtering explanation

Behavior:
1. Hard bounds filter (1K-100K/sqm, 100K min total)
2. IQR filter (k=1.0) - catches mild outliers
3. Percentage filter (50% from median) - catches extreme outliers
4. Deal removed if flagged by ANY method

Example: 12K/sqm deal (60% below 30K median) now caught by percentage backup
even when IQR bounds are permissive due to heterogeneous data.

Tests: All 326 tests pass, manual verification confirms extreme outliers removed.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-11-29 14:44:14 +02:00
Nitzan Pomerantz 8e591423da Change default IQR multiplier to 1.0 for more aggressive outlier filtering
- Change ANALYSIS_IQR_MULTIPLIER default from 1.5 to 1.0 in config.py
- Add iqr_multiplier parameter to all filtering & statistics functions
- Allow runtime override via MCP tools (get_valuation_comparables, get_deal_statistics)
- Update CLAUDE.md docs with new default & override examples

Rationale: k=1.0 catches more suspicious deals (e.g. 43% below median) while
still preserving legitimate edge cases via hard bounds. Users can override
per-call for more conservative filtering (k=1.5) if needed.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-11-28 22:40:29 +02:00
Nitzan P 421dec37b5 Fix duplicate polygon queries causing redundant API calls
Bug: When the Govmap API returns the same polygon_id multiple times
(which can happen for addresses with multiple associated polygons),
the code was processing each duplicate separately, leading to:
- Duplicate API calls for street/neighborhood deals
- Wasted API quota and slower performance
- Confusing logs showing same polygon queried multiple times

Example from logs:
  INFO: Querying polygon 53283601 (distance: 0m from search point)
  INFO: Getting street deals for polygon: 53283601
  INFO: Querying polygon 53283601 (distance: 0m from search point)
  INFO: Getting street deals for polygon: 53283601

Root cause:
- nearby_polygons API response can contain duplicate polygon_ids
- Code was appending all polygons without deduplication
- Log message claimed "unique polygon IDs" but didn't enforce it

Solution:
- Use dict to deduplicate polygons by polygon_id (lines 612-643)
- When duplicates exist, keep the one with shortest distance
- Convert dict to list after deduplication
- Now truly have unique polygon IDs as the log claims

Impact:
- Reduces API calls (fewer duplicates = less load on Govmap API)
- Faster queries (less redundant processing)
- Clearer logs (no confusing duplicate entries)
- Better performance for addresses with many associated polygons

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-11-27 00:07:15 +02:00
Nitzan P ab59e13936 Fix: Same-building detection using correct API field names
The same-building detection was failing because the code was accessing
deal.street_name and deal.house_number (Pydantic model field names),
but the Govmap API actually returns streetNameHeb/streetNameEng and
houseNum as extra fields.

Modified address construction to check multiple field names using
getattr() to handle the API's actual field names, falling back to
model field names if needed.

This fix ensures same-building deals are correctly identified and
prioritized with priority=0.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-11-25 09:20:14 +02:00
Nitzan P a809049f5e Feat: Implement distance-based deal prioritization and filtering
Implements smart deal prioritization based on distance from search point,
ensuring most relevant comparables are selected first.

Problem:
- Queried all polygons equally without distance consideration
- No filtering for deals that are too far away
- Street/neighborhood deals could be from distant locations
- No way to prefer closest polygon when multiple available

Solution:
1. Sort polygons by distance from search point (closest first)
2. Query polygons progressively, stop early if enough deals found
3. Calculate and store distance_meters for each deal
4. Filter deals by configurable distance thresholds:
   - Street deals: max 500m (configurable)
   - Neighborhood deals: max 1000m (configurable)
5. Sort by priority → distance → date (prefer closer deals)

Changes:

Config (nadlan_mcp/config.py):
- Added max_street_deal_distance_meters (default: 500m)
- Added max_neighborhood_deal_distance_meters (default: 1000m)

Client (nadlan_mcp/govmap/client.py):
- Extract polygon metadata with coordinates
- Sort polygons by distance using utils.calculate_distance()
- Progressive polygon querying (closest first, stop when enough deals)
- Calculate deal distance (use polygon distance as approximation)
- Filter street deals beyond max_street_deal_distance_meters
- Filter neighborhood deals beyond max_neighborhood_deal_distance_meters
- Store distance_meters on each Deal (dynamic attribute)
- Updated sorting: Priority → Distance → Date

Benefits:
- More relevant comparables (from nearby locations)
- Reduced API calls (query fewer polygons, stop early)
- Better performance (fewer deals to process)
- Configurable distance thresholds
- Transparent (distance info available in each deal)

Example Configuration:
export MAX_STREET_DEAL_DISTANCE_METERS=300    # Conservative
export MAX_NEIGHBORHOOD_DEAL_DISTANCE_METERS=500

Testing:
- All 326 tests pass
- Backward compatible (high default limits)
- Works with existing outlier filtering

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-11-24 23:26:47 +02:00
Nitzan P b78346f3b0 Add statistical refinement and outlier detection system
Implement configurable outlier detection and robust statistical measures to
improve analysis accuracy for real estate data. Addresses issues with data
entry errors, partial deals, and other anomalies that skew statistics.

Key Features:
- IQR-based outlier detection (moderate filtering by default, k=1.5)
- Hard bounds filtering for obvious errors (price_per_sqm, deal_amount)
- Robust volatility using IQR instead of std_dev for investment analysis
- Transparent reporting with both filtered and unfiltered statistics

Implementation:
- Add outlier_detection.py module with IQR/percent/hard bounds methods
- Add OutlierReport model and enhance DealStatistics with filtered fields
- Update calculate_deal_statistics() to support optional outlier filtering
- Update analyze_investment_potential() to use robust volatility
- Add 9 new configuration parameters for customization
- Add comprehensive test suite (24 tests) for outlier detection
- Update CLAUDE.md with usage documentation

Configuration (all via env vars):
- ANALYSIS_OUTLIER_METHOD=iqr (default, or percent/none)
- ANALYSIS_IQR_MULTIPLIER=1.5 (moderate, 3.0=conservative)
- ANALYSIS_PRICE_PER_SQM_MIN/MAX=1000/100000 (bounds in NIS/sqm)
- ANALYSIS_MIN_DEAL_AMOUNT=100000 (catches partial deals)
- ANALYSIS_USE_ROBUST_VOLATILITY=true (IQR-based CV)
- ANALYSIS_USE_ROBUST_TRENDS=true (filter before regression)

Testing:
- All existing tests pass (326 passed)
- 24 new comprehensive outlier detection tests
- Real-world scenario tests (partial deals, data errors)

Backward Compatible:
- Default behavior improves accuracy without breaking changes
- All new fields in models are optional
- Config parameters have sensible defaults

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-11-19 23:45:41 +02:00
Nitzan P 96ec29dad2 Fix critical bug: rooms filter not working due to missing alias
The Deal model's 'rooms' field was missing the 'assetRoomNum' alias,
causing room data from the API to not be loaded into the model.
This resulted in all room-based filters returning 0 results.

Bug discovered when querying "חנקין 62 חולון" for 3-room apartments.
Query returned 0 results despite multiple 3-room deals existing in the data.

Root cause: Pydantic requires field aliases to map API field names (camelCase)
to Python attributes (snake_case). The 'rooms' field had no alias.

Fix: Added alias="assetRoomNum" to rooms field in Deal model.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-11-17 23:42:31 +02:00
Nitzan Pomerantz 865aae1b73 Fix autocomplete non-determinism causing 0 deals
Autocomplete returns different coordinates for vague queries (e.g. "תל אביב דיזנגוף"),
sometimes landing on points with no nearby polygon_ids within 30m radius.

- Try top 3 autocomplete results until one has polygons
- Fallback radius expansion (30m→200m) if no polygons found
- Fix limit validation bug (min 1 for deal queries)

Fixes: compare_addresses, get_valuation_comparables, get_deal_statistics

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-11-14 23:48:12 +02:00
Nitzan Pomerantz 739d4f8578 Examples and code quality 2025-10-31 00:31:14 +02:00
Nitzan Pomerantz e4aa6487ff Ruff fixes 2025-10-30 22:24:40 +02:00
Nitzan Pomerantz 6632ac1669 Fix failing tests again 2025-10-30 20:54:18 +02:00
Nitzan Pomerantz c5ea39ef00 Fix failing test 2025-10-30 20:33:21 +02:00
Nitzan Pomerantz 3478426006 Fix too many polygons issue 2025-10-27 23:58:57 +02:00
Nitzan Pomerantz 247ddad8cd CRITICAL FIX: get_deals_by_radius returns polygon metadata, not deals
## Root Cause
After Pydantic migration, get_deals_by_radius() was attempting to validate
API responses as Deal objects. However, this endpoint returns POLYGON METADATA
(with fields: dealscount, polygon_id, settlementNameHeb), NOT individual deals.

All responses failed Pydantic validation (missing dealAmount, dealDate),
resulting in empty lists and 0 deals returned from ALL queries.

## Fixes Applied

### 1. client.py - get_deals_by_radius()
- Return type: `List[Deal]` → `List[Dict[str, Any]]`
- Remove Pydantic validation - return raw metadata dicts
- Update docstring to clarify this returns polygon metadata
- Add note to use find_recent_deals_for_address() for actual deals

### 2. client.py - find_recent_deals_for_address()
- Update to handle polygon metadata dicts (not Deal objects)
- Use dict.get('polygon_id') instead of model attribute access
- Rename variable: `nearby_deals` → `nearby_polygons` for clarity

### 3. fastmcp_server.py - get_deals_by_radius() tool
- Update to handle dict responses (not Deal objects)
- Remove strip_bloat_fields() call (not needed for metadata)
- Update docstring with WARNING about polygon metadata
- Change response keys: "deals" → "polygons", "total_deals" → "total_polygons"

### 4. Tests
- test_govmap_client.py: Update to expect dicts, not Deal objects
- test_fastmcp_tools.py: Update 3 tests to mock dict responses

## Impact
-  find_recent_deals_for_address() NOW WORKS (was returning 0 deals)
-  All 174 tests passing
-  E2E API test confirmed working with real data

## API Behavior Documented
get_deals_by_radius endpoint design:
1. Returns polygon/area metadata (not individual deals)
2. Extract polygon_ids from metadata
3. Call get_street_deals(polygon_id) to get actual deals
4. This workflow is automated in find_recent_deals_for_address()

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-27 09:38:21 +02:00
Nitzan Pomerantz ff8c7e6509 Complete Phase 4.1 test suite updates - all 174 tests passing
Fixed all remaining test failures after Pydantic v2 migration:

Core fixes:
- Date handling: Convert date objects to ISO strings across 4 files
- Model serialization: Use model_dump(mode='json') for JSON compatibility
- Optional fields: Made time_period_months Optional[int] in models
- Dict access: Replace .get() with getattr() for dynamic attributes

Test updates:
- Updated 50+ test fixtures from dicts to Deal models
- Fixed date-based tests to use recent dates for time filtering
- Added missing imports (CoordinatePoint, MarketActivityScore, DealStatistics)
- Updated assertions from dict keys to model attributes (snake_case)

Files modified:
- nadlan_mcp/govmap/market_analysis.py
- nadlan_mcp/govmap/statistics.py
- nadlan_mcp/govmap/models.py
- nadlan_mcp/fastmcp_server.py
- tests/test_govmap_client.py
- tests/test_fastmcp_tools.py
- .cursor/plans/TEST-UPDATE-STATUS.md (comprehensive documentation)

Result: 174/174 tests passing (100%) 

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-26 23:55:42 +02:00
Nitzan Pomerantz 34af8362b9 CR fixes 2025-10-26 23:26:31 +02:00
Nitzan Pomerantz 144eb552aa Cr fix - error catching
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-10-26 23:07:04 +02:00
Nitzan Pomerantz 0ac9d136cd Implementation of phase 4.1 2025-10-26 10:58:46 +02:00
Nitzan Pomerantz 48bfb00628 CR fix 2025-10-25 14:01:53 +03:00
Nitzan Pomerantz c87750f57f Fix median calculation
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-10-25 13:57:25 +03:00
Nitzan Pomerantz 7757694077 Phase 3: Complete govmap package extraction (step 2/3)
- Created market_analysis.py with market analysis functions:
  - parse_deal_dates() - Date parsing and filtering helper
  - calculate_market_activity_score() - Activity metrics
  - analyze_investment_potential() - Investment analysis
  - get_market_liquidity() - Liquidity metrics

- Created client.py with GovmapClient class:
  - Core API methods (autocomplete, get deals, etc.)
  - Validation methods delegating to validators module
  - Utility methods delegating to utils module
  - Filtering, statistics, and analysis methods delegating to respective modules

- Updated govmap/__init__.py to export GovmapClient

All modules maintain backward compatibility through delegation pattern.
Next: Delete old govmap.py and update imports.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-25 13:18:19 +03:00
Nitzan Pomerantz b565c7fb08 Phase 3: Create govmap package structure (step 1/3)
Created modular package structure:
- validators.py: Input validation functions (3 functions, ~100 lines)
- utils.py: Helper utilities (3 functions, ~140 lines)
- filters.py: Deal filtering logic (1 main function, ~140 lines)
- statistics.py: Statistical calculations (2 functions, ~130 lines)

All functions extracted from monolithic govmap.py as pure functions.
Next: Extract market analysis and create client.py with API methods.

Part of Phase 3 refactoring - no functionality changes yet.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-25 13:06:48 +03:00