Commit Graph
20 Commits
Author SHA1 Message Date
sgeboers 846e9cf67f fix: import canonical parties from config, simplify theme consistency check 2026-04-05 10:05:46 +02:00
sgeboers bad9cd758d docs: add SVD theme divergence solution doc and validation hook 2026-04-05 09:58:53 +02:00
sgeboers bfe37c6806 fix: align report generation with JSON output for positive/negative separation
Bug: report_per_component used scored[:args.report_top_n] which took
top N by score (all positive for components with only positive scores).
JSON correctly separated positive and negative poles.

Fix: Use same positive/negative separation logic for report as JSON.
2026-04-04 18:19:31 +02:00
sgeboers 33edb334c4 feat: implement exclusive SVD motion assignment with label review report
- Each motion now assigned to exactly one component (highest absolute score)
- Added --exclusive flag (default: True) for backward compatibility
- Added markdown report generation with motion details for label review
- Added --report-top-n for report size (default: 20 per component)
- Updated JSON output with 'exclusive' flag for transparency
2026-04-04 17:55:28 +02:00
sgeboers c9c59dd166 feat(diagnostics): enhance trajectory diagnostic script with real data mode 2026-04-01 01:59:41 +02:00
sgeboers 9f98dbae60 Add debug st.info before st.plotly_chart to diagnose invisible chart 2026-03-31 01:49:38 +02:00
sgeboers 88110b0aaa Fix update_existing_motions: single write connection and module-level duckdb import
Use one DuckDB write connection for the entire update loop instead of
opening/closing per row, wrapped in try/finally for proper cleanup.
Move 'import duckdb' to module level with other imports.
2026-03-29 23:28:40 +02:00
sgeboers be8887f6f8 Add --skip-details, --update-existing flags to download_past_year.py with tests
Enable backfilling body_text for existing motions that lack it (2016-2018 data).
New extract_besluit_id() and update_existing_motions() helpers support the
--update-existing mode, while --no-skip-details enables detail fetching during
normal downloads. Includes 7 tests covering URL parsing, DB update flow, and
argparse wiring.
2026-03-29 23:25:04 +02:00
sgeboers 9daa899885 fix: remove motion title truncation, add SVD JSON generation script
Removes the raw_title[:80] cap on expander labels so full titles show.
Adds scripts/generate_svd_json.py to regenerate top_svd_top_motions.json
from any SVD window after a recompute.
2026-03-25 01:15:42 +01:00
sgeboers d1faf2b3e4 feat(mindmodel): add CLI wrapper, edge-case tests, and manifest schema tests 2026-03-24 22:41:28 +01:00
sgeboers f77875ed54 feat(mindmodel): add CLI wrapper and tests 2026-03-24 21:29:10 +01:00
sgeboers a74e6006f5 feat(mindmodel): add validator and tests 2026-03-24 21:28:17 +01:00
sgeboers 7bd7d0d18c feat(mindmodel): add checks utilities and tests 2026-03-24 21:26:38 +01:00
sgeboers 2efd7ba3a0 feat(mindmodel): add manifest loader and tests 2026-03-24 21:24:58 +01:00
sgeboers eb73275f32 feat(mp-quiz): add MP quiz tab and DB helpers; add design and plan docs 2026-03-24 19:59:25 +01:00
sgeboers b09e580f65 feat: motion content enrichment pipeline hardening
- ai_provider_wrapper: retry/fallback with exponential backoff, None sentinel for failed items
- text_pipeline: use wrapper, return 5-tuple (stored, skipped_existing, skipped_no_text, errors, failed_ids)
- similarity/compute: filter trivial 1.0 matches on identical short titles (<12 chars)
- rerun_embeddings: --retry-missing mode, calls ensure_text_embeddings_for_ids on failed ids
- sync_motion_content: per-ext_id retries, HTTPAdapter pool, --max-body-workers CLI flag, audit on failure
- qa_similarity script: samples motions, writes JSON ledger to thoughts/ledgers/
- All tests green: 61 passed, 2 skipped
2026-03-23 21:31:39 +01:00
sgeboers 2891e9ee70 feat: add StemAtlas Streamlit app, explorer, Docker deployment, blog charts 2026-03-22 22:38:17 +01:00
sgeboers daa22c5e2b feat: complete parliamentary embedding pipeline with full historical coverage
- Add fused (SVD + text) embedding pipeline for annual windows 2016-2026
- Fix store_fused_embedding duplicate bug: DELETE before INSERT (idempotent)
- Add --text-batch-size CLI flag to run_pipeline.py (default 200)
- Add explicit --start-date/--end-date to download_past_year.py
- Backfill mp_votes for all motions (party-level votes, 111k new rows)
- Add similarity cache recompute: 212k rows across 9 annual windows
- Improve ai_provider retry logic, text_pipeline batching
- Improve analysis/political_axis PCA handling and visualizations
- Add diagnostic/utility scripts: compare_svd, generate_compass, inspect_axis, etc.
- Untrack data/motions.db (3.6GB binary), add to .gitignore with outputs/
- Update continuity ledger with full session state
2026-03-22 16:08:06 +01:00
sgeboers 847b783877 fix(pipeline): fix API pagination, add skip_details fast path, bulk mp_votes insert
- _get_voting_records returns (records, besluit_meta) tuple; paginate via Besluit?expand=Stemming (469/mo vs 8400)
- get_motions(skip_details=True) bypasses per-motion detail chain (3 HTTP calls/motion)
- extract_mp_votes rewritten: bulk DataFrame insert (80k rows in 1.9s), includes party-level actors
- run_pipeline.py fixed: pass db_path not db, handle dict/int return types
- download_past_year.py: skip_details=True default, limit-per-chunk default 50000
2026-03-21 23:24:06 +01:00
sgeboers a36e6cba4e feat(pipeline): implement parliamentary embedding pipeline MVP
- Add 4 migration files: mp_votes, mp_metadata, svd_vectors, fused_embeddings
- Extend database.py with 5 new helper methods and table init
- Add pipeline/ package: extract_mp_votes, fetch_mp_metadata, text_pipeline,
  svd_pipeline (with Procrustes alignment), fusion
- Add full test suite (17 tests) covering all pipeline modules and migrations
- Fix Procrustes alignment bug: scipy scale is a norm value, not a multiplier
- Fix DuckDB date type handling in test assertions (datetime.date vs string)
- Remove duckdb.py shim; tests now run against real duckdb + scipy via uv

Ref: thoughts/shared/plans/2026-03-21-parliamentary-embedding-pipeline-plan.md
2026-03-21 22:31:22 +01:00