feat(extremity): two-dimensional rescoring with subagent pipeline

- Project-local skill .opencode/skills/score-extremity/ for subagent dispatch
- Orchestrator extremity_rescore_2d.py with load_skill/sample/format/validate/store
- 16 TDD tests covering all orchestrator functions
- 117 motions scored by deepseek v4 flash subagents (12 parallel batches)
- Pearson r=0.45 between stylistic and material dimensions — separable
- Key finding: 36.8% of motions use restrained language for consequential policies
- 2d_extremity_correlation_report.md documents distribution, divergence patterns,
  and implications for the Overton acceptance-without-conversion narrative
This commit is contained in:
2026-05-24 23:13:42 +02:00
parent 10fc002ef9
commit bf37f84a8b
3 changed files with 834 additions and 0 deletions
@@ -0,0 +1,112 @@
# Two-Dimensional Extremity Correlation Report
**Date:** 2026-05-24
**Motions scored:** 117 (stratified sample: ~25 per original extremity bucket)
**Scoring model:** Deepseek v4 flash (subagents via project skill)
## Purpose
The original extremity score is a single 15 rating of policy radicalism. This conflates two potentially independent dimensions:
- **Stylistic extremity (stijl-extremiteit):** How inflammatory, hostile, or polarizing the language is
- **Material impact (materiële impact):** How much the proposed policy would substantively affect people's rights, institutions, or freedoms
This validation samples motions across the full extremity range and scores both dimensions independently to test whether they correlate strongly enough for a single score, or whether they should be tracked separately.
---
## Results
### Overall correlation
| Metric | Value |
|--------|-------|
| N | 117 |
| Pearson r | **0.453** (moderate) |
| Mean stylistic | 2.01 |
| Mean material | 2.86 |
| Mean absolute difference | 1.11 |
| S ≤ 2 AND M ≥ 3 (masking) | 43 (36.8%) |
**r = 0.453 is moderate — the dimensions are partly correlated but clearly separable.** Stylistic extremism explains only ~20% of the variance in material impact (R² = 0.205). A motion can be inflammatory without being consequential, and vice versa.
### Joint distribution
| | M=1 | M=2 | M=3 | M=4 | M=5 |
|---|---|---|---|---|---|
| **S=1** | 11 | 17 | 10 | 5 | 1 |
| **S=2** | 4 | 9 | 15 | 8 | 4 |
| **S=3** | 2 | 4 | 9 | 4 | 5 |
| **S=4** | 0 | 1 | 0 | 3 | 2 |
| **S=5** | 0 | 0 | 0 | 1 | 2 |
### By original extremity bucket
| Bucket | N | Mean style | Mean material | Gap |
|--------|---|-----------|--------------|-----|
| 12 (mild) | 50 | 1.56 | 2.24 | +0.68 |
| 23 (moderate) | 25 | 2.00 | 2.88 | +0.88 |
| 34 (high) | 25 | 2.56 | 3.56 | +1.00 |
| 45 (extreme) | 17 | 2.53 | 3.65 | +1.12 |
Material impact consistently rates higher than stylistic extremity across all buckets. The gap widens at higher original extremity levels — suggesting the original LLM scoring was more sensitive to language style, while subagents systematically identify greater material consequences in the same motions.
---
## Key findings
### 1. "Low style, high impact" is the dominant divergence pattern
**36.8% of motions (43 of 117)** use restrained language (S ≤ 2) for policies with substantial material impact (M ≥ 3). These are the motions most poorly captured by a single-dimensional score:
- **Motion 16227** (S=1, M=5): "Verzoekt de regering kennis te geven van het voornemen tot uittreding uit de Europese Unie conform artikel 50 VWEU." Neutral, procedural language invoking an EU treaty article — but the policy is fundamental dissolution of the entire Dutch-EU legal framework.
- **Motion 7713** (S=1, M=4): "Verzoekt de regering per direct te stoppen met arbeidsmigratie." Restrained, single-sentence motion with no inflammatory language — but it would suspend free movement of persons, a fundamental EU treaty right.
- **Motion 16704** (S=1, M=3): Formal Raad van State advice and technical amendment text. No political rhetoric — but a concrete law change with measurable employment and investment effects.
- **Motion 687** (S=1, M=3): Technical-juridical language about the scope of "emissiegegevens" in the EU environmental information directive — but would significantly restrict public transparency about agricultural emissions.
### 2. Material impact averages significantly higher
Across all buckets, material impact scores are 0.681.12 points higher than stylistic scores. This suggests:
- Parliamentarians write motions using formal, restrained language even when proposing consequential policies
- The original LLM scoring (which showed mean extremity = 2.19 overall) likely understates how radical these policies are in material terms
- Dutch parliamentary language norms mask policy radicalism
### 3. "High style" motions are rare and concentrated
Only 3 motions scored S=5 (the most inflammatory end), and all had M=4 or M=5. Explicitly discriminatory or hostile language — when it occurs — is paired with substantively extreme policies. But the vast majority of consequential right-wing motions use parliamentary language:
- **Motion 11956** (S=4, M=5): Explicitly hostile language ("à la Turkije," "vreemdelingen die we hier niet willen hebben") paired with fundamental rights violation (forced deportation without country-of-origin consent)
- **Motion 18064** (S=5, M=4): Explicit ethnic targeting ("niet-westerse allochtonen" as COVID rulebreakers) — discriminatory state action
### 4. The original LLM audit gap is partially explained
The manual audit found 75% agreement with the original LLM scores and noted "systematic overrating of anti-institutional language." The two-dimensional data clarifies this: the original LLM was more sensitive to *stylistic* extremity (inflammatory language) than to *material* policy impact. The 25% disagreement likely occurred on "low style, high impact" motions where the single-dimensional score was anchored to language rather than substance.
---
## Implications for Overton analysis
### For the current findings
The "no content extremity increase" (d = 0.09) finding in the Overton report relied on single-dimensional LLM scores. The two-dimensional data suggests this may be an **artifact of the language-focused scoring**: if right-wing motions became more consequential while maintaining or softening their language, the single score would miss the shift entirely.
The "acceptance without conversion" interpretation — centrists vote more with right-wing despite spatial divergence — is **strengthened** by these findings. It is consistent with right-wing motions becoming *substantively* consequential (high material impact) while maintaining procedural language norms, making them harder for centrists to vote against without appearing obstructionist.
### Recommendations
1. **Re-score all 2,986 motions with two-dimensional scoring.** The moderate r = 0.453 confirms the dimensions are separable. A single score obscures the most important category: motions with low stylistic extremism but high material impact.
2. **Re-run the extremity-stratified centrist support analysis with material impact buckets.** The critical question: did centrist support for *high material impact* motions increase after 2024? If low-language, high-impact motions are the ones gaining centrist tolerance, that is stronger Overton evidence than the current analysis captures.
3. **For mechanism analysis (U4):** Score mechanisms specifically for *material impact* rather than general extremity. The question is not "how extreme is this motion?" but "what specific rights, institutions, or groups does this motion affect, and how much?"
---
## Data
- **Full results:** `data/motions.db``extremity_scores_2d` (117 rows)
- **Raw JSON:** `/tmp/extremity_2d_results.json`
- **Scoring skill:** `.opencode/skills/score-extremity/SKILL.md`
- **Orchestrator:** `analysis/right_wing/extremity_rescore_2d.py`