Measuring Semantic Shift and Recovery Across LLM Post-Training and Aggregation
A methods-first paper frame. T^FIELD is one candidate intervention; v0.8.1 is explicitly NO-GO and no quantitative result is claimed.
Research question
How can semantic distribution shift, recovery-specific effects and aggregation coverage in LLM systems be measured without mistaking sampling noise, mapping failure, verbosity or post-hoc selection for genuine contraction?
Primary contribution sought
A reproducible framework for null-calibrated directional semantic shift, recovery lift over a neutral control, and aggregation coverage with false-representation control, while keeping token, semantic, trajectory and visible-output objects separate.
Current primary endpoints
- E1 STAGE: one preregistered DEFICIT_EXCESS contrast, with EXPANSION_EXCESS mandatory companion.
- E2 RECOVERY: RECOVERY_LIFT from paired recovery vs neutral recheck forks.
- E3 TOPOLOGY: AGG_COVERAGE with mandatory REPRESENTATION_FPR.
- E4 R_META: structural-recurrence rate difference versus pooled non-structural controls.
Statistical repair
The v0.8.1 hard-threshold ES_LOSS/REC_GAIN path is historical and NO-GO for confirmation. Raw positive-part deficits are not interpreted without an explicit null calibration. Within-task inference follows the experimental randomization/permutation structure; generalized claims use tasks as the upper-level unit.
Hostile alternative
A non-T^ pipeline using independent sampling, frozen semantic clustering and budgeted coverage selection. If this baseline matches or exceeds T_WEAVE at equal or lower cost without worse false representation, no T^FIELD-specific contribution is claimed.
Publication criterion
The paper remains meaningful under a null or negative result for T^FIELD. The intended contribution is measurement and attribution first. Confirmatory data collection is prohibited until the v0.8.2 freeze gate opens.