PUBLISHED BY ICSAC
When Reduction Becomes Lossy: An Intervention-Relative Test of Explanatory Boundaries
© 2026 Guillaume Vimeney. Published under CC BY 4.0; the author retains copyright.
Abstract
Can a restricted computational model predict well while omitting distinctions required by its explanatory task? We define boundary sufficiency relative to an outcome, a representation, and a declared family of input interventions. Exact sufficiency requires a common response law on every fiber of the retained representation throughout that family; observational predictive accuracy alone does not establish it. We prove that, under a fixed distribution, the excess optimal log loss from discarding distinctions equals conditional mutual information, and separate this irreducible within-regime loss from the transport loss of a predictor fitted under another regime. Three finite-state synthetic systems implement the test. A strongly lumpable four-state chain is an exact negative control: aggregation incurs zero regret under every tested intervention. A non-lumpable chain has observational regret 0.080 nat per outcome, but regret 0.368 nat after aggregate refitting and 0.733 nat when the observational predictor is frozen under uniform microstate intervention. In an agent–context model with the same observational regret, changing context to oppose the agent gives 1.460 nat frozen-predictor regret. Thirty seeded simulations reproduce these patterns with finite data; sensitivity sweeps include the zero-contrast counterexample. The resulting criterion diagnoses a failure of a specified computational boundary, not an ontological failure of complete microdescription. No empirical claim about neural or cognitive mechanisms follows from these constructed examples.
Code and data
https://doi.org/10.5281/zenodo.23042786, hosted by the author. ICSAC links to it and doesn’t host, run or maintain it.
Open review
The full panel report from the Institute's open-review process.
Read full panel review
Open review
This submission was evaluated by a panel of 10 independent advanced AI reviewers scoring six dimensions. Panel consensus was strong consensus.
Aggregate scores
| Dimension | Mean | Per-reviewer |
|---|---|---|
| Domain Fit | 5.0 | 5, 5, 5, 5, 5, 5, 5, 5, 5, 5 |
| Methodological Transparency | 4.9 | 5, 5, 5, 5, 5, 5, 5, 4, 5, 5 |
| Internal Consistency | 4.9 | 5, 5, 5, 5, 5, 5, 5, 4, 5, 5 |
| Citation Integrity | 4.4 | 4, 4, 5, 4, 5, 4, 4, 5, 4, 5 |
| Novelty Signal | 4.3 | 4, 4, 5, 5, 4, 4, 4, 4, 5, 4 |
| AI Provenance Signal | 5.0 | 5, 5, 5, 5, 5, 5, 5, 5, 5, 5 |
Reviewer assessments
Individual reviewer assessments are collapsed by default. Expand any row to read that reviewer's summary and per-dimension justification.
Reviewer 1 — RECOMMEND
Summary: A formally rigorous submission that gives an exact, testable criterion for when a coarse-grained or boundary-restricted model loses explanatory sufficiency relative to a declared intervention family, backed by exact finite-state proofs and reproducible finite-sample verification. Domain fit, methodological transparency, and internal consistency are all strong; citation integrity is solid with one minor unverifiable reference and two citations whose support is somewhat underspecified but not misused.
- Domain Fit (5/5): The submission is a formal/mathematical treatment of coarse-graining, Markov lumpability, and causal-abstraction sufficiency, with exact finite-state proofs (Section 3, the conditional-mutual-information identity R_D^obs = I_D(Y;S|B)) and computational verification (Section 4-5). This is squarely evaluable by a panel versed in information theory, Markov chains, and formal causal modeling; no specialist domain outside the panel's competence is required.
- Methodological Transparency (5/5): Transition matrices are given explicitly (Section 4.1, P_s=[(1-p_s)/2,...]), all parameter values are stated (p=(0.15,0.15,0.85,0.85) for System A, p=(0.10,0.90,0.90,0.10) for System B), seeds are disclosed (base seeds 1000-1029), sample sizes are specified (5,000 training inputs, 20,000 test outcomes per seed, 30 seeds), and the smoothing prior (Beta(1/2,1/2)) is named. The declarations section states the full generator, exact calculations, and machine-readable outputs are archived. The robustness grid (r and delta parameters, Section 4.3/5.5) and the exact numeric results in Table (Section 5.4) permit independent reimplementation.
- Internal Consistency (5/5): The reported numbers are internally coherent: System B's observational regret (0.079881 nat) is correctly derived from H_Bern(0.14)-H_Bern(0.1), and the refit/frozen/transport decomposition in Section 5.2 (0.368064 + 0.365321 ≈ 0.733385) matches the stated chain-rule identity R_i^frozen = R_i^refit + KL term. The A* matched-twin construction in Section 5.3 logically follows from the stated identification argument. The finite-sample results (Section 5.5) are explicitly reconciled with the exact-regret theorem rather than treated as contradictory, and the paper appropriately restricts its claims (Section 6-7) to match what the proofs actually establish.
- Citation Integrity (4/5): Per independent verification, all citations except Potochnik 2017 resolve to real works with matching claim contexts (Brier 1950, Woodward 2003, Kemeny and Snell 1960, Hoel et al. 2013, Craver 2007, Pearl 2009, etc.), and Potochnik 2017 is unverifiable rather than fabricated -- it is used for a general, low-stakes point about idealization serving different aims (Section 2), so an unverifiable citation there is a minor concern, not a load-bearing failure. The misattribution check flags two citations as unclear support: Geiger et al. 2021 ('Neural Abstractions') is cited alongside Rubenstein et al. 2017 and Geiger et al. 2025 for the commuting-under-interventions claim (Section 2), and Shalizi and Moore 2025 is cited for macrostate construction tied to observed distinctions and dynamics (Section 2, Section 6). Both citations sit in a paragraph where the paper itself explicitly narrows its claim relative to the cited work ('Our criterion is deliberately less ambitious', 'our treatment fixes a partition and tests it, rather than claiming to discover a unique optimal one'), which mitigates but does not eliminate the concern.
- Novelty Signal (4/5): The three-way decomposition of observational regret, interventional refit regret, and frozen-predictor transport regret (Section 3), and specifically the A* matched-twin non-identifiability construction (Section 5.3) showing that observational (B,Y) data alone cannot distinguish a boundary-sufficient system from an insufficient one, is a genuinely new formal diagnostic not reducible to existing causal-emergence (Hoel et al. 2013) or causal-abstraction (Geiger et al. 2021) results -- the paper explicitly and correctly distinguishes its estimand from both (Section 6). The contribution is incremental relative to established tools (lumpability, conditional mutual information, proper scoring rules) rather than field-advancing, which caps it short of a 5.
- AI Provenance Signal (5/5): The submission discloses AI assistance ('OpenAI ChatGPT contributed to literature research, synthesis analysis, and editorial review') per the declarations section, which is the disclosed-and-substantive case the rubric scores cleanly rather than penalizes. The text contains specific, verifiable numerical claims throughout (e.g., '0.733385 nat', 'p < 0.001'-style precision without an actual p-value substitute, matched exactly to the underlying exact-enumeration calculations), a dedicated limitations/falsifiability section (Section 7) that states concrete conditions under which the criterion would be falsified, explicit engagement with a competing framework (the causal-emergence literature, Section 6, correctly distinguishing sign and estimand rather than ignoring it), and non-uniform section lengths driven by content. No fabricated citations were found and no generic/templated abstract language is present -- the abstract names exact regret values and specific systems.
Reviewer 2 — RECOMMEND
Summary: A rigorous, formally grounded paper that provides a testable framework for evaluating explanatory boundaries in computational models. The methodology is transparent, the claims are internally consistent, and the work makes a clear conceptual contribution. The panel recommends acceptance.
- Domain Fit (5/5): The submission uses formal mathematical and computational methods (Markov chains, information theory, causal abstraction) to construct falsifiable claims about explanatory boundaries. The work is squarely within the panel's competence to evaluate as it involves theory, computation, and formal modeling. It does not rely on field-specific empirical expertise the panel lacks.
- Methodological Transparency (5/5): The methodology is fully specified: exact finite-state enumeration, explicit transition matrices, parameter values, intervention families, loss functions, and estimation procedures are provided. Code and data are declared available. The distinction between exact expectations and finite-sample Monte Carlo checks is clearly drawn. All assumptions (positivity, invariant response mechanism) are stated.
- Internal Consistency (5/5): Claims follow logically from the formal framework and results. The negative control (System A) behaves exactly as predicted. The non-lumpable case (B) and agent-context case (C) demonstrate the predicted regret decompositions. The observational twin construction (A⋆) is a logically derived consequence that strengthens the identification argument. No contradictions between claims and evidence are present.
- Citation Integrity (4/5): All citations verified as real and used in load-bearing ways (e.g., Kemeny and Snell for lumpability, Woodward for intervention, Craver for mechanistic mapping). The two citations flagged as unverifiable (Potochnik 2017, Geiger et al. 2021) are not central to the technical claims; Potochnik is cited for a general point about idealizations, and Geiger et al. 2021 is cited alongside Geiger et al. 2025 which is verifiable. No evidence of fabrication or systematic misattribution. Score reduced slightly due to the unverifiable citations being used in a supporting rather than load-bearing role.
- Novelty Signal (4/5): The paper presents a novel synthesis: combining Markov lumpability, causal abstraction, and information-theoretic regret decomposition into a testable profile for explanatory boundary errors. The separation of observational, refitted, and frozen-predictor regret is a useful conceptual contribution. The twin construction (A⋆) is a clever identification argument. The work is incremental in that it builds on established concepts (lumpability, intervention, proper scoring rules) but applies them in a new integrated framework.
- AI Provenance Signal (5/5): The prose is specific, technically precise, and contains domain expertise signals (e.g., 'fiberwise condition', 'transport loss', 'strong lumpability', 'Beta(1/2,1/2) smoothing'). The abstract makes concrete claims with numerical results. The methodology section describes actual methods with explicit parameters. The limitations section is substantive. AI assistance is disclosed for literature research and editorial review, not for core analysis. No generic template phrasing, padded word count, or fabricated methodology is detected.
Reviewer 3 — RECOMMEND
Summary: The submission presents a rigorous, novel framework for testing explanatory sufficiency of reduced models under interventions, with fully transparent methodology and sound internal logic. All citations are appropriate, and the work fits squarely within the institute’s scope.
- Domain Fit (5/5): The work employs formal mathematical definitions, Markov chain analysis, information‑theoretic quantities, and synthetic computational experiments to make falsifiable claims. The panel has the expertise to evaluate these methods, so it is fully in scope.
- Methodological Transparency (5/5): All synthetic systems are fully specified (state spaces, transition matrices, intervention policies). Exact enumerations, seed numbers, sample sizes, smoothing priors, and code/data availability are described, enabling independent replication.
- Internal Consistency (5/5): The paper’s definitions (boundary sufficiency, regret decompositions) logically lead to the derived identities and the reported experimental results. No contradictions are observed between theory, methodology, and results.
- Citation Integrity (5/5): All cited works are real and are used appropriately to support claims about proper scoring rules, lumpability, mechanistic explanation, intervention semantics, and causal abstraction. The only citation that cannot be verified from the excerpt (Potochnik 2017) is not fabricated and does not appear to be mis‑attributed.
- Novelty Signal (5/5): The manuscript introduces a new intervention‑relative test of explanatory boundaries, formalizes a boundary‑sufficiency profile, and decomposes loss into observational, refit, and transport components—none of which have been presented together in prior work.
- AI Provenance Signal (5/5): The text contains detailed technical derivations, specific numerical results, and concrete synthetic examples. It does not exhibit generic, filler language or missing substantive content.
Reviewer 4 — RECOMMEND
Summary: This submission presents a rigorous, novel formal framework for testing explanatory sufficiency in reduced models using intervention-relative criteria. The methodology is transparent and reproducible, claims are logically consistent, and the work makes a significant contribution to the theory of causal abstraction and model boundaries. It is recommended for acceptance.
- Domain Fit (5/5): The submission uses formal mathematical and computational methodology to make falsifiable claims about explanatory sufficiency in reduced models. It is squarely within ICSAC's scope of complexity, causal abstraction, and intervention analysis, and the panel is fully competent to evaluate it.
- Methodological Transparency (5/5): Methods are fully specified: exact finite-state systems are defined, response mechanisms are given, interventions are explicitly described, and estimation procedures are detailed. Code and data are synthetic and fully reproducible from the provided generator. Parameters, simulation seeds, and evaluation metrics are all reported.
- Internal Consistency (5/5): The claims follow directly from the formalism and constructed examples. The three systems logically instantiate the theoretical framework, and results are derived from exact enumeration and finite-sample simulations that align with the theory. No contradictions are present.
- Citation Integrity (4/5): All cited works verified as real. One unverifiable citation (Potochnik 2017) is minor and not load-bearing. Misattribution concerns for Geiger et al. 2021 and Shalizi and Moore 2025 exist due to lack of accessible abstracts, but other citations are accurately used and support their claims, particularly Craver, Woodward, Pearl, and lumpability references.
- Novelty Signal (5/5): The paper introduces a novel formal criterion for boundary sufficiency that distinguishes observational from interventional predictive loss. The decomposition of regret into refit and transport components, and the construction of analytically identical observational systems with divergent interventional outcomes, represent a new diagnostic for explanatory boundaries.
- AI Provenance Signal (5/5): The prose is precise, technically specific, and free of padding or generic phrasing. The work demonstrates deep engagement with formalism, counterexamples, and limitations. No red flags for AI provenance: structure is content-driven, and there is no citation stuffing or methodological vagueness.
Reviewer 5 — RECOMMEND
Summary: A rigorous formal treatment of explanatory boundaries using information-theoretic measures and synthetic Markov systems. The work is methodologically transparent, internally consistent, and provides a clear, testable criterion for distinguishing observational success from interventional sufficiency.
- Domain Fit (5/5): The submission uses formal mathematical and computational methodology (Markov chains, information theory, log loss) to make falsifiable claims about explanatory boundaries. The work is entirely within the panel's competence to evaluate.
- Methodological Transparency (5/5): The methodology is exceptionally transparent. The author provides exact transition matrices, specific probability vectors, explicit loss functions, and detailed parameters for Monte Carlo simulations (seeds, sample sizes). The distinction between exact enumeration and finite-sample estimation is clearly handled.
- Internal Consistency (5/5): The claims follow logically from the formal framework. The use of a strongly lumpable chain as a negative control and the construction of the 'matched twin' (A*) to prove the insufficiency of observational data are logically rigorous and well-executed.
- Citation Integrity (5/5): Citations are real and load-bearing. The work correctly situates itself against causal abstraction (Geiger), causal emergence (Hoel), and mechanistic explanation (Craver), using these references to define the specific narrow gap the paper fills rather than as mere veneer.
- Novelty Signal (4/5): The submission provides a novel, testable diagnostic for 'boundary sufficiency' that distinguishes between observational adequacy and interventional invariance. It adds a useful operational layer to the existing discourse on causal abstraction and coarse-graining.
- AI Provenance Signal (5/5): The prose is highly specific, technically dense, and engaged. The methodology is concrete and verifiable, and the results are derived from specific constructed systems rather than generic templates. There are no signs of low-engagement provenance.
Reviewer 6 — RECOMMEND
Summary: The submission gives a formally precise, computationally verified criterion for when a restricted representation loses outcome-relevant information under a declared intervention family, with an exact negative control, a non-lumpable failure case, an agent-context case, and a novel non-identifiability ('twin') construction. Methodology is transparent (exact parameters, seeds, sample sizes) and citations are load-bearing with only minor unresolved relevance flags on two references.
- Domain Fit (5/5): The submission uses formal mathematical methodology (Markov lumpability, conditional mutual information, proper scoring rules) and finite-state computational experiments to make falsifiable claims about when a reduced representation preserves an outcome's response law under intervention. This is core formal/computational epistemology and information theory that the panel can evaluate directly from the derivations and reported numbers in Sections 3-5.
- Methodological Transparency (5/5): The transition mechanism is given in closed form (P_s vector, Section 4.1), all parameter values are stated (p=(0.15,0.15,0.85,0.85) for System A, p=(0.10,0.90,0.90,0.10) for System B, the X/C response table for System C), and the finite-sample protocol specifies seeds (1000-1029), training/test set sizes (5,000/20,000), and the Beta(1/2,1/2) smoothing used for groupwise frequency estimation (Section 4.3). The Declarations section states that the full generator, exact calculations, and machine-readable outputs are included in the archive, satisfying the reproducibility bar for computational work.
- Internal Consistency (5/5): The reported figures are internally verifiable from the stated model: H_Bern(0.1)=0.325083 nat and H_Bern(0.14)=0.404964 nat (Section 5.2) compute correctly from the given Bernoulli parameters, and the chain-rule decomposition R_i^frozen=R_i^refit+KL[...] (Section 3) is used consistently to explain why System B's frozen regret (0.733385 nat) exceeds its refit regret (0.368064 nat) by exactly the reported transport term (0.365321 nat). The System A negative control (zero regret under every tested intervention) and the A* 'observationally indistinguishable sufficient twin' construction (Section 5.3) are used consistently to support the paper's identification claim rather than contradicting it.
- Citation Integrity (4/5): Per the independent verification pass, all 19 DOI-checked citations resolve to real works and Potochnik 2017 is unverifiable-but-not-fabricated. On load-bearing use: Hoel et al. 2013/Hoel 2017 are correctly distinguished from the paper's own criterion in Section 2 and 6 ('effective-information comparisons ... can use different intervention distributions ... cannot establish that complete microscale knowledge has lost information'), Kemeny and Snell 1960 directly supports the strong-lumpability definition used in Section 3, and Woodward 2003/Pearl 2009 are used narrowly for the intervention/invariance and observational-vs-interventional distinctions actually deployed in the formalism. The misattribution flags on Geiger et al. 2021 and Shalizi and Moore 2025 are noted but their cited context (causal-abstraction commuting-under-interventions; macrostate construction tied to observed distinctions and dynamics) matches their titles closely enough that citation-stuffing is not clearly established; this is treated as a minor open concern rather than a load-bearing failure.
- Novelty Signal (4/5): The paper's contribution is a specific, testable decomposition (observational regret vs. refit regret vs. frozen-predictor transport regret, Section 3) rather than a restatement of existing causal-abstraction or causal-emergence results, and the Section 5.3 'A*' twin construction is a genuine identification result: two systems with identical observational (B,Y) law but opposite boundary-sufficiency status. This builds on established tools (Markov lumpability, conditional mutual information, proper scoring rules) rather than introducing a new mathematical object, which bounds the novelty below field-advancing.
- AI Provenance Signal (5/5): The abstract commits to specific, checkable numbers (0.080 nat observational regret, 0.368 nat refit, 0.733 nat frozen, 1.460 nat for the opposing-context case) rather than generic framing, and Section 7 states concrete falsification conditions for the paper's own criterion. The submission explicitly engages competing accounts (Section 2 and 6 distinguish the paper's claim from Hoel et al.'s causal-emergence results and from Craver/Kaplan-Craver mechanistic-adequacy requirements) rather than presenting itself as uncontested. AI assistance is disclosed for 'literature research, synthesis analysis, and editorial review' in the Declarations section, consistent with disclosed-and-substantive use rather than low-engagement generation.
Reviewer 7 — RECOMMEND
Summary: A rigorous, transparent, and internally consistent formal treatment of explanatory boundary errors using intervention-relative sufficiency. The paper provides a clear test with controlled synthetic examples, separates distinct estimands often conflated in the literature, and honestly states its limitations. It is a solid contribution suitable for inclusion in the ICSAC community.
- Domain Fit (5/5): The work uses formal mathematical and computational methodology (Markov chains, information theory, causal abstraction) to make falsifiable claims about explanatory boundaries. It is squarely within ICSAC's scope of complexity science, emergence, and substrate-independence, and the panel can credibly evaluate the formal and computational methods.
- Methodological Transparency (5/5): The methodology is fully specified: exact finite-state enumeration, explicit transition matrices, parameter values, intervention families, loss functions, and estimation procedures. The supplement and code archive are referenced. All assumptions (positivity, known mechanism, perfect measurement) are stated. The finite-sample checks with 30 seeds and sensitivity sweeps are described.
- Internal Consistency (5/5): Claims follow logically from the formal framework and results. The negative control (system A) demonstrates zero regret as predicted. The non-lumpable case (B) and agent-context case (C) show the expected patterns of observational vs. interventional regret. The decomposition of frozen regret into refit regret and transport KL divergence is mathematically derived and empirically confirmed. The twin construction (A⋆) is a coherent identification argument.
- Citation Integrity (4/5): All citations verified as real and used in load-bearing ways (Brier, Gneiting & Raftery for proper scoring rules; Kemeny & Snell for lumpability; Woodward for intervention; Craver & Kaplan for mechanistic explanation; Hoel et al. for causal emergence; Pearl for do-operator). The two citations flagged as unverifiable (Potochnik 2017) or with unclear relevance (Geiger et al. 2021, Shalizi & Moore 2025) are not central to the paper's core claims; the paper's own framework is self-contained. No fabrication detected. Score reduced slightly due to the unverifiable/misattribution flags, but the load-bearing citations are solid.
- Novelty Signal (4/5): The paper presents a novel formal criterion (boundary sufficiency profile) that cleanly separates observational, refitted, and frozen interventional regret. The decomposition of frozen regret into refit and transport terms is a useful conceptual contribution. The twin construction showing observational indistinguishability of sufficient and insufficient models is insightful. The work is not field-advancing (score 5) because it builds on well-known concepts (lumpability, conditional mutual information, causal abstraction) but applies them in a new, carefully integrated way.
- AI Provenance Signal (5/5): The prose is specific, technically precise, and contains domain expertise (e.g., 'fiberwise condition', 'strong lumpability', 'transport loss', 'Beta(1/2,1/2) smoothing'). The abstract makes concrete claims with numerical results. The methodology section describes actual methods with explicit parameters. The limitations section is honest and detailed. The AI assistance disclosure is transparent. No signs of generic template text, padding, or fabricated methodology.
Reviewer 8 — RECOMMEND
Summary: The submission presents a rigorous, reproducible theoretical framework for assessing when coarse‑grained models lose explanatory power under specified interventions. The methodology is transparent, the claims are internally consistent, citations are sound, and the contribution is novel enough to merit acceptance into the community.
- Domain Fit (5/5): The work employs formal mathematical definitions, synthetic finite‑state Markov models, information‑theoretic analysis, and explicit computational experiments, all of which constitute scientific methodology. The panel has the expertise to evaluate these methods without requiring specialized empirical domain knowledge.
- Methodological Transparency (4/5): The manuscript specifies the state spaces, transition matrices, intervention policies, loss functions, seed numbers, sample sizes, and provides a code archive statement. While the description is detailed enough for replication, the lack of a direct code URL or full algorithmic pseudocode prevents a perfect score.
- Internal Consistency (4/5): Definitions (boundary sufficiency, regret decompositions) are introduced and subsequently used to derive the reported results. The logical flow from theorems to synthetic examples is coherent, with no evident contradictions, though some proofs are sketched rather than fully detailed.
- Citation Integrity (5/5): All cited works (e.g., Brier 1950, Gneiting & Raftery 2007, Kemeny & Snell 1960, Craver 2007, Pearl 2009, Hoel et al. 2013, Geiger et al. 2021/2025, Rubenstein et al. 2017) are real and appropriately support the statements made. No fabricated or mis‑attributed citations were detected; the only unverifiable reference (Potochnik 2017) is not essential to the core argument.
- Novelty Signal (4/5): The paper introduces a new "intervention‑relative boundary sufficiency" test and a decomposition of optimal regret into observational, refitted, and frozen components. While building on existing concepts (causal abstraction, lumpability, effective information), the specific formulation and synthetic validation appear novel.
- AI Provenance Signal (5/5): The text is highly specific, contains detailed mathematical derivations, concrete numerical results, and domain‑appropriate citations. There is no evidence of generic filler, padded abstract, or fabricated methodology typical of low‑engagement AI‑generated prose.
Reviewer 9 — RECOMMEND
Summary: This work presents a rigorous, novel framework for testing explanatory sufficiency in coarse-grained models, supported by exact and finite-sample analyses. The methodology is transparent, the claims are logically consistent, and the contribution advances the understanding of causal abstraction in computational systems.
- Domain Fit (5/5): The submission uses formal mathematical and computational methodology to make falsifiable claims about explanatory sufficiency in reduced models, squarely within ICSAC's scope of complexity, causal abstraction, and substrate-independent information processing. The panel can fully evaluate the theoretical and computational content without requiring external empirical expertise.
- Methodological Transparency (5/5): The methods are fully specified: exact finite-state constructions, explicit transition mechanisms, precise definitions of regret and sufficiency, and a reproducible simulation setup with seeds, sample sizes, and estimation procedures. Code and data generation are provided, enabling complete replication.
- Internal Consistency (5/5): The claims follow logically from the formal framework and computational results. Definitions of boundary sufficiency are operationalized in the experiments, and the three-case analysis consistently illustrates the theoretical distinctions. The limitations section acknowledges assumptions without undermining core results.
- Citation Integrity (4/5): All cited works verified as real; no fabrication. Potochnik 2017 is unverifiable but not central to the load-bearing argument. Misattribution concerns are minimal: Geiger et al. 2021 and Shalizi and Moore 2025 lack accessible abstracts, but their cited roles are plausible given the context. Core claims are supported by confirmed sources like Pearl 2009, Woodward 2003, and Kemeny and Snell 1960.
- Novelty Signal (5/5): The paper introduces a novel formal criterion—boundary sufficiency—distinguishing observational from interventional predictive adequacy. The decomposition of regret into refitted and frozen components, and the demonstration of identification failure under observational equivalence, represent a new diagnostic framework for causal abstraction in computational models.
- AI Provenance Signal (5/5): The prose is technically precise, avoids generic phrasing, and engages deeply with domain-specific concepts. No signs of template writing, padding, or disengaged citation use. The structured argument, formal proofs, and targeted simulations reflect substantive engagement with the material.
Reviewer 10 — RECOMMEND
Summary: A rigorous and transparent formal analysis of explanatory boundaries using synthetic Markov systems. The work provides a clear mathematical framework for diagnosing when a reduced computational model fails under intervention, making a solid contribution to the study of coarse-graining and abstraction.
- Domain Fit (5/5): The submission uses formal mathematical and computational methodology (Markov chains, information theory, and synthetic simulations) to make falsifiable claims about explanatory boundaries. The panel can credibly evaluate these formal claims.
- Methodological Transparency (5/5): The methodology is exceptionally transparent. The author provides exact transition matrices, explicit probability distributions, specific seed values for Monte Carlo runs, and a clear decomposition of loss (observational, refit, and frozen). The use of synthetic systems with finite-state enumeration ensures the results are auditable.
- Internal Consistency (5/5): The claims follow logically from the formal framework. The distinction between observational regret and interventional regret is mathematically grounded in conditional mutual information and KL divergence, and the synthetic cases (A, B, and C) consistently demonstrate the predicted behaviors.
- Citation Integrity (5/5): The citations are used in a load-bearing manner to situate the work. The author correctly distinguishes their narrow diagnostic approach from the broader goals of causal abstraction (Geiger et al.) and causal emergence (Hoel et al.), using the references to define the boundaries of their own contribution.
- Novelty Signal (4/5): The submission proposes a novel, testable criterion for 'boundary sufficiency' that separates within-regime loss from transport loss. While it builds on established concepts like Markov lumpability and causal abstraction, the specific operationalization of a 'sufficiency profile' for explanatory boundaries is a meaningful contribution.
- AI Provenance Signal (5/5): The prose is highly specific, technically dense, and deeply engaged with the subject matter. The presence of exact numerical results (e.g., 0.080 nat, 0.368 nat), detailed synthetic system designs, and a nuanced engagement with counterarguments (e.g., the A* construction) provides strong signals of high-engagement human provenance.
Reviews at ICSAC are open and transparent. AI tooling helps the panel draft and structure each review; final acceptance decisions rest with the curation team. Reviews are published alongside acceptance for accountability; individual reviewer identities are abstracted to keep focus on the assessment rather than the tooling behind it.
Review Quality Control audit
A second-pass audit of the panel's own review against the Institute's published rubric.
Read RQC audit
Review Quality Control
Review Quality Control: passed.
This audit quality checks each AI reviewer's assessment for rubric adherence, internal consistency, specificity, and institutional voice. It is published alongside the panel review so the quality of the review process is as auditable as the review itself.
Notes
- No dimension across any of the ten valid reviewers scored at or below 2, and no injection signal was detected in any reviewer.
- The Potochnik 2017 unverifiable citation and the Geiger et al. 2021 / Shalizi and Moore 2025 misattribution flags are handled consistently across all ten reviewers as minor, non-load-bearing concerns — no drift or inconsistency in how the panel treated this shared caveat.
Reviewer Quality Control Audit
| Reviewer | Rubric Adherence | Internal Consistency | Specificity | Tone |
|---|---|---|---|---|
| Reviewer 1 | 5/5 | 5/5 | 5/5 | 5/5 |
| Reviewer 2 | 5/5 | 5/5 | 5/5 | 5/5 |
| Reviewer 3 | 5/5 | 5/5 | 4/5 | 5/5 |
| Reviewer 4 | 5/5 | 5/5 | 4/5 | 5/5 |
| Reviewer 5 | 5/5 | 5/5 | 4/5 | 5/5 |
| Reviewer 6 | 5/5 | 5/5 | 5/5 | 5/5 |
| Reviewer 7 | 5/5 | 5/5 | 5/5 | 5/5 |
| Reviewer 8 | 5/5 | 5/5 | 4/5 | 5/5 |
| Reviewer 9 | 5/5 | 5/5 | 4/5 | 5/5 |
| Reviewer 10 | 5/5 | 5/5 | 5/5 | 5/5 |
Reviewer 1
- Rubric Adherence (5/5): All six dimensions (domain_fit, methodological_transparency, internal_consistency, citation_integrity, novelty_signal, ai_provenance_signal) are scored on the 1-5 scale with one justification each, correctly named.
- Internal Consistency (5/5): The 4/5 citation_integrity score is supported by a justification describing one unverifiable-but-not-fabricated reference and two minor misattribution flags treated as non-load-bearing; the RECOMMEND recommendation matches the uniformly strong per-dimension narrative.
- Specificity (5/5): Justifications cite exact section numbers, the conditional-mutual-information identity, specific transition-matrix parameters, seed ranges, and precise numeric results (e.g., 0.733385 nat decomposition).
- Tone (5/5): Institutional third person throughout ('the submission', 'the paper'), no emojis, no pleasantries, findings stated directly.
Reviewer 2
- Rubric Adherence (5/5): All six required dimensions present, correctly named, 1-5 scale respected, one justification each.
- Internal Consistency (5/5): The 4/5 scores on citation_integrity and novelty_signal are each explained with a specific reason (unverifiable citation in a supporting role; incremental synthesis of established concepts), and the summary's 'clear conceptual contribution' framing matches the RECOMMEND call.
- Specificity (5/5): Cites named systems (System A, B, twin A*), specific citation roles (Kemeny and Snell for lumpability, Woodward for intervention, Craver for mechanistic mapping), and concrete technical terms rather than generic praise.
- Tone (5/5): Consistent institutional voice, no first-person lapses, no emojis, no softening language.
Reviewer 3
- Rubric Adherence (5/5): All six dimensions scored with correct names and scale, one justification per dimension.
- Internal Consistency (5/5): All-high scores (5s across the board) are each backed by a specific supporting claim, and the summary's characterization as 'rigorous, novel' with no reservations is consistent with the absence of any dimension below 5.
- Specificity (4/5): Justifications name specific systems and cited authors (Brier 1950, Gneiting & Raftery) but lean more on category-level description ('state spaces, transition matrices, intervention policies') than on exact numeric results or section citations found in other reviewers.
- Tone (5/5): Institutional phrasing ('the panel', 'the work') throughout, no emojis or pleasantries.
Reviewer 4
- Rubric Adherence (5/5): All six dimensions present with correct names and 1-5 scale.
- Internal Consistency (5/5): The 4/5 citation_integrity score is explained by the same unverifiable/misattribution concerns cited elsewhere in the panel, consistent with the otherwise strong justification set and the RECOMMEND summary.
- Specificity (4/5): References specific authors (Craver, Woodward, Pearl, lumpability literature) and named systems, but several justifications ('claims follow directly from the formalism') are more general than the numerically anchored justifications seen in other reviewers.
- Tone (5/5): Consistent institutional third-person voice, no emojis or cushioning language.
Reviewer 5
- Rubric Adherence (5/5): All six dimensions scored, correctly named, 1-5 scale used consistently.
- Internal Consistency (5/5): The 4/5 novelty_signal score is justified with a specific reason (adds an operational layer to existing discourse rather than a wholly new object), consistent with the overall RECOMMEND summary.
- Specificity (4/5): Cites specific methodological elements (exact transition matrices, probability vectors, Monte Carlo seeds) and named authors used in a load-bearing way, though with less exact numeric-result citation than the first-seat reviewers.
- Tone (5/5): Institutional voice maintained throughout, no emojis or pleasantries.
Reviewer 6
- Rubric Adherence (5/5): All six dimensions scored with correct names and scale, one justification each.
- Internal Consistency (5/5): The 4/5 citation_integrity and novelty_signal scores are each supported with the same specific reasoning as the corresponding first-pass reviewer, and the summary's qualified praise ('minor unresolved relevance flags') matches the RECOMMEND recommendation.
- Specificity (5/5): Justifications reference exact entropy values (H_Bern(0.1)=0.325083 nat), specific sections, and named constructions (System A negative control, A* twin), mirroring the specificity of the first-pass first-seat reviewer.
- Tone (5/5): Institutional third person throughout, no emojis or pleasantries.
Reviewer 7
- Rubric Adherence (5/5): All six dimensions present, correctly named, 1-5 scale respected.
- Internal Consistency (5/5): The 4/5 citation_integrity and novelty_signal scores are explained consistently with the corresponding first-pass reviewer, and the summary matches the RECOMMEND call with no contradictory language.
- Specificity (5/5): Cites specific technical terms ('fiberwise condition', 'Beta(1/2,1/2) smoothing'), named systems A/B/C, and specific citation roles rather than generic praise.
- Tone (5/5): Consistent institutional voice, no emojis or softening hedges.
Reviewer 8
- Rubric Adherence (5/5): All six dimensions scored with correct names and 1-5 scale.
- Internal Consistency (5/5): The 4/5 methodological_transparency and internal_consistency scores are each explained with a specific caveat (no direct code URL; some proofs sketched rather than fully detailed), consistent with the RECOMMEND summary's otherwise positive framing.
- Specificity (4/5): Names specific cited authors and works (Brier 1950, Kemeny & Snell 1960, Hoel et al. 2013) but several justifications describe methodology at a categorical level ('state spaces, transition matrices, intervention policies') rather than citing exact figures or section numbers.
- Tone (5/5): Institutional phrasing maintained, no emojis or pleasantries.
Reviewer 9
- Rubric Adherence (5/5): All six dimensions scored, correctly named, 1-5 scale used.
- Internal Consistency (5/5): The 4/5 citation_integrity score is explained by the same unverifiable/misattribution reasoning used across the panel, and the summary's assessment aligns with the RECOMMEND recommendation without contradiction.
- Specificity (4/5): References specific cited authors (Pearl 2009, Woodward 2003, Kemeny and Snell 1960) and the three-case analysis, but several justifications remain at the level of general methodological description rather than citing exact numeric results.
- Tone (5/5): Institutional third-person voice throughout, no emojis or cushioning.
Reviewer 10
- Rubric Adherence (5/5): All six dimensions present with correct names and 1-5 scale, one justification each.
- Internal Consistency (5/5): The 4/5 novelty_signal score is supported by a specific reason (builds on established concepts but with a meaningful operationalization), consistent with the RECOMMEND summary.
- Specificity (5/5): Cites exact numeric results (0.080 nat, 0.368 nat), named synthetic systems A/B/C, and the A* construction rather than generic phrasing.
- Tone (5/5): Institutional voice maintained, no emojis or pleasantries.
Review Quality Control is an internal ICSAC audit of the panel review itself. The four dimensions above are published as part of ICSAC's open review commitment.
How to cite this paper
The science is the author's, cited at its DOI as an article of Persistence; the curation record is the Institute's. See /how-to-cite for the full model, including the annual print edition.
The work — cite this when discussing the paper's findings
Plain text (APA-style)
Vimeney, G. (2026). When Reduction Becomes Lossy: An Intervention-Relative Test of Explanatory Boundaries. Persistence, 1, Article 001. https://doi.org/10.67697/icsac.2026.001
BibTeX
@article{vimeney2026when,
author = {Vimeney, Guillaume},
title = {{When Reduction Becomes Lossy: An Intervention-Relative Test of Explanatory Boundaries}},
journal = {Persistence},
volume = {1},
pages = {001},
year = {2026},
publisher = {Institute for Complexity Science and Advanced Computing},
doi = {10.67697/icsac.2026.001},
url = {https://doi.org/10.67697/icsac.2026.001},
}RIS
TY - JOUR AU - Vimeney, Guillaume PY - 2026 TI - When Reduction Becomes Lossy: An Intervention-Relative Test of Explanatory Boundaries T2 - Persistence VL - 1 SP - 001 PB - Institute for Complexity Science and Advanced Computing DO - 10.67697/icsac.2026.001 UR - https://doi.org/10.67697/icsac.2026.001 ER -
The curation record — cite this when discussing the review itself
Plain text (APA-style)
ICSAC Curation Panel (2026). Curation record for "When Reduction Becomes Lossy: An Intervention-Relative Test of Explanatory Boundaries" by Vimeney, G. Institute for Complexity Science and Advanced Computing. https://icsacinstitute.org/publications/when-reduction-becomes-lossy
BibTeX
@misc{vimeney2026when-record,
author = {{ICSAC Curation Panel}},
title = {{Curation record for "When Reduction Becomes Lossy: An Intervention-Relative Test of Explanatory Boundaries"}},
year = {2026},
publisher = {Institute for Complexity Science and Advanced Computing},
url = {https://icsacinstitute.org/publications/when-reduction-becomes-lossy},
note = {Panel reviews, RQC audit, and curation decision for Vimeney, G.},
}RIS
TY - GEN AU - ICSAC Curation Panel PY - 2026 TI - Curation record for "When Reduction Becomes Lossy: An Intervention-Relative Test of Explanatory Boundaries" PB - Institute for Complexity Science and Advanced Computing UR - https://icsacinstitute.org/publications/when-reduction-becomes-lossy N1 - Panel reviews, RQC audit, and curation decision for Vimeney, G. ER -