arXiv is now an independent nonprofit! Learn more
License: CC BY-NC-SA 4.0
arXiv:2609.03973v1 [cs.AI] 03 Sep 2026

Common-Witness Certificates and Sharp Feature Bounds for Counterfactual Image Auditing

Usef Faghihi ††thanks: Department of Mathematics and Computer Science, Université du Québec à Trois-Rivières, Trois-Rivières, Quebec, Canada (). Both authors contributed equally: each contributed 50% of the theoretical work and 50% of the experimental work. Email: usef.faghihi@uqtr.ca    Amir Saki11footnotemark: 1
Abstract

Structured image editors may satisfy every regional plausibility constraint separately although no single latent explanation is compatible with the complete output. We formulate this local-to-global failure through a common-witness grade and its witness nerve. The framework separates auditing from causal identification: shared exogeneity alone permits every coupling of the regime marginals, whereas a scientifically justified witness relation yields exactly the relation-supported couplings and hence sharp feature-level partial-identification bounds. For quasiconvex regional losses, classical Helly theory gives finite incompatibility certificates. For the labeled witness structures generated by the audit, we derive a heterogeneous action-stratified certificate, a sharp tolerant certificate for finite witness atlases, and a blocker-hypergraph repair formula. Fractional Helly theory converts dense local compatibility into a large jointly coherent subset. Simultaneous confidence regions for regime marginals provide finite-sample outer confidence bounds for the complete identified interval; Bonferroni–Clopper–Pearson bands give a nonasymptotic multinomial construction. In controlled MNIST rotations and finite paired studies on Morpho-MNIST and smallNORB, archived summaries report coherent-pair acceptance of 0.9593–0.9778, while regional-action patchworks retain local acceptance of 0.9702–0.9804 but have global acceptance of 0–0.1138. Synthetic studies exercise sharp bounds, certificate recovery, and structured computation up to 10410^{4} feature states and 10510^{5} regional constraints. The method audits a prespecified feature relation; it does not identify unrestricted pixel-level counterfactuals.

keywords
counterfactual images, partial identification, Helly theorem, optimal transport, finite-sample inference, infeasibility certificates
††runningheads: Common-Witness Image Auditing / U. Faghihi and A. Saki
MSC
52A35, 62G15, 62P30, 68T45, 90C05

1 Introduction

Counterfactual image editing asks what the image of the same unit would have been under another intervention. In an augmented structural causal model (ASCM), both potential images are evaluated at the same exogenous state. That semantic convention does not reveal their joint distribution. Pan and Bareinboim show that unrestricted pixel-level counterfactuals are generally not identified from image–label samples, even when a latent causal diagram is supplied [21]. Their feature-based relaxation is therefore a bound on declared care-set quantities, not recovery of the true pixel coupling.

We study a complementary auditing failure. An edited image can satisfy each protected-region constraint with a different latent explanation while no one explanation satisfies all regions simultaneously. Formally, a witness for each region need not be one witness for every region. This global-witness gap is invisible to an audit that aggregates independent local passes.

The gap becomes scientifically informative only when the witness family is anchored outside the candidate pair. Examples include a validated renderer, paired interventions, or a design guarantee. A similarity score fitted only to unpaired images is not, by itself, cross-world information. We encode an externally justified witness family by regional costs, use their common sublevel intersections to audit an image pair, and project the resulting admissibility rule to a prespecified finite feature. The projection is then an explicit assumption in a support-only partial-identification model; it is not inferred from the causal graph.

The contribution is a certificate-to-inference pipeline:

  1. 1.

    We turn regional costs into a graded feasibility complex. For the labeled witness structures produced by this audit, we derive exact certificate-size and repair formulas: a sharp q⁡(s+1)q(s+1) certificate for a finite atlas with qq witnesses and ss exceptional roles, a blocker-hypergraph description of minimal explanations, and a heterogeneous action-stratified certificate with size ∑a(da+1)\sum_{a}(d_{a}+1). Classical Helly, union-of-convex-set, fractional-Helly, and tolerance results supply the underlying intersection principles [6, 2, 8, 20, 15].

  2. 2.

    We separate this audit from identification. A boundary proposition records that shared exogeneity alone admits every coupling. Once an external relation is declared, its relation-supported couplings give the exact identified set in a stated support-only two-regime feature model. No coupling inside that set is silently selected.

  3. 3.

    We propagate simultaneous marginal confidence regions through the coupling program to obtain finite-sample outer confidence bounds for the complete oracle interval. Controlled and finite paired studies illustrate the local-to-global separation, while synthetic programs check the predicted interval and computational behavior. We report the empirical information boundary explicitly. The accompanying public repository contains the implementation, tests, configurations, and retained outputs; external datasets must be obtained from their original sources.

The main line is

regional costs\displaystyle\text{regional costs} ⟶common-witness certificate⟶declared feature relation\displaystyle\longrightarrow\ \text{common-witness certificate}\ \longrightarrow\ \text{declared feature relation}
⟶sharp coupling bounds⟶outer confidence bounds.\displaystyle\longrightarrow\ \text{sharp coupling bounds}\ \longrightarrow\ \text{outer confidence bounds}.

Technical extensions, full protocols, and secondary analyses are indexed in the Supplementary Material.

2 Related work and positioning

Pan and Bareinboim establish the nonidentification boundary for counterfactual image editing and propose feature-level causal bounds [21]; later work develops a disentangled causal latent editor [22]. Our object is different: an externally anchored same-witness audit followed by a support-only partial-identification analysis. It neither replaces their construction nor weakens their impossibility theorem.

Partial identification of joint potential-outcome functionals has a long history [18, 3, 26, 11, 10, 27]. Coupling and transport methods make the remaining cross-world ambiguity explicit [25, 23, 7]. The linear program and its transport dual are established tools. The contribution here is their placement after an auditable, externally sourced relation, together with a precise support-only sharpness statement.

Helly’s theorem and fractional Helly theory control intersections of convex families [6, 14, 13]. Amenta and Eckhoff–Nischke treat unions of convex components and the generalized pigeonhole mechanism [2, 8]; tolerance variants are also established [20, 15]. Our certificate theorems are exact labeled specializations for the witness structures arising in the audit, with sharpness and blocker-based repair. We do not claim that finite-feature bounding, transport duality, or Helly theory is new.

Finally, Theorem 7 is confidence-set propagation. Its importance here is the estimand: it covers the entire identified interval. Clopper–Pearson bands are conservative but nonasymptotic [5]; more efficient simultaneous multinomial regions can be substituted if their joint coverage is proved [12, 24, 19].

3 Setting and the common-witness audit

3.1 Feature target and identification boundary

Let X∈{0,1}X\in\{0,1\} denote an intervention and Ix∈ℐI_{x}\in\mathcal{I} the corresponding potential image, where ℐ\mathcal{I} is a standard Borel space. We prespecify a measurable finite feature

Ψ:ℐ→𝒴,𝒴={1,…,K},Yx=Ψ⁡(Ix),\Psi:\mathcal{I}\to\mathcal{Y},\qquad\mathcal{Y}=\{1,\ldots,K\},\qquad Y_{x}=\Psi(I_{x}),

and assume that the single-world laws μx​(y)=ℙ⁡(Yx=y)\mu_{x}(y)=\mathbb{P}(Y_{x}=y) are identified from randomized intervention data or another valid causal argument. Image samples alone do not generally provide this identification.

Write πy​y′=ℙ⁡(Y0=y,Y1=y′)\pi_{yy^{\prime}}=\mathbb{P}(Y_{0}=y,Y_{1}=y^{\prime}). Our targets are bounded linear functionals

Hh​(π)=∑y,y′h⁡(y,y′)​πy​y′.H_{h}(\pi)=\sum_{y,y^{\prime}}h(y,y^{\prime})\pi_{yy^{\prime}}. (1)

Transition probabilities, average feature changes, and numerators of conditional feature queries have this form. A conditional query with a positive identified denominator μ0​(B)\mu_{0}(B) is obtained by dividing the corresponding numerator by that fixed quantity.

Proposition 1 (Shared-exogeneity saturation).

Let μ0,μ1\mu_{0},\mu_{1} be probability laws on a common standard Borel space. Every coupling π∈Π⁡(μ0,μ1)\pi\in\Pi(\mu_{0},\mu_{1}) is the joint potential-outcome law of a two-regime structural model with one exogenous variable shared across interventions.

Proof.

On the probability space with law π\pi, let U=(U0,U1)U=(U_{0},U_{1}) be the coordinate map and set f⁡(x,u0,u1)=uxf(x,u_{0},u_{1})=u_{x}. Then Y=f⁡(X,U)Y=f(X,U) satisfies Yx=UxY_{x}=U_{x} and ℒ⁡(Y0,Y1)=π\mathcal{L}(Y_{0},Y_{1})=\pi. Both worlds use the same realization of UU.

Proposition 1 is a boundary lemma, not a novelty claim. It shows that using the same witness twice cannot by itself overcome the causal hierarchy. Information enters only through restrictions on admissible exogenous states, response functions, or cross-world pairs.

3.2 Regional losses, grade, and nerve

Let 𝒲\mathcal{W} be a declared witness space and let cr​(w,i,j)≥0c_{r}(w;i,j)\geq 0, r∈[m]={1,…,m}r\in[m]=\{1,\ldots,m\}, measure the violation of role rr by witness ww for the factual–candidate pair (i,j)(i,j). When (i,j)(i,j) is fixed or clear from context, we suppress it from the notation. Thus, for example, cr​(w)c_{r}(w) denotes cr​(w,i,j)c_{r}(w;i,j); the same convention applies to all witness sets, grades, and tolerant quantities defined below.

The ASCM response type UU and the auditing witness ww are distinct objects unless an external scientific argument identifies them. At tolerance ε\varepsilon, put

Arε​(i,j)={w∈𝒲:cr​(w,i,j)≤ε}.A_{r}^{\varepsilon}(i,j)=\{w\in\mathcal{W}:c_{r}(w;i,j)\leq\varepsilon\}. (2)

Compactness and lower semicontinuity are imposed whenever an attained minimizer is used; the measurable version is given in Supplement S1.

Definition 2 (Common-witness grade and nerve).

For nonempty S⊆[m]S\subseteq[m], define

F⁡(S,i,j)=infw∈𝒲maxr∈S⁡cr​(w,i,j),ρ⁡(i,j)=F⁡([m],i,j).F(S;i,j)=\inf_{w\in\mathcal{W}}\max_{r\in S}c_{r}(w;i,j),\qquad\rho(i,j)=F([m];i,j). (3)

The witness nerve at scale ε\varepsilon is

Kε​(i,j)={∅}∪{∅≠S⊆[m]:⋂r∈SArε​(i,j)≠∅}.K_{\varepsilon}(i,j)=\{\varnothing\}\cup\left\{\varnothing\neq S\subseteq[m]:\bigcap_{r\in S}A_{r}^{\varepsilon}(i,j)\neq\varnothing\right\}. (4)

If T⊆ST\subseteq S, then F⁡(T,i,j)≤F⁡(S,i,j)F(T;i,j)\leq F(S;i,j). Consequently KεK_{\varepsilon} is a simplicial complex, maximal faces are maximal jointly explainable role sets, and minimal nonfaces are minimal incompatibility explanations. Under attainment, (i,j)(i,j) has one global witness exactly when ρ⁡(i,j)≤ε\rho(i,j)\leq\varepsilon.

To permit a prespecified number of regional exceptions, define

F[s]​(S,i,j)=infw∈𝒲minD⊆S|D|≤s⁡maxr∈S∖D​cr​(w,i,j),max⁡∅:=0,F^{[s]}(S;i,j)=\inf_{w\in\mathcal{W}}\min_{\begin{subarray}{c}D\subseteq S\\ |D|\leq s\end{subarray}}\max_{r\in S\setminus D}c_{r}(w;i,j),\qquad\max\varnothing:=0, (5)

and ρs=F[s]​([m])\rho_{s}=F^{[s]}([m]). The integer ss is a within-pair role budget; it is not a probability mass of inadmissible feature pairs.

3.3 From image witnesses to a feature relation

The audit induces the existential feature projection

ℛε={(y,y′):there is a lift (i,j) with Ψ⁡(i)=y, Ψ⁡(j)=y′,and ρ⁡(i,j)≤ε}.\mathcal{R}_{\varepsilon}=\left\{(y,y^{\prime}):\begin{array}[]{l}\text{there is a lift $(i,j)$ with $\Psi(i)=y$, $\Psi(j)=y^{\prime}$,}\\[-2.84526pt] \text{and $\rho(i,j)\leq\varepsilon$}\end{array}\right\}. (6)

The causal model must separately assume ℙ{(Y0,Y1)∈ℛε}=1\mathbb{P}\{(Y_{0},Y_{1})\in\mathcal{R}_{\varepsilon}\}=1. The relation is therefore an externally declared Layer-3 restriction, not a consequence of its definition. An existential feature projection can also be an outer relaxation of the image-level problem; it is lossless only under a feature-saturation or conditional-fiber condition (Supplement S3). Without an external anchor, the honest default is ℛε=𝒴2\mathcal{R}_{\varepsilon}=\mathcal{Y}^{2}.

4 Combinatorial certificates and repair

We first give the finite-atlas result used directly by the experiments, then place its convex analogue in the classical Helly context.

4.1 Finite atlases

Theorem 3 (Exact finite-atlas tolerant certificate).

Suppose 𝒲\mathcal{W} is finite with q=|𝒲|q=|\mathcal{W}| and 0≤s<m0\leq s<m. Then

ρs​(i,j)=max∅≠S⊆[m]|S|≤min⁡{q⁡(s+1),m}⁡F[s]​(S,i,j).\rho_{s}(i,j)=\max_{\begin{subarray}{c}\varnothing\neq S\subseteq[m]\\ |S|\leq\min\{q(s+1),m\}\end{subarray}}F^{[s]}(S;i,j). (7)

Thus tolerant incompatibility has a certificate using at most q⁡(s+1)q(s+1) roles. The bound is best possible.

Proof.

For fixed ww, arrange its mm costs in nonincreasing order, c[1]​(w)≥⋯≥c[m]​(w)c_{[1]}(w)\geq\cdots\geq c_{[m]}(w). Deleting at most ss roles leaves the smallest possible maximum c[s+1]​(w)c_{[s+1]}(w), hence

ρs=minw∈𝒲⁡c[s+1]​(w).\rho_{s}=\min_{w\in\mathcal{W}}c_{[s+1]}(w).

For every ww, choose TwT_{w} containing s+1s+1 roles with its largest costs and put S⋆=⋃wTwS_{\star}=\bigcup_{w}T_{w}. Then |S⋆|≤q⁡(s+1)|S_{\star}|\leq q(s+1). The (s+1)(s+1)st largest cost of each ww on S⋆S_{\star} equals its (s+1)(s+1)st largest cost on [m][m], so F[s]​(S⋆)=ρsF^{[s]}(S_{\star})=\rho_{s}. Monotonicity gives the reverse inequality for every subfamily.

For sharpness, take m=q⁡(s+1)m=q(s+1) and partition the roles into disjoint blocks BwB_{w} of size s+1s+1. Fix 0≤b<d0\leq b<d and set cr​(w)=dc_{r}(w)=d for r∈Bwr\in B_{w} and cr​(w)=bc_{r}(w)=b otherwise. The complete family has tolerant grade dd. Every proper role set omits a role from some BwB_{w}; for that witness, at most ss retained roles have cost dd, so its tolerant grade is at most bb. Thus all q⁡(s+1)q(s+1) roles can be necessary.

The proof is an exact consequence of the labeled finite-atlas structure. We do not claim that tolerant Helly theory is new; generic tolerance transfers and modern tolerance complexes are studied in [20, 15]. The audit-specific value of Theorem 3 is its sharp certificate and the explicit repair formula below.

Proposition 4 (Blocker duality and exact repair).

For fixed ε\varepsilon, let

Bwε={r:cr​(w,i,j)>ε},Gwε=[m]∖Bwε.B_{w}^{\varepsilon}=\{r:c_{r}(w;i,j)>\varepsilon\},\qquad G_{w}^{\varepsilon}=[m]\setminus B_{w}^{\varepsilon}.

Then

F[s](S;i,j)≤ε⟺∃w∈𝒲:|S∩Bwε|≤s.F^{[s]}(S;i,j)\leq\varepsilon\quad\Longleftrightarrow\quad\exists w\in\mathcal{W}:\ |S\cap B_{w}^{\varepsilon}|\leq s. (8)

Minimal ss-tolerant incompatibility explanations are the inclusion-minimal (s+1)(s+1)-fold transversals of the bad-role hypergraph. Each has at most q⁡(s+1)q(s+1) roles. For s=0s=0,

Kε=⋃w∈𝒲2Gwε,sε⋆:=min⁡{s:ρs≤ε}=minw∈𝒲⁡|Bwε|.K_{\varepsilon}=\bigcup_{w\in\mathcal{W}}2^{G_{w}^{\varepsilon}},\qquad s_{\varepsilon}^{\star}:=\min\{s:\rho_{s}\leq\varepsilon\}=\min_{w\in\mathcal{W}}|B_{w}^{\varepsilon}|. (9)

Proof.

For fixed ww, all retained costs are at most ε\varepsilon after at most ss deletions exactly when at most ss members of SS belong to BwεB_{w}^{\varepsilon}. This proves (8); negating it yields the multiple-transversal description. Choosing s+1s+1 bad roles for every witness gives the size bound. Finally, |Bwε||B_{w}^{\varepsilon}| is exactly the number of deletions required for witness ww, and minimizing over ww proves (9).

Rowwise order selection computes ρs\rho_{s}, one valid certificate, and the exact repair number in O⁡(q​m)O(qm) selection time (or O⁡(q​m​log⁡m)O(qm\log m) by sorting). Enumeration of all minimal transversals can still be exponential [1].

4.2 Convex and action-stratified atlases

For a compact convex witness set of affine dimension dd, continuous quasiconvex role losses have convex closed sublevel sets. The classical Helly theorem then gives

ρ⁡(i,j)=max∅≠S⊆[m]|S|≤d+1⁡F⁡(S,i,j).\rho(i,j)=\max_{\begin{subarray}{c}\varnothing\neq S\subseteq[m]\\ |S|\leq d+1\end{subarray}}F(S;i,j). (10)

Indeed, at the maximum local grade every d+1d+1 sublevel sets intersect, so the complete family intersects. Equation (10) is a direct application of Helly’s theorem, not a new intersection theorem.

Many audits have a discrete requested action and continuous nuisance parameters. The resulting witness space is a labeled disjoint union rather than one convex set.

Theorem 5 (Heterogeneous action-stratified certificate).

Suppose 𝒲=⨆a=1q𝒲a\mathcal{W}=\bigsqcup_{a=1}^{q}\mathcal{W}_{a}, where 𝒲a\mathcal{W}_{a} is nonempty, compact, convex, and has affine dimension dad_{a}. Assume every role loss is continuous and quasiconvex on each stratum. Define

Fa​(S)=minw∈𝒲a⁡maxr∈S​cr​(w),F⁡(S)=mina⁡Fa​(S),F_{a}(S)=\min_{w\in\mathcal{W}_{a}}\max_{r\in S}c_{r}(w),\qquad F(S)=\min_{a}F_{a}(S),

and H=min⁡{m,∑a=1q(da+1)}H=\min\{m,\sum_{a=1}^{q}(d_{a}+1)\}. Then

F⁡([m])=max∅≠S⊆[m]|S|≤H⁡F⁡(S).F([m])=\max_{\begin{subarray}{c}\varnothing\neq S\subseteq[m]\\ |S|\leq H\end{subarray}}F(S). (11)

The bound is sharp for every dimension list when m=∑a(da+1)m=\sum_{a}(d_{a}+1).

Proof.

Let ρa=Fa​([m])\rho_{a}=F_{a}([m]). Applying (10) within stratum aa gives Sa⊆[m]S_{a}\subseteq[m], |Sa|≤da+1|S_{a}|\leq d_{a}+1, with Fa​(Sa)=ρaF_{a}(S_{a})=\rho_{a}. Put S⋆=⋃aSaS_{\star}=\bigcup_{a}S_{a}. Monotonicity gives

ρa=Fa​(Sa)≤Fa​(S⋆)≤Fa​([m])=ρa\rho_{a}=F_{a}(S_{a})\leq F_{a}(S_{\star})\leq F_{a}([m])=\rho_{a}

for every aa. Hence F⁡(S⋆)=mina⁡ρa=F⁡([m])F(S_{\star})=\min_{a}\rho_{a}=F([m]) and |S⋆|≤H|S_{\star}|\leq H.

For sharpness, create one block of da+1d_{a}+1 roles (a,j)(a,j) per stratum and take 𝒲a=Δda\mathcal{W}_{a}=\Delta^{d_{a}}. For x∈Δdbx\in\Delta^{d_{b}} set

c(a,j)​(b,x)={(da+1)​xj,a=b,0,a≠b.c_{(a,j)}(b,x)=\begin{cases}(d_{a}+1)x_{j},&a=b,\\ 0,&a\neq b.\end{cases}

The complete block has grade one in its stratum, so the full-family grade is one. If (a,j)(a,j) is removed, choose stratum aa and vertex eje_{j}; every remaining cost is zero. Thus every proper subfamily has grade zero.

Classical results for unions of convex sets already imply closely related q⁡(d+1)q(d+1) Helly numbers [2, 8]. The content of Theorem 5 is the audit-specialized exact grade equality, heterogeneous dimension sum, and sharp labeled construction, not a claim of a fundamentally new general Helly theorem.

The witness nerve also quantifies approximate coherence. Let h=d+1h=d+1 and let Nh​(ε)N_{h}(\varepsilon) count its feasible hh-faces. If DεD_{\varepsilon} is the largest number of roles explained by one witness and tε=m−Dεt_{\varepsilon}=m-D_{\varepsilon}, Kalai’s exact fractional-Helly bound gives

Nh​(ε)≤(mh)−(d+tεh).N_{h}(\varepsilon)\leq\binom{m}{h}-\binom{d+t_{\varepsilon}}{h}. (12)

In particular, if at least an α\alpha fraction of the hh-subsets are feasible, one witness explains at least

⌈[1−(1−α)1/(d+1)]​m⌉\left\lceil\bigl[1-(1-\alpha)^{1/(d+1)}\bigr]m\right\rceil (13)

roles [14, 13, 9]. The complete proof and exact integer inversion are in Supplement S4. Dense local compatibility yields a large coherent core, not global compatibility.

If estimated losses obey max⁡supwr⁡|c^r​(w,i,j)−cr​(w,i,j)|≤δ\max_{r}\sup_{w}|\widehat{c}_{r}(w;i,j)-c_{r}(w;i,j)|\leq\delta, then every tolerant grade changes by at most δ\delta and

Kε−δ[s]⊆K^ε[s]⊆Kε+δ[s].K^{[s]}_{\varepsilon-\delta}\subseteq\widehat{K}^{[s]}_{\varepsilon}\subseteq K^{[s]}_{\varepsilon+\delta}. (14)

This is a deterministic robustness statement; statistical use requires an independently justified uniform error bound. Finite-net outer approximation and exact nerve recovery away from critical grades are in Supplement S4.

5 Sharp feature bounds

For a declared relation ℛ⊆𝒴2\mathcal{R}\subseteq\mathcal{Y}^{2}, let

Πℛ(μ0,μ1)={π∈ℝ+K×K:∑y′πy​y′=μ0(y),∑yπy​y′=μ1(y′),πy​y′=0outside ℛ}.\Pi_{\mathcal{R}}(\mu_{0},\mu_{1})=\left\{\pi\in\mathbb{R}_{+}^{K\times K}:\sum_{y^{\prime}}\pi_{yy^{\prime}}=\mu_{0}(y),\quad\sum_{y}\pi_{yy^{\prime}}=\mu_{1}(y^{\prime}),\quad\pi_{yy^{\prime}}=0\ \text{outside }\mathcal{R}\right\}. (15)

Equivalently, π​𝟏=μ0\pi\mathbf{1}=\mu_{0} fixes row sums and π⊤​𝟏=μ1\pi^{\top}\mathbf{1}=\mu_{1} fixes column sums. The transpose is needed because left multiplication by π⊤\pi^{\top} sums the original matrix down its rows, producing one total for each column.

The support-only feature model contains every two-regime feature SCM with these marginals and ℙ{(Y0,Y1)∈ℛ}=1\mathbb{P}\{(Y_{0},Y_{1})\in\mathcal{R}\}=1, with no other response function, latent-DAG, or full-image restriction.

Theorem 6 (Sharp support-only identified interval).

If Πℛ​(μ0,μ1)\Pi_{\mathcal{R}}(\mu_{0},\mu_{1}) is nonempty, then the sharp identified set for (1) in the support-only feature model is

[minπ∈Πℛ​(μ0,μ1)⁡Hh​(π),maxπ∈Πℛ​(μ0,μ1)⁡Hh​(π)].\left[\min_{\pi\in\Pi_{\mathcal{R}}(\mu_{0},\mu_{1})}H_{h}(\pi),\max_{\pi\in\Pi_{\mathcal{R}}(\mu_{0},\mu_{1})}H_{h}(\pi)\right]. (16)

Both endpoints and every intermediate value are attainable.

Proof.

The feasible set is a nonempty compact convex polytope and HhH_{h} is linear, so its image is a closed interval with attained endpoints. Every model in the declared class induces a feasible coupling. Conversely, Proposition 1 realizes every feasible coupling as a shared-exogenous feature SCM, and mixtures realize all intermediate values.

Sharpness is relative to the displayed model. If a fixed full-image law, latent DAG, or nonsaturated image relation imposes further restrictions, the feature program can be only outer. This qualification prevents an audit relation from being mistaken for identification of the Pan–Bareinboim pixel counterfactual.

If the external science supports only a violation-mass budget τ\tau, replace hard support by

Πℛ(τ)​(μ0,μ1)={π∈Π⁡(μ0,μ1):π⁡(ℛc)≤τ}.\Pi_{\mathcal{R}}^{(\tau)}(\mu_{0},\mu_{1})=\{\pi\in\Pi(\mu_{0},\mu_{1}):\pi(\mathcal{R}^{c})\leq\tau\}. (17)

The same compactness argument makes its optimized interval sharp in the corresponding budget model and nested in τ\tau. The smallest feasible budget is the Hall deficiency

δℛ​(μ0,μ1)=maxA⊆𝒴⁡{μ0​(A)−μ1​(ℛ⁡(A))},\delta_{\mathcal{R}}(\mu_{0},\mu_{1})=\max_{A\subseteq\mathcal{Y}}\{\mu_{0}(A)-\mu_{1}(\mathcal{R}(A))\},

by max-flow/min-cut and the finite Hall–Strassen criterion [25]. These established transport results, their duals, and the Polish-space extension are collected in Supplements S2–S3. The parameter τ\tau is sensitivity input, not a probability learned from unpaired images.

For sparse ℛ\mathcal{R}, the two endpoint programs use one variable per allowed edge and 2​|ℛ|2|\mathcal{R}| nonzeros in the marginal equality matrix. The relation is first checked for feasibility; infeasibility is reported rather than hidden by renormalization.

6 Finite-sample outer inference

Let 𝒞α\mathcal{C}_{\alpha} be any random simultaneous confidence region for the two regime marginals satisfying

ℙ{(μ0,μ1)∈𝒞α}≥1−α.\mathbb{P}\{(\mu_{0},\mu_{1})\in\mathcal{C}_{\alpha}\}\geq 1-\alpha. (18)

For a fixed known relation define the union of compatible couplings

𝒫^α​(ℛ)=⋃(ν0,ν1)∈𝒞αΠℛ​(ν0,ν1).\widehat{\mathcal{P}}_{\alpha}(\mathcal{R})=\bigcup_{(\nu_{0},\nu_{1})\in\mathcal{C}_{\alpha}}\Pi_{\mathcal{R}}(\nu_{0},\nu_{1}). (19)
Theorem 7 (Outer coverage of the complete oracle interval).

Let [L⋆,U⋆][L^{\star},U^{\star}] be the sharp interval over Πℛ​(μ0,μ1)\Pi_{\mathcal{R}}(\mu_{0},\mu_{1}). If the random endpoints are measurable, then

ℙ{[L⋆,U⋆]⊆[infπ∈𝒫^α​(ℛ)Hh(π),supπ∈𝒫^α​(ℛ)Hh(π)]}≥1−α.\mathbb{P}\left\{[L^{\star},U^{\star}]\subseteq\left[\inf_{\pi\in\widehat{\mathcal{P}}_{\alpha}(\mathcal{R})}H_{h}(\pi),\sup_{\pi\in\widehat{\mathcal{P}}_{\alpha}(\mathcal{R})}H_{h}(\pi)\right]\right\}\geq 1-\alpha. (20)

If the random feasible set is empty, reporting the vacuous payoff range preserves the guarantee.

Proof.

On the event in (18), every oracle-feasible coupling belongs to (19). Minimizing over the larger set cannot increase the lower endpoint, and maximizing cannot decrease the upper endpoint.

A distribution-free concrete choice uses independent within-regime samples. If Nx​(y)N_{x}(y) is the count in cell yy among nxn_{x} observations, construct a two-sided Clopper–Pearson interval for each of the M=2​KM=2K marginal cells at cellwise noncoverage α/M\alpha/M. In beta-quantile notation its endpoints are

ℓx​(y)\displaystyle\ell_{x}(y) ={0,Nx​(y)=0,Beta−1⁡(α2​M,Nx​(y),nx−Nx​(y)+1),Nx​(y)>0,\displaystyle=\begin{cases}0,&N_{x}(y)=0,\\ \operatorname{Beta}^{-1}\!\left(\frac{\alpha}{2M};N_{x}(y),n_{x}-N_{x}(y)+1\right),&N_{x}(y)>0,\end{cases} (21)
ux​(y)\displaystyle u_{x}(y) ={1,Nx​(y)=nx,Beta−1⁡(1−α2​M,Nx​(y)+1,nx−Nx​(y)),Nx​(y)<nx.\displaystyle=\begin{cases}1,&N_{x}(y)=n_{x},\\ \operatorname{Beta}^{-1}\!\left(1-\frac{\alpha}{2M};N_{x}(y)+1,n_{x}-N_{x}(y)\right),&N_{x}(y)<n_{x}.\end{cases} (22)

Each count is marginally binomial, so Bonferroni gives simultaneous coverage despite dependence among cells within a multinomial sample [5]. Here, exact means finite-sample coverage of at least the nominal level, not equality or shortest possible width. Hoeffding bands provide a simpler alternative. The inference target is the complete oracle interval, not one selected coupling.

If a random outer relation satisfies ℙ⁡(ℛ⋆⊆ℛ^+)≥1−β\mathbb{P}(\mathcal{R}_{\star}\subseteq\widehat{\mathcal{R}}^{+})\geq 1-\beta, the same containment argument and a union bound give coverage at least 1−α−β1-\alpha-\beta when the program uses ℛ^+\widehat{\mathcal{R}}^{+}. Sample splitting alone does not establish this outer-relation property. Compatibility tests and independent finite-panel variants, with their required sampling assumptions, are in Supplement S4.

7 Audit and optimization pipeline

The operational procedure keeps its information sources separate:

  1. 1.

    Prespecify Ψ\Psi, protected roles, allowed descendants, the witness atlas, tolerance, and any role or relation-violation budget.

  2. 2.

    For a candidate pair compute ρ\rho or ρs\rho_{s} and return a short blocking-role certificate and exact repair count when the audit fails.

  3. 3.

    Project the externally justified audit to a feature relation and state explicitly whether that projection is exact or outer.

  4. 4.

    Estimate the two single-world marginals and solve the lower and upper support or budget transport programs, using simultaneous marginal bands when finite-sample coverage is required.

  5. 5.

    Return the identified interval, outer confidence interval, and incompatibility explanation. Selecting a point inside the interval requires a separate declared decision rule.

For nonconvex neural witness spaces not represented by a verified finite atlas, a verified global optimizer would be needed; the certificates above do not validate an arbitrary local neural search.

8 Numerical studies

The numerical studies address three questions that correspond directly to the theory: whether regional plausibility can coexist with global incompatibility, whether a declared relation can sharpen feature-level bounds without concealing infeasibility or misspecification, and whether the resulting optimization and certificate computations remain tractable in structured large instances. The studies are not presented as evidence that an unrestricted pixel-level counterfactual is identified.

Status of the numerical evidence

The accompanying public repository contains the implementation, tests, configurations, retained outputs, and integrity manifests supporting the numerical studies. External datasets are not redistributed and must be obtained from their original sources. The repository documents the scope of the retained evidence and the limitations of exact historical and cross-platform replay.

8.1 Common-witness audits

Controlled rotation audit.

We first use the official MNIST training and test partitions [16] in a controlled renderer experiment. For a source image SS, one angle

A∈{−20,−10,0,10,20}​ degreesA\in\{-20,-10,0,10,20\}\text{ degrees}

is applied to the whole image, followed by clipped independent Gaussian measurement noise with standard deviation 0.020.02. The witness atlas is the same declared set of five renderer responses. For rectangular region PrP_{r}, the discrepancy of angle aa is

cr​(a)=[|Pr|−1​∑p∈Pr{I1​(p)−Rotate⁡(I0,a)​(p)}2]1/2.c_{r}(a)=\left[|P_{r}|^{-1}\sum_{p\in P_{r}}\{I_{1}(p)-\operatorname{Rotate}(I_{0},a)(p)\}^{2}\right]^{1/2}.

The local and common-witness scores are

L=maxr⁡mina​cr​(a),ρ=mina⁡maxr​cr​(a).L=\max_{r}\min_{a}c_{r}(a),\qquad\rho=\min_{a}\max_{r}c_{r}(a).

Thus LL permits a different angle in each region, whereas ρ\rho requires one angle to explain every region. A split of 3,0003{,}000 training images sets the 0.9750.975 order-statistic threshold and a disjoint 3,0003{,}000-image split assesses relation violations. The archived protocol evaluates the official 10,00010{,}000 test images at three seeds and four fixed partitions (4,9,16,4,9,16, and 6464 regions). These are twelve configurations, not 120,000120{,}000 independent test units: the same 10,00010{,}000 images recur. The principal negative control selects angles separately by region, so it is constructed to satisfy the local quantifier while violating the common-angle quantifier. Because every split uses the same specified renderer, this is held-out validation within a controlled mechanism, not independent scientific validation of a causal relation.

Separately fitted editor and witness.

The second study uses paired responses supplied by Morpho-MNIST [4] and smallNORB [17]. Morpho-MNIST provides index-matched plain, thin, and thick digits; these are benchmark-generated transformations rather than physical interventions. smallNORB provides physical toy objects photographed under factorially varied pose and illumination. Lighting 00 is paired with lightings 1,3,1,3, and 55, holding object, pose, and camera fixed. These are matched photographs, not observations of the same stochastic unit in two counterfactual worlds.

An action-conditional latent-residual editor is fitted on units disjoint from a separate PCA–ridge witness. The editor is an experimental vehicle rather than a claimed architectural contribution. The witness divides each image into a fixed 4×44\times 4 grid and normalizes each regional discrepancy by an action-specific held-out threshold qaq_{a}. For normalized costs c~r​(a)=cr​(a)/qa\widetilde{c}_{r}(a)=c_{r}(a)/q_{a}, the relevant scores are

L=maxr⁡mina​c~r​(a),ρfree=mina⁡maxr​c~r​(a),ρreq=maxr⁡c~r​(areq).L=\max_{r}\min_{a}\widetilde{c}_{r}(a),\qquad\rho_{\mathrm{free}}=\min_{a}\max_{r}\widetilde{c}_{r}(a),\qquad\rho_{\mathrm{req}}=\max_{r}\widetilde{c}_{r}(a_{\mathrm{req}}).

A strictly positive qaq_{a} is required; the protocol uses a positive calibration quantile, and any zero quantile would require a positive floor fixed before evaluation. A score at most one is accepted. The free score asks whether some declared action explains the complete image; only the requested-action score tests compliance with the requested edit. The negative control alternates paired actions across the sixteen regions. It can therefore be locally plausible even though no single action explains the image.

The relation panel contains 4,0004{,}000 Morpho-MNIST base digits and ten smallNORB physical objects. The editor evaluation uses a disjoint 6,0006{,}000 Morpho-MNIST base digits and fifteen smallNORB physical objects, with three prespecified editor seeds. Repeated actions and views are dependent measurements of the same base digit or object and are not counted as new independent units. In particular, the smallNORB threshold was calibrated from only five physical objects and is an empirical rule, not a distribution-free conformal guarantee.

Study Evaluation unit Coherent Patch local Patch global AUROC
MNIST rotations 10,00010{,}000/setting 0.97090.9709–0.97770.9777 0.97020.9702–0.98040.9804 00–0.00180.0018 0.99880.9988–1.00001.0000
Morpho-MNIST 4,0004{,}000 digits 0.9777500.977750 0.9767500.976750 0.1137500.113750 0.9872130.987213
smallNORB 1010 objects 0.9592590.959259 0.9728400.972840 00 1.0000001.000000
Table 1: Archived aggregate common-witness results. Coherent is the free-angle global acceptance for controlled MNIST and requested-action acceptance for the paired studies; patch global uses the free common-action score. MNIST entries are ranges over twelve seed–partition configurations that reuse the same test images. The smallNORB AUROC of 11 is complete separation on a fixed ten-object panel, not a population-perfect guarantee.

Table 1 shows the intended local-to-global separation in all three studies. In the controlled renderer, patchworks retain approximately the same local acceptance as coherent pairs but almost never admit one global angle. The learned-witness study shows the same qualitative separation, although the Morpho-MNIST free-action global acceptance of 0.1137500.113750 is materially above zero. On the disjoint editor sets, the conditional editor’s archived mean whole-image SSIM is 0.9031310.903131 versus 0.8705580.870558 for conditional ridge on Morpho-MNIST, and 0.9860640.986064 versus 0.9858210.985821 on smallNORB. We report the corresponding improvements only to appropriate precision, 0.03260.0326 and 0.000240.00024: the latter is practically very small.

The smallNORB aggregate also conceals important object-level uncertainty. Requested-action acceptance ranges from 0.7798350.779835 to 11 across ten objects. Zero global patchwork acceptances among ten objects has one-sided 95%95\% Clopper–Pearson upper limit 0.25890.2589. Consequently, the reported AUROC 11 supports finite-panel separation of the constructed control but does not establish population-perfect detection or a population patchwork acceptance below 0.200.20.

8.2 Sharp bounds and finite-sample outer inference

A three-state synthetic feature provides a direct numerical check of the support-constrained coupling program. With only the two regime marginals, the sharp target interval is [0.35,0.80][0.35,0.80]. Imposing the declared one-step relation narrows it to [0.55,0.70][0.55,0.70]. A misspecification control places 0.200.20 of the true coupling mass outside that relation, and the constrained interval then fails to cover the true target. The negative control is essential: narrowing is created by the added relation, not by the observed marginals alone.

For finite-sample inference, independent samples from each regime are drawn at

n∈{100,300,1000,3000,10000,30000},n\in\{100,300,1000,3000,10000,30000\},

with 300300 repetitions per sample size. The archived Hoeffding outer intervals cover the complete oracle identified interval in all 1,8001{,}800 repetitions, with mean width decreasing from 0.92370.9237 to 0.20750.2075. A later analysis applies the exact Bonferroni–Clopper–Pearson construction in (21)–(22) to the same archived cell counts. Its reported mean widths decrease from 0.74130.7413 to 0.18610.1861, reductions of 10.3%10.3\%–27.8%27.8\% relative to Hoeffding, and all 1,8001{,}800 intervals again cover. These successes are an implementation check under the stated simulation law, not evidence of exact nominal calibration; the coverage guarantee follows from the theorem. The exact-band comparison is post-confirmatory and descriptive because it was specified after the Hoeffding outcomes had been inspected.

Refer to caption
Figure 1: Archived synthetic summaries. Left: all pairs can be compatible while the complete witness intersection is empty. Middle and right: observed coverage of the complete oracle interval and contraction of the Hoeffding outer interval as the regime sample sizes increase.

The controlled audit-to-bound panel also records a necessary negative result. It fixes ten images per digit, for N=100N=100 units, and the raw witness relation contains 0.00970.0097 of the N2N^{2} candidate pairs. The hard-support transport program is infeasible and is reported as such; no renormalization or silent relaxation is used. With the predeclared empirical violation budget τ=0.08\tau=0.08, the archived interval is [0.1517,0.2600][0.1517,0.2600], compared with [0.0600,0.8200][0.0600,0.8200] without the image relation, and contains the hidden paired target 0.21000.2100. However, balancing the panel by digit violates the i.i.d. premise of the intended pooled binomial guarantee. A post-hoc stratified sensitivity calculation uses τ=0.12\tau=0.12 and gives [0.1183,0.2975][0.1183,0.2975], also containing 0.21000.2100. Neither relaxed result is a confirmatory 95%95\% confidence statement; a new independent panel with the stratified procedure fixed in advance would be required.

8.3 Structured computational checks

The archived computations use sparse edge variables for banded transport and short dual certificates for repeated-simplex minimax instances. These tests check the implementations against known optima and residual conditions; they do not claim comparable scaling for dense arbitrary relations, face enumeration, or nonconvex neural witness optimization.

Structured problem Largest instance Representation Time (s)
Sparse coupling 10,00010{,}000 states 109,970109{,}970 edges 18.52118.521
Affine minimax, d=4d=4 100,000100{,}000 roles 55-role certificate 0.2510.251
Affine minimax, d=8d=8 100,000100{,}000 roles 99-role certificate 0.3850.385
Affine minimax, d=16d=16 100,000100{,}000 roles 1717-role certificate 0.7860.786
Helly-tight 1,0241{,}024 roles dimension 1,0231{,}023 9.7049.704
Finite-atlas sharpness 1212 roles 4,0944{,}094 proper faces 0.1920.192
Table 2: Largest successful cases in the archived hardware-specific scaling summary. Relations and affine certificate families are supplied rather than learned. Times are descriptive and hardware-specific.

All 3535 archived scaling cases are reported as successful in 34.8534.85 seconds total with peak resident memory 0.4830.483 GiB. At 10,00010{,}000 states, sparse transport replaces 10810^{8} dense state pairs by 109,970109{,}970 edge variables; the reported marginal residuals are below 1.4×10−181.4\times 10^{-18}. Across the affine cases, the largest reported primal, stationarity, or duality error is 3.34×10−163.34\times 10^{-16}. Constructed leave-one-out atlases with m=4,…,12m=4,\ldots,12 accept every nonempty proper role set and reject the full set, attaining the finite-atlas certificate bound. These are numerical verification and structured-scaling results, not an empirical claim about naturally occurring high-order image obstructions.

8.4 Empirical scope

The experiments validate the declared computations and local-to-global failure mode on controlled or finite paired response families. They do not establish the scientific correctness of an arbitrary relation, identify an unrestricted counterfactual image, or infer a joint multi-regime image law. Morpho-MNIST is synthetic, and smallNORB has only ten physical objects in the relation panel. The reported large-state computations rely on supplied sparse or repeated structure.

9 Scope, information boundary, and limitations

We first clarify the scope of the results. Unrestricted pixel-level counterfactuals are generally not identified from unpaired regime marginals. Accordingly, the framework targets sharp bounds for prespecified finite-dimensional image features under a declared support restriction. This is the identified object of the analysis rather than an approximation to an otherwise identified pixel-level counterfactual.

The witness atlas and the cross-world relation are additional scientific inputs, not consequences of the observed regime marginals. The witness atlas determines the coherence question being audited, whereas the relation determines the admissible feature couplings used for partial identification. If either is misspecified, the audit may reject coherent pairs or the identified interval may exclude the true target. Sample splitting can assess empirical performance but cannot by itself establish the causal validity of these inputs. Likewise, violation budgets are sensitivity parameters unless they are independently calibrated.

A further information boundary arises from feature projection. Projecting an image-level relation onto a coarse feature space can discard image-level restrictions. The feature-level and image-level analyses coincide only under the feature-saturation or conditional-fiber conditions stated in the Supplementary Material.

The current experiments provide controlled evaluations on MNIST, Morpho-MNIST, and smallNORB, together with synthetic and computational studies. They validate the predicted local–global separation, sharp-bound calculations, finite-sample coverage, and structured computational scaling in these settings. Evaluation on broader natural-image domains and empirical validation of the multi-regime extension are left for future work.

10 Conclusion

Local regional plausibility need not imply one globally coherent counterfactual explanation. The common-witness grade records this distinction as a feasibility complex. Finite and action-stratified witness structures yield short exact certificates and repair counts; an externally justified feature relation then narrows, but does not select within, the coupling of the regime marginals. Optimizing over all compatible couplings gives sharp feature bounds in the stated support-only model, and simultaneous marginal confidence regions give finite-sample outer coverage of the complete oracle interval. The resulting pipeline is mathematically explicit about where information enters and where ambiguity remains. Its validity in a new application rests on prespecification, external validation of the witness relation, and reproducible evidence at the correct independent unit.

Code and data availability

The implementation code, experiment runners, tests, protocols, retained aggregate outputs, and reproduction instructions are available here. External datasets are not redistributed and must be obtained from the original sources cited in the Supplementary Material. Documented reproduction limitations are provided in the repository.

References

  • [1] E. Amaldi, M. E. Pfetsch, and Jr. Trotter (2003) On the maximum feasible subsystem problem, IISs and IIS-hypergraphs. Mathematical Programming 95 (3), pp. 533–554. External Links: Document Cited by: §4.1.
  • [2] N. Amenta (1996) A short proof of an interesting helly-type theorem. Discrete & Computational Geometry 15 (4), pp. 423–427. External Links: Document Cited by: item 1, §2, §4.2.
  • [3] A. Balke and J. Pearl (1997) Bounds on treatment effects from studies with imperfect compliance. Journal of the American Statistical Association 92 (439), pp. 1171–1176. External Links: Document Cited by: §2.
  • [4] D. C. Castro, J. Tan, B. Kainz, E. Konukoglu, and B. Glocker (2019) Morpho-MNIST: quantitative assessment and diagnostics for representation learning. Journal of Machine Learning Research 20 (178), pp. 1–29. External Links: Link Cited by: §8.1.
  • [5] C. J. Clopper and E. S. Pearson (1934) The use of confidence or fiducial limits illustrated in the case of the binomial. Biometrika 26 (4), pp. 404–413. External Links: Document Cited by: §2, §6.
  • [6] L. Danzer, B. Grünbaum, and V. Klee (1963) Helly’s theorem and its relatives. In Convexity, Proceedings of Symposia in Pure Mathematics, Vol. 7, pp. 101–180. Cited by: item 1, §2.
  • [7] L. De Lara, A. González-Sanz, N. Asher, L. Risser, and J. Loubes (2024) Transport-based counterfactual models. Journal of Machine Learning Research 25 (136), pp. 1–59. External Links: Link Cited by: §2.
  • [8] J. Eckhoff and K. Nischke (2009) Morris’s pigeonhole principle and the helly theorem for unions of convex sets. Bulletin of the London Mathematical Society 41 (4), pp. 577–588. External Links: Document Cited by: item 1, §2, §4.2.
  • [9] J. Eckhoff (1985) An upper-bound theorem for families of convex sets. Geometriae Dedicata 19 (2), pp. 217–227. External Links: Document Cited by: §4.2.
  • [10] Y. Fan, E. Guerre, and D. Zhu (2017) Partial identification of functionals of the joint distribution of potential outcomes. Journal of Econometrics 197 (1), pp. 42–59. External Links: Document Cited by: §2.
  • [11] Y. Fan and S. S. Park (2010) Sharp bounds on the distribution of treatment effects and their statistical inference. Econometric Theory 26 (3), pp. 931–951. External Links: Document Cited by: §2.
  • [12] L. A. Goodman (1965) On simultaneous confidence intervals for multinomial proportions. Technometrics 7 (2), pp. 247–254. External Links: Document Cited by: §2.
  • [13] G. Kalai (1984) Intersection patterns of convex sets. Israel Journal of Mathematics 48 (2–3), pp. 161–174. External Links: Document Cited by: §2, §4.2.
  • [14] M. Katchalski and A. Liu (1979) A problem of geometry in RnR^{n}. Proceedings of the American Mathematical Society 75 (2), pp. 284–288. External Links: Document Cited by: §2, §4.2.
  • [15] M. Kim and A. Lew (2023) Leray numbers of tolerance complexes. Combinatorica 43, pp. 985–1006. External Links: Document Cited by: item 1, §2, §4.1.
  • [16] Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner (1998) Gradient-based learning applied to document recognition. Proceedings of the IEEE 86 (11), pp. 2278–2324. External Links: Document Cited by: §8.1.
  • [17] Y. LeCun, F. J. Huang, and L. Bottou (2004) Learning methods for generic object recognition with invariance to pose and lighting. In Proceedings of the 2004 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Vol. 2, pp. 97–104. External Links: Document Cited by: §8.1.
  • [18] C. F. Manski (1990) Nonparametric bounds on treatment effects. American Economic Review 80 (2), pp. 319–323. Cited by: §2.
  • [19] W. L. May and W. D. Johnson (2000) Constructing two-sided simultaneous confidence intervals for multinomial proportions for small counts in a large number of cells. Journal of Statistical Software 5 (6), pp. 1–24. External Links: Document Cited by: §2.
  • [20] L. Montejano and D. Oliveros (2011) Tolerance in helly-type theorems. Discrete & Computational Geometry 45 (2), pp. 348–357. External Links: Document Cited by: item 1, §2, §4.1.
  • [21] Y. Pan and E. Bareinboim (2024) Counterfactual image editing. In Proceedings of the 41st International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 235, pp. 39087–39101. External Links: Link Cited by: §1, §2.
  • [22] Y. Pan and E. Bareinboim (2025) Counterfactual image editing with disentangled causal latent space. In Advances in Neural Information Processing Systems, Vol. 38. External Links: Link Cited by: §2.
  • [23] S. T. Rachev and L. Rüschendorf (1998) Mass transportation problems, volume i: theory. Springer, New York. External Links: Document Cited by: §2.
  • [24] C. P. Sison and J. Glaz (1995) Simultaneous confidence intervals and sample size determination for multinomial proportions. Journal of the American Statistical Association 90 (429), pp. 366–369. External Links: Document Cited by: §2.
  • [25] V. Strassen (1965) The existence of probability measures with given marginals. Annals of Mathematical Statistics 36 (2), pp. 423–439. External Links: Document Cited by: §2, §5.
  • [26] J. Tian and J. Pearl (2000) Probabilities of causation: bounds and identification. Annals of Mathematics and Artificial Intelligence 28, pp. 287–313. External Links: Document Cited by: §2.
  • [27] J. Zhang, J. Tian, and E. Bareinboim (2022) Partial counterfactual identification from observational and experimental data. In Proceedings of the 39th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 162, pp. 26548–26558. External Links: Link Cited by: §2.