Parlay Identification of Ising Couplings
Abstract
A contract that pays only when two events both occur is called a parlay. Its price supplies joint information absent from the two individual prices. We ask what this information identifies, and how repeated resolved outcomes quantify its uncertainty. An Ising model describes joint outcomes through individual-event biases and pair interactions. For two events, positive probabilities for all four outcomes determine its parameters uniquely. In a larger network, the same calculation measures effective pair dependence, which can include interactions through other events. Selected pair observations cannot establish a missing direct interaction without additional information. We give recovery formulas under full observations or justified support restrictions, and a test for global quote compatibility. For repeated aligned outcomes, we derive the maximum-likelihood estimator and confidence sets that allow zero counts. Further bounds cover quote error, changes in the outcome law, and certified numerical evaluation. Synthetic examples distinguish inference accuracy from market-maker loss. Prices describe a pricing law. Interpreting them as physical probabilities requires separate calibration evidence.
The author has a commercial interest in systems of the kind this paper describes.
1 Introduction
1.1 What joint prices add
Two individual event prices leave their joint probability undetermined. Consider contracts paying \$1 when Team A or Team B wins its next game, and zero otherwise. The games may share conditions that affect both results. Suppose their prices are p_i=.65 and p_j=.40, as in Example 3.3. Under independence, a contract paying only when both teams win has price p_ip_j=.26. Such a joint contract is a parlay; its component events are its legs. A price of .31 or .21 is also compatible with those individual prices. The first puts more probability on joint success than independence does, and the second puts less.
Write \pi_{ij} for the parlay price. The four outcome probabilities are then \pi_{ij}, p_i-\pi_{ij}, p_j-\pi_{ij}, and 1-p_i-p_j+\pi_{ij}. They correspond respectively to both wins, only Team A winning, only Team B winning, and neither winning. Their nonnegativity restricts the parlay price to the interval [.05,.40] in this example. Thus one joint price completes the pair table, while either individual price alone leaves the dependence unresolved.
These probabilities first describe prices. Write Q for a coherent pricing law: one probability distribution that reproduces all the quoted contract prices. Write P for the physical outcome law, the distribution that generates resolved events. Encode a win by \omega_i=+1 and a loss by \omega_i=-1. The quoted marginal prices are p_i=Q(\omega_i=+1) and p_j=Q(\omega_j=+1). Market scoring rules elicit beliefs under their preference assumptions [2]. Equality Q=P requires a separate calibration assumption and evidence from resolved outcomes [26]. Each result applies to its named law. The symbol \mathbb{P} denotes that specified law, using Q for quote inversion and P for outcome estimation.
Resolved outcomes supply another route to the same four-cell table. An aligned round records both results from one specified occurrence of the pair, such as the two games in one week. Repeated rounds estimate how often each joint result occurs. Both legs can resolve without a traded parlay, and every round supplies one observation regardless of ticket count. The fixed-law estimates below require independent rounds with a common pair law. Trade signals and settled observations therefore have different sampling requirements.
The distinction between pair dependence and network interaction is equally necessary. Two results can depend on a shared third event even when their model has no direct interaction. The pair table measures their resulting association. Recovering the direct interaction requires additional observations or an independently justified restriction on the network. The paper makes these requirements explicit before drawing any network conclusion.
ParlayMarket [3] learns a shared Ising pricing law from trade targets. The question here is what specified probabilities and resolved observations determine, and with what uncertainty. The classical two-by-two table already supplies the exact pair inverse and its sampling interpretation [29, 30, 4]. Its use within a larger event model requires the compatibility, embedding, and observation distinctions developed below.
1.2 A joint model with pair interactions
For several events, a model must assign consistent probabilities to all joint outcomes. Rolling correlations, factor models, copulas, and Bayesian networks provide different descriptions of dependence. Here we use a model whose log probability contains individual-event terms and pair-interaction terms. For M binary events e_1, \ldots, e_M with outcomes \omega_i \in \{-1, +1\}, it has joint distribution \mathbb{P}(\omega_1, \ldots, \omega_M) \;=\; \frac{1}{Z} \exp\!\left(\sum_i h_i \omega_i + \sum_{i<j} J_{ij}\omega_i\omega_j\right), \tag{1} Here Z normalizes the probabilities to sum to one. The field h_i sets an event’s conditional bias, and the coupling J_{ij} measures its interaction with another event. An event index is also called a site, and its signed outcome a spin. This is the Ising model [5], also expressed by Besag’s auto-logistic model [22] and the Boltzmann machine [23]. It is an exponential family: its log probabilities are linear in fixed outcome statistics, apart from normalization. Three properties explain its use here:
The Ising density (1) is the pairwise maximum-entropy distribution on \{-1, +1\}^M with specified marginals and pairwise correlations; among distributions matching those moments it assumes the least additional structure.
The parameters (h_i, J_{ij}) have direct conditional interpretations: 2h_i + 2\sum_{j \neq i} J_{ij}\omega_j is the conditional log-odds of e_i given the other outcomes, and 4J_{ij} is the conditional log odds ratio of the pair (e_i, e_j) given the rest. Neither is a marginal quantity; Remark 2.2 states why h_i is not the marginal half-logit.
The pair (h, J) determines a consistent joint distribution over all 2^M outcomes, enabling the pricing of parlays and other joint-outcome contracts.
The pairwise maximum-entropy form omits higher-order log-linear interactions. It cannot represent a three-event law whose genuine joint interaction remains invisible in every pair table. Its two-event family nevertheless represents every positive pair law. Copulas, vine copulas, factor models, and directed graphical models impose different dependence assumptions. We use identifiability to mean that the stated observations determine one parameter value within the stated model. The results establish that property and quantify estimation error under the Ising parameterisation.
1.3 From pair probabilities to inference
The pair inverse starts with four probabilities and extracts their signed logarithmic contrasts. Its network interpretation then depends on what was observed beyond the pair. Sampling uncertainty enters when the probabilities themselves are estimated from resolved rounds. The following results address these successive questions.
- (1)
-
Identifiability (Section 3). For M = 2, the triple (p_i, p_j, \pi_{ij}) of two marginals and the joint-outcome probability determines (h_i, h_j, J_{ij}) uniquely through an explicit closed form, and the admissible region of parlay prices is exactly the interior of the Fréchet-Hoeffding region (Theorem 3.1).
- (2)
-
Embedding bias (Section 4). Applied to a pair inside a network of M > 2 coupled events, the closed form reads the direct coupling plus a two-hop correction through shared neighbours and a second correction carried by the correlations among those neighbours. Theorem 4.1 gives both terms and an unconditional fourth-order remainder bound, and an exactly solvable three-site chain confirms the two-hop constant and the remainder’s order.
- (3)
-
Estimation (Section 6). The two-site model is saturated: the closed form on empirical joint-outcome counts is the maximum-likelihood estimator. Theorem 6.3 computes its Fisher information and identifies its efficient variance as Woolf’s log-odds-ratio variance; Theorem 6.4 gives a finite-sample concentration bound whose constant is computed from the two-site moments.
- (4)
-
Robustness (Section 8). Theorem 8.1 propagates an L^\infty perturbation of the quoted triple to the inferred parameters: noise of size \eta moves each of the three inferred parameters by at most 2\eta/\kappa, where \kappa is the minimum of all four cell probabilities throughout the admitted uncertainty region.
Section 3.2 tests whether all quoted moments admit one joint law. Section 5 gives constructive direct recovery and proves why selected moments cannot certify a missing edge. Section 6.3 constructs count-based uncertainty sets, including zero-count samples. Section 7 separates sampling error from changes in the outcome law. Its certified arithmetic in Section 7.1 preserves interval inclusion through numerical evaluation.
Global compatibility and sufficient-information recovery determine which joint model the observations support. The confidence sets retain this distinction when observations are finite or their law changes. Computation results control approximation within a specified law. Section 10 evaluates these different quantities in synthetic examples.
2 The Ising Model for Binary Event Markets
2.1 Outcomes and conditional interactions
Consider M binary event types e_1, \ldots, e_M with outcomes \omega_i \in \{-1, +1\}; \omega_i = +1 denotes e_i occurring. A settled round is one joint realisation of the type vector: round r resolves (\omega_1^{(r)}, \ldots, \omega_M^{(r)}) as a single draw from the joint distribution below, and the sites index the types, not the individual resolutions. The two markets of Section 1 are types in this sense, with one round for each aligned pair of games. A type that resolves once and never again supplies exactly one round; Section 11, limitation (L9), records that one round supplies no useful precision guarantee for the estimates of Section 6.
The model assigns each outcome a real score before converting scores to probabilities. The negative of that score is called the Hamiltonian. Its sign convention agrees with the statistical-physics literature.
Definition 2.1 (Ising Hamiltonian).
The Ising Hamiltonian for M events is H(\omega) \;=\; -\sum_{i=1}^{M} h_i \omega_i \;-\; \sum_{1 \leq i < j \leq M} J_{ij}\, \omega_i \omega_j, \tag{2} where h_i \in \mathbb{R} is the external field on site i and J_{ij} \in \mathbb{R} is the pairwise interaction; J_{ij} > 0 means the events e_i, e_j tend to co-occur, J_{ij} < 0 means they tend to anti-correlate.
The coupling sign describes conditional association when all other outcomes are fixed. In a larger network, marginal association can also reflect paths through other sites, as Section 4 shows. The joint distribution is the Boltzmann density \mathbb{P}(\omega_1, \ldots, \omega_M) \;=\; \frac{1}{Z} \exp\!\bigl(-H(\omega)\bigr) \;=\; \frac{1}{Z} \exp\!\left(\sum_i h_i \omega_i + \sum_{i<j} J_{ij}\omega_i\omega_j\right), \tag{3} with partition function Z = \sum_{\omega \in \{-1,+1\}^M} \exp(-H(\omega)).
A marginal averages over the other outcomes; a conditional probability holds them fixed. For 0<p<1, the odds are p/(1-p) and their logarithm is \mathop{\mathrm{logit}}(p)=\log(p/(1-p)). The next calculation explains why converting a marginal price to log-odds does not usually recover its field.
Remark 2.2 (Fields are not read off marginals).
Suppose the market on event e_i quotes marginal price p_i, and that price identifies the Ising marginal \mathbb{P}(\omega_i = +1) = p_i. It is tempting to read the field off that marginal as \tfrac{1}{2}\mathop{\mathrm{logit}}(p_i). That is not the field of a coupled model. The field is a conditional quantity: from (3), \mathop{\mathrm{logit}}\mathbb{P}(\omega_i = +1 \mid \omega_{-i}) \;=\; 2h_i \;+\; 2\sum_{j \neq i} J_{ij}\omega_j , \tag{4} the auto-logistic form of Besag [22], so 2h_i is the intercept of e_i’s conditional log-odds and not its marginal log-odds. The two agree only where the conditioning does no work. For M = 2, summing the Boltzmann weights over \omega_j gives \mathop{\mathrm{logit}}(p_i) \;=\; 2h_i \;+\; \ln\frac{\cosh(h_j + J_{ij})}{\cosh(h_j - J_{ij})}, \tag{5} so h_i = \tfrac{1}{2}\mathop{\mathrm{logit}}(p_i) holds exactly when \cosh(h_j + J_{ij}) = \cosh(h_j - J_{ij}), that is, exactly when h_j J_{ij} = 0: the coupling vanishes, or the neighbour’s field does. Outside that set the half-logit is biased, and the bias is not small. At (h_i, h_j, J_{ij}) = (0.4, 0.6, 0.7) it returns 0.737 against a true field of 0.400; at (0.3, -0.5, 0.9) it returns -0.044 and so reports the wrong sign.
The second term of (5) records the neighbour’s contribution to the marginal probability. Recovering the field requires accounting for that contribution. This is why Theorem 3.1 inverts the joint law rather than the two marginals, and why the parlay price appears in all three coordinates of (7) rather than in J_{ij} alone. The independent case is the consistency check: setting \pi_{ij} = p_ip_j in (7) returns J_{ij} = 0 and h_i = \tfrac{1}{2}\mathop{\mathrm{logit}}(p_i), recovering the naive form exactly where it is valid.
The factor of 1/2 throughout arises because the Ising convention uses \omega_i \in \{-1, +1\} rather than \{0, 1\}. Under the \{0, 1\} convention with \sigma_i = (\omega_i + 1)/2 the coupling scales to \tilde J_{ij} = 4 J_{ij} and the field becomes \tilde h_i = 2h_i - 2\sum_{j \neq i} J_{ij}, the shift arising because \omega_i\omega_j = 4\sigma_i\sigma_j - 2\sigma_i - 2\sigma_j + 1. This recovers the logistic parameterisation, in which (4) reads \mathop{\mathrm{logit}}\mathbb{P}(\sigma_i = 1 \mid \sigma_{-i}) = \tilde h_i + \sum_{j \neq i} \tilde J_{ij}\sigma_j. The convention change is algebraic. The probability in this formula belongs to the specified pricing or outcome law. Quote perturbations and their uncertainty propagate through Section 8.
3 Identifiability from Marginal and Parlay Prices
3.1 Closed-form inversion and exact identifiability
A parlay trade on pair (i, j) pays \$1 if \omega_i = +1 and \omega_j = +1 jointly, zero otherwise; an admissible quote supplies the joint pricing probability \pi_{ij} \;=\; \mathbb{P}(\omega_i = +1, \omega_j = +1). The triple (p_i, p_j, \pi_{ij}) determines all four cell probabilities. Taking logarithms converts the Ising weights into a linear system. Signed sums cancel the normalizing constant and isolate the two fields and the coupling. The proof calls these sign patterns the Walsh basis; each basis function is a product of selected spins. Positivity is essential because the inverse takes the logarithm of every cell.
Theorem 3.1 (Exact identifiability of the two-event Ising model).
For marginals p_i, p_j \in (0,1) define the admissible parlay region \mathcal{A}(p_i, p_j) \;:=\; \bigl\{ \pi \in (0, 1) \;:\; \max(0, p_i + p_j - 1) \;<\; \pi \;<\; \min(p_i, p_j) \bigr\}, \tag{6} and call (p_i, p_j, \pi_{ij}) an admissible triple when (p_i, p_j) \in (0,1)^2 and \pi_{ij} \in \mathcal{A}(p_i, p_j). Then:
- (i)
-
The map \Phi : (p_i, p_j, \pi_{ij}) \mapsto (h_i, h_j, J_{ij}) defined on the admissible triples by \begin{aligned} h_i &= \tfrac{1}{4}\ln\!\frac{\pi_{ij}\,(p_i-\pi_{ij})}{(p_j-\pi_{ij})(1-p_i-p_j+\pi_{ij})}, \\ h_j &= \tfrac{1}{4}\ln\!\frac{\pi_{ij}\,(p_j-\pi_{ij})}{(p_i-\pi_{ij})(1-p_i-p_j+\pi_{ij})}, \\ J_{ij} &= \tfrac{1}{4}\ln\!\frac{\pi_{ij}\,(1-p_i-p_j+\pi_{ij})}{(p_i-\pi_{ij})(p_j-\pi_{ij})}, \end{aligned} \tag{7} is a bijection from the set of admissible triples onto \mathbb{R}^3. The two-event Ising parameters (h_i, h_j, J_{ij}) are therefore uniquely determined by the triple (p_i, p_j, \pi_{ij}) inside the admissible region.
Every one of the three coordinates depends on the parlay price \pi_{ij}, the fields included. A field is not recoverable from its own marginal alone: inverting a coupled model requires the joint observation. Remark 2.2 states the exact obstruction and the special case in which the marginal half-logit is correct.
- (ii)
-
\mathcal{A}(p_i, p_j) is exactly the set of parlay prices the two-event Ising family realises: every admissible triple corresponds to a well-defined two-event Ising distribution, and no triple outside this region does.
- (iii)
-
At fixed marginals (p_i, p_j), the Fréchet-Hoeffding bounds \max(0, p_i + p_j - 1) \leq \pi \leq \min(p_i, p_j) are attained in the limit J_{ij} \to \pm\infty; the strict inequalities in \mathcal{A} correspond to finite J_{ij}.
- (iv)
- \pi_{ij} = p_i p_j iff J_{ij} = 0 (independence). \pi_{ij} > p_i p_j iff J_{ij} > 0; \pi_{ij} < p_i p_j iff J_{ij} < 0.
Proof. The closed form. The four-state Boltzmann distribution at M = 2 gives \begin{align*} \mathbb{P}(+1, +1) &= Z^{-1}\exp(h_i + h_j + J_{ij}), & \mathbb{P}(+1, -1) &= Z^{-1}\exp(h_i - h_j - J_{ij}),\\ \mathbb{P}(-1, +1) &= Z^{-1}\exp(-h_i + h_j - J_{ij}), & \mathbb{P}(-1, -1) &= Z^{-1}\exp(-h_i - h_j + J_{ij}), \end{align*} so the cross-ratio satisfies \frac{\mathbb{P}(+1,+1)\,\mathbb{P}(-1,-1)}{\mathbb{P}(+1,-1)\,\mathbb{P}(-1,+1)} \;=\; \exp(4 J_{ij}). Substituting \mathbb{P}(+1,+1) = \pi_{ij}, \mathbb{P}(+1, -1) = p_i - \pi_{ij}, \mathbb{P}(-1, +1) = p_j - \pi_{ij}, \mathbb{P}(-1, -1) = 1 - p_i - p_j + \pi_{ij} gives the J_{ij} formula in (7).
(i) The Walsh basis. Write the four cell probabilities as \mathbb{P}(+1,+1) = \pi_{ij}, \mathbb{P}(+1,-1) = p_i - \pi_{ij}, \mathbb{P}(-1,+1) = p_j - \pi_{ij}, \mathbb{P}(-1,-1) = 1 - p_i - p_j + \pi_{ij}, each strictly positive on \mathcal{A}; the map from admissible triples to cell vectors is an affine bijection onto the interior of the probability simplex on the four sign patterns. Taking logarithms of the Boltzmann weights gives the linear system \ln\mathbb{P}(\omega_i,\omega_j) \;=\; h_i\omega_i + h_j\omega_j + J_{ij}\omega_i\omega_j - \ln Z , whose coefficient matrix over the four sign patterns has columns 1, \omega_i, \omega_j, \omega_i\omega_j: the Walsh functions on \{-1,+1\}^2, orthogonal, forming a 4 \times 4 Hadamard matrix. Given a positive cell vector, the signed quarter-contrasts of its logarithms therefore return a unique (h_i, h_j, J_{ij}), and they are exactly the three expressions of (7). Conversely, given (h_i, h_j, J_{ij}) \in \mathbb{R}^3, the normalised Boltzmann weights form a positive cell vector whose quarter-contrasts return the same parameters, because the contrasts annihilate the constant Walsh direction carried by \ln Z; and any positive cell vector with the same contrasts has a logarithm differing by a multiple of the constant direction, so normalisation makes the two vectors equal. The two maps are mutually inverse bijections between the interior of the simplex and \mathbb{R}^3, and composing with the affine bijection above makes \Phi a bijection onto \mathbb{R}^3. Each parameter is determined by all three observations jointly, none by its own marginal alone: the second contrast recovers h_i, the third h_j, the fourth J_{ij}.
Monotonicity. Given (p_i, p_j), the map \pi \mapsto J via (7) has derivative \frac{dJ}{d\pi} \;=\; \frac{1}{4} \cdot \frac{d}{d\pi}\ln\!\frac{\pi(1 - p_i - p_j + \pi)}{(p_i - \pi)(p_j - \pi)} \;=\; \frac{1}{4}\!\left[\frac{1}{\pi} + \frac{1}{1 - p_i - p_j + \pi} + \frac{1}{p_i - \pi} + \frac{1}{p_j - \pi}\right], which is strictly positive on \mathcal{A} because each of the four denominators is positive exactly on \mathcal{A}: J is strictly increasing in \pi. At the lower boundary \pi \downarrow \max(0, p_i + p_j - 1), one of \{\pi,\, 1 - p_i - p_j + \pi\} tends to zero and J \to -\infty; at the upper boundary \pi \uparrow \min(p_i, p_j), one of \{p_i - \pi,\, p_j - \pi\} tends to zero and J \to +\infty.
(ii) Immediate from (i): the forward map lands inside the admissible region because the Boltzmann cells are strictly positive, and (i) exhibits its inverse on all of the region. A triple outside the closed Fréchet region has a negative cell and corresponds to no joint distribution at all.
(iii) The boundary limits computed under Monotonicity above.
(iv) At J_{ij} = 0, the Boltzmann factors are \exp(\pm h_i \pm h_j) and the distribution factorises to \mathbb{P}(\omega_i, \omega_j) = \mathbb{P}(\omega_i)\mathbb{P}(\omega_j), so \pi_{ij} = p_i p_j. Strict monotonicity of J(\pi) from (i) gives the signed equivalences. ◻
Applied to a pair (i, j) inside a network of M > 2 coupled events, the closed form (7) evaluated on the true pair marginal returns the effective pair coupling J^{\mathrm{eff}}_{ij} = J^*_{ij} + \Delta_{ij}, with J^* the network interaction matrix and \Delta_{ij} second order in the couplings; Section 4 computes \Delta_{ij}.
Remark 3.2 (Fréchet-Hoeffding bounds).
The admissible region \mathcal{A}(p_i, p_j) is exactly the interior of the Fréchet-Hoeffding compatibility region for two Bernoulli margins: the strict inequalities arise because every finite-parameter Ising law has strictly positive cells. A market that quotes (p_i, p_j, \pi_{ij}) outside the closed Fréchet region is inconsistent with every joint distribution on \{-1, +1\}^2. A boundary table is compatible and lies outside the finite two-site family. With fixed interior singleton marginals, its coupling tends to a signed infinity. Section 6.3 defines output and uncertainty for boundary observations.
Example 3.3 (Numerical illustration).
Suppose p_1 = 0.65, p_2 = 0.40. The admissible parlay region is \mathcal{A} = (0.05, 0.40).
Independence: \pi_{12} = p_1 p_2 = 0.26 gives J_{12} = \tfrac{1}{4}\ln[0.26 \cdot 0.21/(0.39 \cdot 0.14)] = 0 exactly.
Positive coupling: \pi_{12} = 0.31 gives J_{12} = \tfrac{1}{4}\ln[0.31 \cdot 0.26/(0.34 \cdot 0.09)] \approx +0.24.
Negative coupling: \pi_{12} = 0.21 gives J_{12} = \tfrac{1}{4}\ln[0.21 \cdot 0.16/(0.44 \cdot 0.19)] \approx -0.23.
Infeasibility: any quote \pi_{12} < 0.05 lies below the Fréchet lower bound p_1 + p_2 - 1 = 0.05 and is inconsistent with every joint distribution on \{-1, +1\}^2; the boundary quote \pi_{12} = 0.05 is the degenerate law with \mathbb{P}(-1,-1) = 0 of Remark 3.2, outside the Ising family.
3.2 Global compatibility of quoted moments
Several valid pair tables can contradict one another when assigned to the same events. A moment is the expectation of a specified outcome statistic. For spins, singleton and pair moments are respectively averages of x_i and x_ix_j. The pair moment is a raw expectation. Covariance subtracts the product of singleton means, and Pearson correlation further divides by the two standard deviations. Let T(x) collect the available statistics at a full outcome x\in\{-1,+1\}^M. Let m be their quote-implied moment vector, using \mathbb{E}[x_i]=2p_i-1 and \mathbb{E}[x_ix_j]=4\pi_{ij}-2p_i-2p_j+1. The unknowns below are the probabilities of all full outcomes. Searching over these probabilities tests joint compatibility and, when necessary, measures the smallest adjustment to the quoted moments.
Proposition 3.4 (Finite-state feasibility and projection).
The moments m admit a joint law exactly when nonnegative numbers w_x solve \sum_xw_x=1,\qquad\sum_xw_xT(x)=m. For an inconsistent vector, the program \min_{w_x\geq0,\,\sum_xw_x=1}\left\|\sum_xw_xT(x)-m\right\|_\infty returns the smallest maximum coordinate adjustment and an explicit compatible law. An interior feasible moment vector has a finite maximum-entropy parameter. Boundary feasibility can require a law with zero cells.
Proof. The weights are precisely a joint probability table. Their moment image is the convex hull of the feature vectors. The simplex is compact, so the continuous projection objective attains its minimum. The interior exponential-family representation follows from entropy maximization. A boundary moment vector has a supporting hyperplane and forces zero probability outside its equality set. ◻
For three spins with all singleton means zero, set every pair moment to -.9. Each pair has positive cells (.025,.475,.475,.025). But every configuration satisfies x_1x_2+x_1x_3+x_2x_3\geq-1, while the requested sum is -2.7. The exact projection residual in these raw moment coordinates is 17/30 in maximum norm. Every correction must increase the correlation sum by 1.7, so one coordinate changes by at least 17/30. Uniform mass 1/6 on the six configurations other than +++, --- attains the bound, with each correlation -1/3. This projection is a boundary law. Mixing it with uniform mass gives a specified positive approximation and a larger recorded residual.
Exact enumeration makes the finite-state feasibility problem explicit. Floating-point solvers report their tolerance and their probability witness. An outer relaxation can certify inconsistency if even the relaxation is infeasible. Feasibility of a relaxation alone leaves joint compatibility unresolved. A nonnegative normalized probability witness establishes compatibility to its checked residual, regardless of how it was found.
4 Embedding Bias at M > 2
The pair inverse remains exact for a pair (i,j) inside a network of M>2 events. Every positive pair table has a two-site Ising representation. The difference lies in the parameter it recovers: marginalizing the other events can change the pair’s coupling. We call the result the effective coupling and reserve direct interaction for the coefficient in the full network law. Their difference is the embedding bias relative to the direct interaction.
A shared neighbour already produces this difference. In the chain i\!-\!k\!-\!j, the two endpoints can be associated even when J^*_{ij}=0. Example 4.2 computes this effect exactly. The theorem separates the contribution of shared neighbours from correlations between different neighbours, then bounds the remaining terms. The environment consists of the other sites after deleting the target pair; its m_k are mean spins. Its proof sums out the other sites and applies the same four-cell logarithmic contrast as the pair inverse. The logarithm of the environment’s exponential moment organizes this calculation into cumulants, its successive derivatives at zero.
Theorem 4.1 (Two-site embedding bias).
Let J^* \in \mathbb{R}^{M \times M} be the true symmetric interaction matrix with J_{\max} = \max_{k, l} |J^*_{kl}|, let D = \max_i |\{k : J^*_{ik} \neq 0\}| be the maximum degree of the interaction graph, and let D_{\mathrm{sh}} = |\{k \neq i, j : J^*_{ik} J^*_{jk} \neq 0\}| \leq D be the number of neighbours the pair (i, j) shares. Let the environment distribution be the Boltzmann distribution of the sites k \neq i, j with the pair deleted — fields h_k and couplings J^*_{kl} for k, l \neq i, j only — and write m_k and \mathop{\mathrm{Cov}} for its marginals and covariances. Put \chi_{\mathrm{env}} \;:=\; \max_{k \neq i, j} \;\sum_{l \neq i, j, k} \bigl|\mathop{\mathrm{Cov}}(\omega_k, \omega_l)\bigr| , and collect the couplings incident on the pair into the vectors u := (J^*_{ik})_{k \neq i,j} and v := (J^*_{jk})_{k \neq i,j}, so that \|u\|_1, \|v\|_1 \leq D J_{\max}. Suppose the closed-form estimator (7) is applied to the pair (i, j) with \pi_{ij} = \mathbb{P}_{J^*}(\omega_i = +1, \omega_j = +1), the true M-site joint, and let \hat J_{ij}^{(M>2)} denote the result. Then \hat J_{ij}^{(M>2)} \;=\; J^*_{ij} \;+\; \Delta_{ij}, \qquad \Delta_{ij} \;=\; \sum_{k \neq i, j} J^*_{ik} J^*_{jk}\,(1 - m_k^{2}) \;+\; E_{ij} \;+\; R_{ij}, \tag{8} where the two-hop sum and the environment-correlation term E_{ij} \;:=\; \sum_{\substack{k, l \neq i, j\\ k \neq l}} J^*_{ik} J^*_{jl}\, \mathop{\mathrm{Cov}}(\omega_k, \omega_l) \tag{9} together carry the whole second cumulant of the environment, and the two error terms satisfy, with no condition on the couplings, |E_{ij}| \;\leq\; D\, J_{\max}^2\, \chi_{\mathrm{env}}, \qquad |R_{ij}| \;\leq\; \tfrac{7}{48}\bigl(\|u + v\|_1^4 + \|u - v\|_1^4\bigr) \;\leq\; \tfrac{14}{3}\,(D J_{\max})^4 . \tag{10} The remainder collects the fourth and higher even cumulants of the environment: odd cumulant orders cancel identically in the coupling contrast, so no third-order term exists beyond E_{ij}. E_{ij} vanishes when the environment sites are pairwise uncorrelated, in particular for the single-neighbour chain of Example 4.2. Under the influence condition below, E_{ij} is bounded by a third-order term with quadratic degree dependence. Without that condition, the covariance bound in (10) applies directly. The size of \chi_{\mathrm{env}}, the reason the two-hop sum is stated at the environment marginals, and a worked family in which E_{ij} dominates are recorded in Remark 4.5.
Proof. Write the full M-site Hamiltonian as H = H_{ij} + H_{\text{other}} + H_{\text{coupling}} where \begin{align*} H_{ij}(\omega_i, \omega_j) &= -h_i \omega_i - h_j \omega_j - J^*_{ij}\omega_i \omega_j,\\ H_{\text{other}}(\omega_{-ij}) &= -\sum_{k \neq i, j} h_k \omega_k - \sum_{\{k,l\} \cap \{i,j\} = \emptyset} J^*_{kl}\omega_k\omega_l,\\ H_{\text{coupling}}(\omega) &= -\sum_{k \neq i, j} (J^*_{ik}\omega_k) \omega_i - \sum_{k \neq i, j} (J^*_{jk}\omega_k) \omega_j. \end{align*} Marginalising over \omega_{-ij}, the reduced two-site distribution is \mathbb{P}(\omega_i, \omega_j) \;\propto\; \exp\!\bigl(h_i \omega_i + h_j \omega_j + J^*_{ij}\omega_i\omega_j\bigr)\, \exp\!\bigl(K(t(\omega_i, \omega_j))\bigr), \qquad t(\omega_i, \omega_j) \;:=\; u\,\omega_i + v\,\omega_j , where K(t) := \ln \mathbb{E}_{\mathrm{env}}\bigl[\exp\bigl(\sum_{k \neq i,j} t_k \omega_k\bigr)\bigr] is the cumulant generating function of the environment spins, the expectation taken under the distribution \propto \exp(-H_{\text{other}}). The closed form (7) recovers exactly the coupling of the two-site density it is applied to (Theorem 3.1), namely the quarter-contrast of the log-cells, in which the field terms and the normalisation cancel: \begin{align*} \hat J_{ij}^{(M>2)} \;&=\; J^*_{ij} \;+\; \tfrac{1}{4}\bigl[K(u+v) + K(-u-v) - K(u-v) - K(-u+v)\bigr]\\ &=\; J^*_{ij} \;+\; \tfrac{1}{2}\bigl[E(u+v) - E(u-v)\bigr], \end{align*} where E(t) := \tfrac{1}{2}[K(t) + K(-t)] is the even part of K. This identity is exact. Every odd-order cumulant is an odd function of t and is annihilated by the even part, so \Delta_{ij} = \tfrac{1}{2}[E(u+v) - E(u-v)] carries only even cumulant orders; the odd part of K enters only the field contrasts, renormalising h_i to h_i + \sum_k J^*_{ik} m_k at first order.
Second cumulant. With \Sigma := \mathop{\mathrm{Cov}} the environment covariance matrix, the quadratic term of E is \tfrac{1}{2} t^\top \Sigma\, t, and \tfrac{1}{2}\Bigl[\tfrac{1}{2}(u+v)^\top \Sigma\, (u+v) - \tfrac{1}{2}(u-v)^\top \Sigma\, (u-v)\Bigr] \;=\; u^\top \Sigma\, v \;=\; \sum_{k \neq i, j} J^*_{ik} J^*_{jk}\,(1 - m_k^{2}) \;+\; E_{ij}, the diagonal of \Sigma giving the two-hop sum and the off-diagonal giving (9).
Remainder. R_{ij} = \tfrac{1}{2}[E_{\geq 4}(u+v) - E_{\geq 4}(u-v)] with E_{\geq 4}(t) := E(t) - \tfrac{1}{2}t^\top \Sigma\, t. Taylor’s theorem with integral remainder at fourth order gives, for the even part, E_{\geq 4}(t) \;=\; \int_0^1 \frac{(1-s)^3}{6}\; \frac{D^4 K(st) + D^4 K(-st)}{2}\,[t^{\otimes 4}]\; ds , and D^4 K(w)[t^{\otimes 4}] is the fourth cumulant of X_t := \sum_k t_k \omega_k under the exponential tilt of the environment by w. For any random variable Y with |Y| \leq L: \mathop{\mathrm{Var}}(Y) \leq L^2 and \|Y - \mathbb{E}Y\|_\infty \leq 2L, hence |\kappa_4(Y)| \;=\; \bigl|\mathbb{E}(Y - \mathbb{E}Y)^4 - 3 \mathop{\mathrm{Var}}(Y)^2\bigr| \;\leq\; (2L)^2 \mathop{\mathrm{Var}}(Y) + 3 \mathop{\mathrm{Var}}(Y)^2 \;\leq\; 7 L^4 . With L = \|t\|_1, an upper bound valid under every tilt, and \int_0^1 (1-s)^3/6\; ds = 1/24, this gives |E_{\geq 4}(t)| \leq \tfrac{7}{24}\,\|t\|_1^4, hence the first bound of (10); \|u \pm v\|_1 \leq \|u\|_1 + \|v\|_1 \leq 2 D J_{\max} gives the second.
The E_{ij} bound. At most D indices k carry J^*_{ik} \neq 0, so |E_{ij}| \;\leq\; \sum_{k \,:\, J^*_{ik} \neq 0} |J^*_{ik}|\; J_{\max} \sum_{l \neq i, j, k} \bigl|\mathop{\mathrm{Cov}}(\omega_k, \omega_l)\bigr| \;\leq\; D\, J_{\max}^2\, \chi_{\mathrm{env}} . \quad\square ◻
Example 4.2 (Exact check on the three-site chain).
Take the chain i\!-\!k\!-\!j with J^*_{ij} = 0, J^*_{ik} = J^*_{jk} = J, and h = 0. Summing over \omega_k gives \mathbb{P}(\omega_i, \omega_j) \propto \cosh\bigl(J(\omega_i + \omega_j)\bigr), so the estimator reads \hat J_{ij}^{(M>2)} \;=\; \frac{1}{4}\ln\frac{\cosh^2(2J)}{\cosh^2(0)} \;=\; \frac{1}{2}\ln\cosh(2J) \;=\; J^2 - \tfrac{2}{3}J^4 + O(J^6). The leading term of Theorem 4.1 is J^*_{ik}J^*_{jk}(1 - m_k^{2}) = J^2 — the environment is the single unmagnetised site k — matching sign and coefficient, and the exact remainder -\tfrac{2}{3}J^4 + O(J^6) is fourth order, as the theorem asserts. With a one-site environment, \chi_{\mathrm{env}} = 0 and E_{ij} = 0 exactly. Here u and v each have the single entry J, so the remainder bound reads |R_{ij}| \leq \tfrac{7}{48}(2J)^4. At J = 0.3: exact value 0.08507, leading term 0.0900, remainder -0.00493 against the bound 0.0189. At J = 0.25: exact value 0.06006, leading term 0.06250, remainder -0.00244 against the bound 0.00911. A pair with a shared neighbour looks more coupled than J^*_{ij} = 0: the path produces a positive effective coupling despite the zero direct interaction.
The expansion becomes informative when environmental correlations can themselves be bounded. The matrix below bounds how much changing one spin changes another site’s conditional probability. Its row-sum condition controls repeated propagation of these influences. The proof couples two sampling chains, meaning that it constructs their updates on one probability space to bound their disagreement.
Proposition 4.3 (Uniform environmental influence).
On the environment obtained by deleting the target pair, set C_{kl}=\tanh|J^*_{kl}|\ (k\ne l),\quad C_{kk}=0,\quad \alpha=\|C\|_\infty<1,\quad G=(I-C)^{-1}. For every finite environmental field vector, |\operatorname{Cov}(\omega_k,\omega_l)|\leq G_{kl},\qquad \chi_{\mathrm{env}}\leq\max_k\sum_{l\ne k}G_{kl}\leq\frac{\alpha}{1-\alpha}. The complete embedding bias satisfies |\Delta_{ij}|\leq \sum_k|u_kv_k|+\sum_{k\ne l}|u_k|G_{kl}|v_l| \leq D_{\mathrm{sh}}J_{\max}^2+DJ_{\max}^2\frac{\alpha}{1-\alpha}. \tag{11}
Proof. The conditional probability at site k is (1+\tanh(h_k+\sum_lJ^*_{kl}\omega_l))/2. Flipping site l changes it by at most \tanh|J^*_{kl}|=C_{kl}. Changing only field h_l by \varepsilon changes that site’s conditional probability by at most |\varepsilon|/2.
Couple the two finite random-scan heat-bath chains using the same update site and maximal coupling there. A stationary coupling exists. Its marginals are the two Gibbs laws because each marginal chain is irreducible. If q_k denotes its disagreement probability, stationarity gives q\leq Cq+|\varepsilon|e_l/2. Iterating and using C^n\to0 gives q\leq|\varepsilon|Ge_l/2. Thus |m_k(h+\varepsilon e_l)-m_k(h)|\leq|\varepsilon|G_{kl}. Differentiation of the finite partition sum yields \partial_{h_l}m_k=\operatorname{Cov}(\omega_k,\omega_l). The row sum of G-I=\sum_{n\geq1}C^n is at most \alpha/(1-\alpha). This is the finite-state influence comparison underlying the general Gibbs comparison theorem [35].
For total bias, integrate the mixed derivative of the environmental log moment-generating function: \Delta_{ij}=\frac14\int_{-1}^{1}\int_{-1}^{1} u^\top\operatorname{Cov}_{h+su+tv}(\omega)v\,ds\,dt. The four corner terms equal the coupling contrast in Theorem 4.1. Every tilt has the same matrix C. Diagonal covariances are at most one, and off-diagonal entries obey the uniform bound. Integration proves (11). ◻
Corollary 4.4 (Certified size of the embedding bias).
Write x=DJ_{\max}. If x\leq1/2, then \chi_{\mathrm{env}}\leq\frac{x}{1-x}\leq2x,\qquad |E_{ij}|\leq2D^2J_{\max}^3,\qquad |\Delta_{ij}|\leq\frac{DJ_{\max}^2}{1-x}\leq2DJ_{\max}^2. For |J^*_{ij}|\geq J_{\max}/2>0, the total relative bias is at most 2x/(1-x)\leq4x.
Proof. The inequality \tanh z\leq z gives \alpha\leq x. Substitute in Proposition 4.3 and use D_{\mathrm{sh}}\leq D. ◻
At D=D_{\mathrm{sh}}=5 and J_{\max}=.02, the complete bias is at most .002223. An independently established bound B yields J^*_{ij}\in[J^{\mathrm{eff}}_{ij}-B,J^{\mathrm{eff}}_{ij}+B]. An effective estimate whose entire uncertainty interval exceeds B certifies a positive direct interaction under those structural assumptions. The environmental bounds require independent support and interaction information. Effective pair estimates alone do not provide it.
Remark 4.5 (Environment marginals, correlations, and the size of E_{ij}).
Environment against network marginals. Theorem 4.1 states the two-hop sum at the environment marginals m_k, not at the full-network marginals m_k^*, because the substitution changes its error order.
Conditional on (\omega_i, \omega_j), the environment is exponentially tilted by t(\omega_i, \omega_j), a tilt of \ell^1-size at most \|u\|_1 + \|v\|_1 \leq 2 D J_{\max}. Each derivative \partial m_k / \partial t_l is a spin covariance bounded by one, and m_k^* averages the tilted marginals over the pair’s law, so |m_k^* - m_k| \leq 2 D J_{\max}. Replacing m_k by m_k^* moves the two-hop sum by up to 4 D_{\mathrm{sh}} D J_{\max}^3 — third order in the couplings, one order larger than R_{ij}.
A version of (8) written at m_k^* would carry that substitution error inside its remainder. The environment form avoids that substitution. (The expansion in the proof is of the environment generating function, not Plefka’s expansion of the free energy at fixed magnetisations [27]; in Plefka’s series it is the second order that yields the Onsager reaction correction of Thouless–Anderson–Palmer [12, 28].)
The size of \chi_{\mathrm{env}}. The unconditional bound is \chi_{\mathrm{env}}\leq M-3. Proposition 4.3 gives the sharper certified bound when \alpha<1. The mean-field matrix (I-J^*)^{-1} belongs to the approximation developed in Section 9 and supplies no exact covariance certificate.
An environmental correlation example. Take i and j each joined to n common neighbours which themselves form a clique, all couplings equal to J, all fields zero, so that J^*_{ij} = 0 and D = n + 1. At n = 8, J = 0.055 (so D J_{\max} = 0.495) the two-hop sum is 0.02420 and \Delta_{ij} exceeds it by 0.01315: the excess is E_{ij} = 0.01341 plus the remainder R_{ij} = -2.5 \times 10^{-4}.
At n = 14, J = 0.033: excess 0.01040, with E_{ij} = 0.01048 and R_{ij} = -8.2 \times 10^{-5}. At n = 100, J = 0.00495: E_{ij} = 0.00231 against a two-hop sum of 0.00245 — the correlation term comparable to the two-hop sum itself. Held at D J_{\max} \approx 1/2, the ratio of E_{ij} to D J_{\max}^3 grows linearly in D across the three cases (9.0, 19, 188), consistent with the O(D^2) count. The remainder stays negligible throughout.
5 Which Observations Identify Direct Interactions
The embedding calculation explains why a pair table can conceal the source of dependence. We now ask which additional information separates direct interactions from effects of the other events. Full observations and known network structure provide two constructive answers.
5.1 Full observations and a known tree
Holding every other outcome fixed removes the effect of averaging over hidden configurations. The conditional four-cell contrast can then isolate a direct interaction. On a known tree, there is only one path between any two sites, so consistent edge tables also determine the full law. Here a tree is a connected interaction graph with no cycles; E(T) is its edge set and d_i its vertex degrees.
Proposition 5.1 (Full observations and known trees).
A positive full-vector law in the pairwise Ising family identifies every direct interaction through J_{ij}=\frac14\log \frac{P(+,+,y)P(-,-,y)}{P(+,-,y)P(-,+,y)} for any fixed configuration y of the remaining sites. On a known tree T, positive consistent singleton and edge marginals reconstruct the law as P_T(x)=\frac{\prod_{\{i,j\}\in E(T)}p_{ij}(x_i,x_j)} {\prod_i p_i(x_i)^{d_i-1}}. Each tree-edge direct interaction equals its edge-table quarter-log odds ratio.
Proof. In the conditional four-cell contrast, fields, normalization, and all terms involving y cancel. The remaining contrast is 4J_{ij}. Root the tree. The displayed product becomes one root marginal times one normalized child-given-parent conditional for each edge. Successive leaf summation proves normalization and recovers the specified marginals. Expand each positive edge log-table in its two-spin basis. Its pair coefficient is the quarter-log odds ratio. Singleton denominator terms change only fields. ◻
Complete sufficient moments also identify the unrestricted pairwise model by strict convexity of its log partition function. Exact likelihood fitting uses those moments. Conditional logistic likelihood uses aligned full-vector observations and estimates the same direct interactions under its model [22, 15]. Interaction screening provides a sparse inverse-Ising comparison with its own sampling assumptions [16]. The conditional-table formula is a population identification proof. Sparse or tree structure can reduce the data and computation required for its estimation.
5.2 Selected pairs and missing interactions
Suppose quotes are available only for a subset \mathcal{P}\subseteq\binom{[M]}{2} of pairs, where [M]=\{1,\ldots,M\}. The support of J is the set of pairs with nonzero direct interaction. Knowing that this support lies in \mathcal P is additional information about the model. The same selected quotes can identify a model under this restriction and fail to identify an unrestricted network. The next proof makes the difference precise through the number and independence of the observed moments.
Proposition 5.2 (Identifiability from partial parlay coverage).
Let all M marginals be quoted and let \mathcal{P} \subseteq \binom{[M]}{2} be the set of pairs with quoted parlay prices.
- (i)
-
If the support of J is contained in \mathcal{P}, the quotes determine (h, (J_{ij})_{(i,j) \in \mathcal{P}}) — for every \mathcal{P}, connected or not.
- (ii)
- If J is unrestricted, identification fails for every proper \mathcal{P}.
Proof. (i) The statistics \{\omega_i\}_{i \in [M]} \cup \{\omega_i \omega_j : (i,j) \in \mathcal{P}\} are distinct Walsh functions on \{-1,+1\}^M, hence linearly independent with no nonzero combination constant, so the family is a minimal exponential family and its mean-parameter map is injective [18, Thm. 1.13]. The quoted marginals and parlay prices are an affine bijection of exactly those mean parameters: \mathbb{E}[\omega_i] = 2p_i - 1 and \mathbb{E}[\omega_i\omega_j] = 4\pi_{ij} - 2p_i - 2p_j + 1. (ii) The parameters have dimension M + \binom{M}{2}, the observations dimension M + |\mathcal{P}| < M + \binom{M}{2}, and a continuous injection from an open subset of the parameter space into the observation space would contradict invariance of domain. ◻
Failure to identify the full model leaves a more specific question. Could selected moments at least prove that an unobserved edge must exist? For interior feasible moments, their maximum-entropy explanation always uses only the selected interactions. The following result therefore rules out that conclusion from those observations alone.
Proposition 5.3 (Selected moments cannot certify a missing edge).
Let a selected-moment vector lie in the interior of the convex hull of the singleton and selected-pair statistics. Its maximum-entropy law has only those selected pair interactions. Thus every such observation admits an explanation whose interaction support lies inside the selected set.
Proof. A finite minimal exponential family maps its natural parameters onto the interior of its mean polytope [18, 36]. Its entropy maximizer is positive and has log density equal to a constant plus a linear combination of the constrained statistics. Unselected pair coefficients are zero. Therefore the observation alone cannot rule out that support. ◻
A concrete witness uses three zero-field spins. In one law, J_{02}=.4, J_{12}=.6, and J_{01}=0. The second law has only J_{01}=\operatorname{arctanh}(\tanh(.4)\tanh(.6))=.2069563793. Both laws have singleton means zero and the same (0,1) pair law. Both have full support. A detector that uses only those observations returns Nonidentifiable. Global quote inconsistency is a different finding. Proposition 3.4 tests compatibility of observations with any joint law.
Open Problem 1 (Additional observations for support recovery).
Choose new pair measurements or aligned full-vector observations that separate competing support models at a declared cost. The target is finite-sample recovery under stated support and sampling assumptions. A proposed acquisition rule must distinguish the two laws above using genuinely additional observations.
5.3 Acquiring additional observations
An observation budget must specify which independent rounds it purchases. Accuracy for a selected pair and recovery of network support are different objectives. The following question concerns how simultaneous pair estimation changes the required sample size.
Open Problem 2 (Sample complexity for simultaneous pair selection).
Theorem 6.4 gives a per-pair concentration bound. Simultaneously estimating all \binom{M}{2} interactions from a market with N_p resolutions on each pair requires a union bound scaling the sample complexity by O(\log M). When only a sparse subset of pairs is relevant (mean network degree D), can adaptive acquisition of independent observations across pairs achieve the oracle rate N_p \geq C(\theta^{\mathrm{eff}}) \epsilon^{-2} \log(D M / \delta)? The \ell_1-regularised estimators of [15, 16] answer this for spin-sample observations; the parlay-price version is open.
For a simpler allocation problem, suppose each pair’s variance coefficient is already known. The next proposition gives the best continuous allocation for a weighted sum of asymptotic variances. Here the index e denotes a pair, and w_e below is its specified decision weight.
Proposition 5.4 (Known-variance allocation).
For positive constants a_e and continuous counts n_e>0 with \sum_e n_e=N, the minimum of \sum_e a_e/n_e is \frac{(\sum_e\sqrt{a_e})^2}{N},\qquad n_e=\frac{N\sqrt{a_e}}{\sum_f\sqrt{a_f}}.
Proof. Cauchy–Schwarz gives (\sum_e\sqrt{a_e})^2\leq(\sum_e a_e/n_e)(\sum_e n_e). Equality requires n_e proportional to \sqrt{a_e}, which gives the displayed allocation. ◻
For coupling estimation, a_e=w_e/(16)\sum_s1/p_{e,s} is the known asymptotic variance coefficient. With a=(1,4,9) and N=600, the allocation (100,200,300) has objective .06. Equal allocation has objective .07. These counts are separately purchased independent rounds. A full-vector round supplies several pair observations simultaneously and needs a different acquisition-cost model.
Open Problem 3 (Adaptive observation allocation).
Extend the known-variance solution to unknown variances, integer counts, heterogeneous costs, shared full-vector observations, and finite-sample guarantees. An adaptive rule must account for the selection induced by its observations. Trading volume alone cannot increase the count of independently resolved rounds.
5.4 When pair interactions are insufficient
Some joint laws require interactions among three or more events. A complete set of conjunction probabilities determines such a law by inclusion–exclusion. The next transform generalizes the four-cell reconstruction and makes its information requirement explicit. The indicator X_i equals one when event i occurs and zero otherwise.
Proposition 5.5 (Complete higher-order reconstruction).
For k binary indicators, let u_A=P(X_i=1\text{ for all }i\in A) for every subset A, with u_\emptyset=1. The cell with ones exactly on B has probability p_B=\sum_{D\subseteq B^c}(-1)^{|D|}u_{B\cup D}. The moments are feasible exactly when these cells are nonnegative. For positive cells and spin form s\in\{-1,1\}^k, every saturated interaction is \theta_A=2^{-k}\sum_s\Bigl(\prod_{i\in A}s_i\Bigr)\log p(s),\qquad A\ne\emptyset.
Proof. Expand \prod_{j\notin B}(1-X_j) to obtain the first identity. Summation gives normalized cells and recovers each conjunction moment. The spin products form an orthogonal basis under uniform measure. Applying that basis transform to \log p proves the second identity. ◻
The complete table needs every required lower-order moment as well as the highest-order conjunction. A triple quote alone does not provide that table.
Open Problem 4 (Partial higher-order observation).
Determine sparse higher-order interactions from incomplete conjunction measurements with explicit uncertainty and acquisition cost. The complete-table transform above supplies the baseline. Partial observation, model selection, and high-dimensional efficiency remain separate tasks.
6 Estimation from Aligned Outcomes
6.1 The saturated maximum-likelihood estimator
The closed-form inversion (7) takes the pair law as given. Resolved rounds instead supply counts in its four cells. The maximum-likelihood estimator (MLE) selects the parameter that makes those recorded counts most probable. The model fits every positive four-cell table, so it is called saturated. Consequently, when all counts are positive, the empirical frequencies themselves maximize the likelihood.
This is the classical 2\times2 log-linear exponential family [29][30, §9.2].
The statistics (\omega_i,\omega_j,\omega_i\omega_j) are sufficient: their sums retain all dependence of the likelihood on the parameter. Woolf’s log-odds-ratio variance describes its large-sample precision [4]. Section 6.2 also gives a finite-sample bound. Both results require the sampling and target conditions stated next.
Assumption 6.1 (Regularity).
(R1) The pair law has four positive cell probabilities. (R2) The recorded rounds are independent and share that pair law. Event identities, alignment, selection, and the observation window define the target population. All contracts on one round share its single observation. The data-selection and prediction rules remain fixed over this batch. (R3) The three effective parameters are those of the unique saturated two-site law. Direct network recovery additionally requires the observation and support conditions stated in Section 5.1. A causal interpretation requires its own model. Pair-law inference itself remains valid with an unobserved common cause.
Definition 6.2 (Aligned outcome observation model).
Let the pair type (i, j) resolve in N_p settled rounds over a window. Each round contributes exactly one observation, the realised joint outcome (\omega_i, \omega_j) \in \{-1, +1\}^2, however many parlay contracts were written on that round; open interest does not enter N_p. The observation record binds both event identities, their round identity, resolution rule, and observation window. Deduplication uses the round identity. The N_p observations are independent draws from one pair law under Assumption 6.1. Contract listing is unnecessary. Selection rules must justify that common law. Let N_{++}, N_{+-}, N_{-+}, N_{--} count the four joint outcomes with N_p = \sum N_{\cdot\cdot}. Under the two-site Ising model with parameter \theta = (h_i, h_j, J_{ij}), the log-likelihood is \ell(\theta) \;=\; \sum_{s_i, s_j \in \{\pm 1\}} N_{s_i s_j}\bigl(h_i s_i + h_j s_j + J_{ij} s_i s_j\bigr) \;-\; N_p \ln Z(\theta), \tag{12} with Z(\theta) = \sum_{s_i, s_j} \exp(h_i s_i + h_j s_j + J_{ij} s_i s_j).
The estimand. The pair marginal of the M-site distribution is itself a two-site Ising law: the family is saturated on the four cells, so exactly one parameter triple fits it. Write \theta^{\mathrm{eff}} = (h_i^{\mathrm{eff}}, h_j^{\mathrm{eff}}, J^{\mathrm{eff}}_{ij}) for that triple. At M = 2 it is (h_i, h_j, J_{ij}); at M > 2, Theorem 4.1 gives J^{\mathrm{eff}}_{ij} = J^*_{ij} + \Delta_{ij} with J^* the network interaction matrix. Every theorem of this section and Section 6.2 is a statement about \theta^{\mathrm{eff}}, the parameter of the law that generates the observations: the concentration below is around J^{\mathrm{eff}}_{ij}, and the embedding bias \Delta_{ij} does not shrink with N_p. The expansion in Section 4 requires environmental information beyond this pair law. Direct interactions are recovered through the additional observations or support information of Section 5.1.
The Fisher information measures local likelihood curvature and determines the estimator’s asymptotic covariance. Estimating the fields together with the coupling reduces the information available for that coupling. The Schur complement below removes this field contribution from the corresponding matrix block. The proof derives the variance both from this matrix and directly from the four empirical cell frequencies. An output labelled Boundary records an absent empirical cell and retains its uncertainty set, as Section 6.3 specifies.
Theorem 6.3 (Saturated MLE: closed form, asymptotic normality, Woolf variance).
Let the truth \theta^{\mathrm{eff}} be interior, with moments m_i^* = \mathbb{E}[\omega_i], m_j^* = \mathbb{E}[\omega_j], C_{ij}^* = \mathbb{E}[\omega_i\omega_j] and cell probabilities p^*_{s_i s_j} = \tfrac{1}{4}(1 + m_i^* s_i + m_j^* s_j + C_{ij}^* s_i s_j), under Assumption 6.1. Then:
- (i)
-
(The MLE is the closed form.) On the event that all four counts are positive, the log-likelihood (12) is uniquely maximised at the \theta whose cell probabilities equal the empirical frequencies N_{s_i s_j}/N_p; in particular \hat J_{ij} \;=\; \frac{1}{4}\,\ln\frac{N_{++}\,N_{--}}{N_{+-}\,N_{-+}}, \tag{13} the closed form (7) evaluated on empirical frequencies. No optimisation step exists.
- (ii)
-
(Consistency and asymptotic normality.) For these limits, extend the estimate measurably to zero-count samples, for example by the zero vector. Operational output at those samples remains Boundary. \hat\theta_{N_p} \xrightarrow{a.s.} \theta^{\mathrm{eff}}, and \sqrt{N_p}\,\bigl(\hat\theta_{N_p} - \theta^{\mathrm{eff}}\bigr) \;\xrightarrow{d}\; \mathcal{N}\bigl(0,\; I(\theta^{\mathrm{eff}})^{-1}\bigr), where I(\theta^{\mathrm{eff}}) = \mathop{\mathrm{Cov}}_{\theta^{\mathrm{eff}}}(T) with T = (\omega_i, \omega_j, \omega_i\omega_j): I(\theta^{\mathrm{eff}}) \;=\; \begin{pmatrix} 1 - m_i^{*2} & C_{ij}^* - m_i^* m_j^* & m_j^* - m_i^* C_{ij}^* \\ C_{ij}^* - m_i^* m_j^* & 1 - m_j^{*2} & m_i^* - m_j^* C_{ij}^* \\ m_j^* - m_i^* C_{ij}^* & m_i^* - m_j^* C_{ij}^* & 1 - C_{ij}^{*2} \end{pmatrix}. \tag{14} I(\theta) is positive-definite at every interior \theta: the family is minimal, so \mathop{\mathrm{Cov}}_\theta(T) is nonsingular whenever all four cell probabilities are positive. No condition on the size of the couplings is involved.
- (iii)
- (Efficient variance is Woolf’s variance.) Partition I(\theta^{\mathrm{eff}}) = \begin{pmatrix} A & b \\ b^\top & I_{33} \end{pmatrix} with A \in \mathbb{R}^{2 \times 2} the field-sector block, b = (m_j^* - m_i^* C_{ij}^*,\; m_i^* - m_j^* C_{ij}^*)^\top, I_{33} = 1 - C_{ij}^{*2}, and write S := I_{33} - b^\top A^{-1} b for the Schur complement. The asymptotic variance of \hat J_{ij} is \mathop{\mathrm{Var}}_{\infty}(\hat J_{ij}) \;=\; \frac{[I(\theta^{\mathrm{eff}})^{-1}]_{33}}{N_p} \;=\; \frac{1}{N_p\, S} \;=\; \frac{1}{16\, N_p}\left(\frac{1}{p^*_{++}} + \frac{1}{p^*_{+-}} + \frac{1}{p^*_{-+}} + \frac{1}{p^*_{--}}\right), \tag{15} the variance Woolf derived for the sample log odds ratio [4]; \hat J_{ij} is one quarter of the sample log odds ratio. Moreover S \leq I_{33}, with equality iff b = 0, which at interior \theta^{\mathrm{eff}} holds iff m_i^* = m_j^* = 0. Whenever (m_i^*, m_j^*) \neq (0, 0), the marginal bound 1/[N_p(1 - C_{ij}^{*2})] is unattainable: the attainable variance exceeds it by the factor I_{33}/S > 1 (13\% in the regime of Remark 12.2).
Proof. (i) The model is a minimal regular exponential family (Section 12.1), so the log-likelihood is strictly concave (Hessian -N_p I(\theta) \prec 0) and its maximiser solves the moment equations \bar T = \mathbb{E}_\theta[T]. The mean-parameter map \theta \mapsto \mathbb{E}_\theta T is a bijection onto the interior of the convex hull of the T-values, equivalently onto the strictly positive cell-probability vectors; with all four counts positive the empirical distribution is such a vector, so the MLE fits it exactly. Applying the cross-ratio identity from the proof of Theorem 3.1 to the fitted cells gives (13).
(ii) By the strong law, the empirical frequencies converge a.s. to the positive cells p^*_{s_i s_j}; eventually all four counts are positive, and continuity of the inverse mean-parameter map gives \hat\theta_{N_p} \to \theta^{\mathrm{eff}} a.s. For normality: \sqrt{N_p}(\bar T - \mathbb{E}T) \xrightarrow{d} \mathcal{N}(0, \mathop{\mathrm{Cov}}(T)), and the delta method through the smooth inverse map, with Jacobian \partial\theta/\partial\mu = I(\theta^{\mathrm{eff}})^{-1} (exponential-family duality [18]), gives asymptotic covariance I^{-1}\, I\, I^{-1} = I(\theta^{\mathrm{eff}})^{-1} [17]. The entries of I = \mathop{\mathrm{Cov}}(T) follow from \omega_k^2 \equiv 1: \mathop{\mathrm{Var}}(\omega_i) = 1 - m_i^{*2}, \mathop{\mathrm{Var}}(\omega_i\omega_j) = 1 - C_{ij}^{*2}, \mathop{\mathrm{Cov}}(\omega_i, \omega_i\omega_j) = \mathbb{E}[\omega_i^2\omega_j] - m_i^* C_{ij}^* = m_j^* - m_i^* C_{ij}^*, and so on. Minimality (affine independence of the three statistics on the four-point support) gives \mathop{\mathrm{Cov}}_\theta(T) \succ 0 at every interior \theta.
(iii) [I^{-1}]_{33} = S^{-1} is the Schur-complement identity (Lemma 12.1). For the Woolf form, apply the delta method directly to \hat J_{ij} = \tfrac{1}{4}(\ln \hat p_{++} + \ln \hat p_{--} - \ln \hat p_{+-} - \ln \hat p_{-+}) on the multinomial: the gradient is \tfrac{1}{4}(1/p_{++},\, -1/p_{+-},\, -1/p_{-+},\, 1/p_{--}), the multinomial covariance is \mathop{\mathrm{diag}}(p) - p p^\top, and the p p^\top term vanishes because the gradient’s inner product with p is \tfrac{1}{4}(1 - 1 - 1 + 1) = 0; what remains is \tfrac{1}{16}\sum_{\text{cells}} 1/p^*_{s_i s_j}. Equality of the two expressions for the same asymptotic variance gives S^{-1} = \tfrac{1}{16}\sum 1/p^*_{s_i s_j}; Section 12.2 records the direct algebraic verification. For the equality locus: A^{-1} \succ 0 gives b^\top A^{-1} b \geq 0 with equality iff b = 0; and b = 0 forces m_j^* = m_i^* C_{ij}^* and m_i^* = m_j^* C_{ij}^*, hence m_i^*(1 - C_{ij}^{*2}) = 0, hence m_i^* = m_j^* = 0 at interior \theta^{\mathrm{eff}} (|C_{ij}^*| < 1); the converse is immediate. ◻
6.2 Non-asymptotic concentration
Theorem 6.3(ii) describes a limit as the number of rounds grows. A finite-sample guarantee must also control the probability of an inaccurate estimate at the actual sample size. The proof first bounds errors in the three empirical moments using their bounded values in \{-1,+1\}. It then propagates those errors through the smooth inverse map. This requires a local region that stays away from every zero-cell boundary, which explains the radius in the theorem.
Theorem 6.4 (Non-asymptotic concentration of \hat J_{ij}, local to \theta^{\mathrm{eff}}).
Under the hypotheses of Theorem 6.3, with S and A as in Theorem 6.3(iii), set C(\theta^{\mathrm{eff}}) \;=\; \frac{18}{S^2\,\lambda_{\min}(A)^2}, \tag{16} with \lambda_{\min}(A) > 0 the smallest eigenvalue of the field-sector Fisher block. There is a radius \epsilon_0(\theta^{\mathrm{eff}}) > 0, defined in Step 3 of the proof and positive at every interior \theta^{\mathrm{eff}}, such that for every \epsilon \in (0, \epsilon_0(\theta^{\mathrm{eff}})) and \delta \in (0, 1), \mathbb{P}\bigl(\text{zero count or }|\hat J_{ij, N_p} - J^{\mathrm{eff}}_{ij}|>\epsilon\bigr) \;\leq\; \delta \quad\text{whenever}\quad N_p \;\geq\; \frac{C(\theta^{\mathrm{eff}})}{\epsilon^2}\,\log\!\frac{6}{\delta}. \tag{17} The relevant boundary consists of all zero-cell facets: 1+s_i m_i^*+s_jm_j^*+s_i s_j C_{ij}^*=0 for some (s_i,s_j). The radius stays positive at each interior law. In the failure event, the displayed estimate is evaluated only when all counts are positive. The rate \epsilon^{-2}\log(1/\delta) is the parametric MLE rate for a fixed regular exponential family [18]; Remark 6.5 shows by a two-point test on the two-site family that no estimator improves it by more than a constant.
Proof. Step 1: Hoeffding on the sufficient-statistic averages. Each \bar T_k := N_p^{-1}\sum_r T_k(\omega^{(r)}), k = 1, 2, 3, is an average of i.i.d. \{-1,+1\}-valued variables, so Hoeffding’s inequality gives \mathbb{P}(|\bar T_k - \mathbb{E}[T_k]| > t) \leq 2 \exp(-N_p t^2/2). A union bound over the three components gives, with probability at least 1 - \delta, \|\bar T - \mu^*\|_\infty \;\leq\; t(N_p, \delta) \;:=\; \sqrt{\frac{2}{N_p}\log\frac{6}{\delta}}, \qquad \mu^* := \mathbb{E}_{\theta^{\mathrm{eff}}}[T].
Step 2: the J-row of the inverse Fisher matrix at \theta^{\mathrm{eff}}. By the block-inverse formula [21, §0.8], [I^{-1}]_{3,\cdot} = \bigl(-S^{-1}(A^{-1}b)^\top,\; S^{-1}\bigr), so \|[I^{-1}]_{3, \cdot}\|_2^2 \;=\; S^{-2}\bigl(1 + \|A^{-1} b\|_2^2\bigr). Each entry of b is a covariance of \{-1,+1\}-valued variables, so |b_k| \leq 1 by Cauchy-Schwarz and \|b\|_2 \leq \sqrt{2}; hence \|A^{-1} b\|_2 \leq \sqrt{2}/\lambda_{\min}(A). With \lambda_{\min}(A) \leq A_{11} = 1 - m_i^{*2} \leq 1, \|[I^{-1}]_{3, \cdot}\|_2^2 \;\leq\; S^{-2}\Bigl(1 + \frac{2}{\lambda_{\min}(A)^2}\Bigr) \;\leq\; \frac{3}{S^2\,\lambda_{\min}(A)^2}, and the combined bound is strict at every interior \theta^{\mathrm{eff}}: equality in the second inequality forces \lambda_{\min}(A) = 1, hence A = I and b = 0, which makes the first inequality strict.
Step 3: the radius. The mean-parameter map \mu(\theta) = \mathbb{E}_\theta[T] is a C^\infty diffeomorphism onto \mathcal{M}, the interior of the convex hull of the T-values, with \partial\theta/\partial\mu = I(\theta)^{-1} (exponential-family duality [18]). The function \mu \mapsto \|[I(\theta(\mu))^{-1}]_{3,\cdot}\|_2 is continuous and, by Step 2, strictly below \sqrt{3}/(S\,\lambda_{\min}(A)) at \mu^*; writing B_\rho for the closed \rho-ball around \mu^*, \rho_0(\theta^{\mathrm{eff}}) \;:=\; \sup\Bigl\{\rho > 0 \;:\; B_\rho \subset \mathcal{M}\;\text{ and }\; \sup_{\mu \in B_\rho} \bigl\|[I(\theta(\mu))^{-1}]_{3,\cdot}\bigr\|_2 \leq \tfrac{\sqrt{3}}{S\,\lambda_{\min}(A)}\Bigr\} is therefore positive. Set \epsilon_0(\theta^{\mathrm{eff}}) := \sqrt{3}\,\rho_0(\theta^{\mathrm{eff}})/(S\,\lambda_{\min}(A)).
Step 4: propagation. Take N_p \geq C(\theta^{\mathrm{eff}})\,\epsilon^{-2}\log(6/\delta) with \epsilon < \epsilon_0(\theta^{\mathrm{eff}}). On the Step-1 event, \|\bar T - \mu^*\|_2 \leq \sqrt{3}\, t(N_p, \delta) \leq \epsilon\, S\,\lambda_{\min}(A)/\sqrt{3} < \rho_0(\theta^{\mathrm{eff}}). Choose an admissible radius \rho strictly between \|\bar T-\mu^*\|_2 and \rho_0. Such a radius exists because the admissible radii form a downward-closed interval. The segment lies in B_\rho, all empirical cells are positive, and the MLE exists. The mean-value inequality along the segment gives \begin{align*} \bigl|\hat J_{ij, N_p} - J^{\mathrm{eff}}_{ij}\bigr| \;&\leq\; \sup_{\text{segment}} \bigl\|[I^{-1}]_{3,\cdot}\bigr\|_2 \,\cdot\, \|\bar T - \mu^*\|_2\\ &\leq\; \frac{\sqrt{3}}{S\,\lambda_{\min}(A)} \cdot \sqrt{3}\, t(N_p, \delta) \;=\; \frac{3\, t(N_p, \delta)}{S\,\lambda_{\min}(A)} \;\leq\; \epsilon, \end{align*} where the last inequality substitutes t(N_p, \delta) at the stated N_p. This is (17). ◻
Remark 6.5 (The rate is tight).
Write \theta_0 = \theta^{\mathrm{eff}} = (h_i, h_j, J^{\mathrm{eff}}_{ij}) and \theta_1 = (h_i, h_j, J^{\mathrm{eff}}_{ij} + 3\epsilon), both interior since \Theta = \mathbb{R}^3. With A the log-partition function of Section 12.1, \mathrm{KL}(\mathbb{P}_{\theta_0} \,\|\, \mathbb{P}_{\theta_1}) = A(\theta_1) - A(\theta_0) - \langle \theta_1 - \theta_0, \nabla A(\theta_0) \rangle = \tfrac{1}{2}(3\epsilon)^2\,\mathop{\mathrm{Var}}_{\bar\theta}(\omega_i\omega_j) \leq \tfrac{9}{2}\epsilon^2 for some \bar\theta on the segment, because \nabla^2 A(\theta) = \mathop{\mathrm{Cov}}_\theta(T) and \omega_i\omega_j \in \{-1, +1\}. Let \tilde J be any estimator built from N_p i.i.d. observations with \mathbb{P}_\theta(|\tilde J - J_{ij}(\theta)| > \epsilon) \leq \delta at both \theta = \theta_0 and \theta = \theta_1. The acceptance sets \{|\tilde J - J_{ij}(\theta_0)| \leq \epsilon\} and \{|\tilde J - J_{ij}(\theta_1)| \leq \epsilon\} are disjoint, so the event E := \{|\tilde J - J_{ij}(\theta_0)| > \epsilon\} satisfies 2\delta \;\geq\; \mathbb{P}^{\otimes N_p}_{\theta_0}(E) + \mathbb{P}^{\otimes N_p}_{\theta_1}(E^{c}) \;\geq\; \tfrac{1}{2}\exp\bigl(-N_p\,\mathrm{KL}(\mathbb{P}_{\theta_0} \,\|\, \mathbb{P}_{\theta_1})\bigr) \;\geq\; \tfrac{1}{2}\exp\bigl(-\tfrac{9}{2}N_p\epsilon^2\bigr) by the Bretagnolle–Huber inequality [20, Lemma 2.6], hence N_p \geq \tfrac{2}{9}\,\epsilon^{-2}\log\tfrac{1}{4\delta} for every \delta < 1/4. No estimator, the MLE included, meets the accuracy target at both points with fewer observations, so the \epsilon^{-2}\log(1/\delta) form of (17) is optimal up to the constant C(\theta^{\mathrm{eff}}).
Remark 6.6 (Comparison with the asymptotic variance).
The asymptotic variance of Theorem 6.3(iii) is 1/(N_p S); the finite-sample constant C(\theta^{\mathrm{eff}}) = 18/(S^2 \lambda_{\min}(A)^2) carries the extra factor 18/(S \lambda_{\min}(A)^2), the price of non-asymptotic control through Hoeffding and a worst-case union bound. In the zero-marginal locus m_i^* = m_j^* = 0: b = 0, A = \begin{pmatrix} 1 & C_{ij}^* \\ C_{ij}^* & 1 \end{pmatrix}, \lambda_{\min}(A) = 1 - |C_{ij}^*|, S = 1 - C_{ij}^{*2}, and C(\theta^{\mathrm{eff}}) = 18/[(1 - C_{ij}^{*2})^2(1 - |C_{ij}^*|)^2].
Example 6.7 (Sample complexity at \epsilon = 0.01).
For the regime (m_i^*, m_j^*, C_{ij}^*) = (0.3, 0.3, 0.2) of Remark 12.2: S \approx 0.847, \lambda_{\min}(A) = 0.80, so C(\theta^{\mathrm{eff}}) = 18/(0.847^2 \cdot 0.80^2) \approx 39.2. At \epsilon = \delta = 0.01: N_p \geq 2{,}507{,}478 independent rounds. At \epsilon = \delta = 0.05: N_p \geq 39.2 \cdot 400 \cdot \log 120 \approx 7.5 \times 10^4. These are sufficient-sample bounds from the Hoeffding route. Both targets lie below the radius of Theorem 6.4: at this \theta^{\mathrm{eff}} the row norm \|[I^{-1}]_{3,\cdot}\|_2 = 1.24 sits well under the Step-3 cap \sqrt{3}/(S\,\lambda_{\min}(A)) = 2.56, and the analytic bound below gives a row norm at most 1.432 throughout the radius-0.02 ball, so \rho_0 \geq 0.02 and \epsilon_0 \geq \sqrt{3} \cdot 0.02 / (S\,\lambda_{\min}(A)) = 0.051 > 0.05. The central-limit approximation calibrated to the same target, N_p \approx z_{0.995}^2/(S \epsilon^2) \approx 7.8 \times 10^4 at \epsilon = \delta = 0.01, is smaller by a factor of about 32; the gap is the price of an explicit finite-sample guarantee.
For completeness, write a_s=(s_i,s_j,s_i s_j), c_s=s_i s_j, and p_{\min}=\min_s p_s(\mu^*). Since p_s(\mu)=(1+a_s\cdot\mu)/4, \nabla J(\mu)=\frac1{16}\sum_s\frac{c_s a_s}{p_s(\mu)},\qquad \|\nabla^2J(\mu)\|_{\mathrm{op}}\leq\frac{3}{16k_\rho^2},\quad k_\rho=p_{\min}-\frac{\sqrt3\rho}{4}>0. The Hessian bound follows from \|a_sa_s^\top\|_{\mathrm{op}}=3. Integrating it gives the row-norm bound \|\nabla J(\mu^*)\|_2+3\rho/(16k_\rho^2). At the example’s cells (.45,.20,.20,.15), \rho=.02 gives k_\rho=.1413397 and a bound of 1.431917.
6.3 Confidence sets with finite and zero counts
The preceding concentration constant depends on the unknown population law. A cell confidence set can instead be constructed directly from the observed counts. It contains the true four-cell vector with a declared probability and remains defined when a cell has no observations. The construction bounds each cell count by a binomial interval, then intersects the four intervals with the probability simplex. That simplex is the set of nonnegative cell vectors whose coordinates sum to one.
Proposition 6.8 (Simultaneous cell confidence set).
Let n>0 independent rounds with the same four-cell law p^* have cell counts k_s, and fix 0<\delta<1. For each cell, take the Clopper–Pearson binomial interval [L_s,U_s] with each tail probability \delta/8 [34]. Define \mathcal C_n=\{p\in[0,1]^4:\ \sum_s p_s=1,\ L_s\leq p_s\leq U_s\}. Then P(p^*\in\mathcal C_n)\geq1-\delta. For a positive population law, the image of its positive cells under the inverse (7) is a simultaneous parameter confidence set. For a contrast \theta_c=\tfrac14\sum_s c_s\log p_s, where c_s\in\{-1,1\}, a conservative coordinate enclosure is \frac14\log\frac{\prod_{c_s=1}L_s}{\prod_{c_s=-1}U_s} \ \leq\ \theta_c\ \leq\ \frac14\log\frac{\prod_{c_s=1}U_s}{\prod_{c_s=-1}L_s}. A zero lower product denotes an unbounded endpoint. The simplex constraint can tighten these enclosures.
Proof. Each count has binomial marginal law with success probability p_s^*. Binomial-tail inversion gives failure probability at most \delta/4 for its interval. A union bound over four cells proves simultaneous coverage. Dependence between the four counts does not affect that bound. The true cell vector sums to one, so intersection with the simplex preserves the event. Monotonicity of the logarithm proves each enclosure on positive cells. ◻
For 0<k_s<n, the endpoints are beta quantiles with parameters (k_s,n-k_s+1) and (k_s+1,n-k_s). Set L_s=0 at k_s=0 and U_s=1 at k_s=n. An absent empirical cell returns Boundary and the confidence set. The finite unregularized MLE does not exist there. An optional estimate (k_s+\alpha)/(n+4\alpha), with declared \alpha>0, returns Regularized. Its finite inverse is distinct from the unregularized MLE and retains the confidence set. The vector (0,.2,.3,.5) is a boundary law with moments (-.6,-.4,0).
6.4 Independent rounds, pooling, and decisions
A data record counts independent resolved rounds for each estimand. It records the horizon, selection rule, missing resolutions, and any shared shocks. Exchangeability alone does not establish independence. Pooling requires a common pair law and independent units, or a replacement dependence-aware guarantee. Five independent strata with the same pair law can contribute 520 rounds each to a 2600-round estimate. Repeated tickets on one result contribute one round.
The illustrative concentration bound requires 75{,}065 rounds at \epsilon=\delta=.05, and 2{,}507{,}478 at .01. At one round per week, 75{,}000 rounds represent 1442.3 years using 52 weeks per year. Practical decisions therefore use the available sample size and its reported uncertainty. The exact cell construction supports modest samples without asserting a narrow interval. Section 10 reports synthetic coverage and shrinkage error at 104, 520, 2600, and 5000 rounds. A positive decision witness uses counts (2250,1000,1000,750). Its 95\% coupling enclosure is [.06428,.19745], which supports a positive effective coupling. Empirical utility requires prospective resolutions, a declared decision loss, and comparison at the actual data horizon.
7 Inference under a Changing Outcome Law
A longer observation window reduces sampling error when its law stays fixed. When the law changes, that window also combines different distributions. Its empirical frequencies estimate their average, whose relation to the current law depends on every cell probability. The fields determine base rates and can change this average while the interaction remains fixed. The following example shows why a bound on coupling drift alone is insufficient.
Example 7.1 (Pooling independent regimes).
Take equally many independent rounds from two regimes. Within each regime, the two Bernoulli variables are independent. Their success probabilities are (.8,.8) in the first regime and (.2,.2) in the second. In cell order (++,+-,-+,--), the laws and their average are p^{(1)}=(.64,.16,.16,.04),\qquad p^{(2)}=(.04,.16,.16,.64),\qquad \overline p=(.34,.16,.16,.34). Both regime interactions equal zero. The pooled effective coupling equals J(\overline p)=\frac12\log\frac{17}{8}=.3768859012\ldots. Thus a bound on changes in the interaction alone cannot control pooling bias.
Let \mathcal{F}_s contain the information available through round s; the increasing sequence (\mathcal{F}_s) is a filtration. Let Z_s be the aligned outcome pair at that round. Write p_s(z)=\mathbb{P}(Z_s=z\mid\mathcal{F}_{s-1}) for its conditional four-cell law. These laws can depend on earlier observations. For a deterministic window W=(t,N), define \widehat p_W(z)=\frac1N\sum_{s=t-N+1}^t\mathbf1\{Z_s=z\},\qquad \overline p_W(z)=\frac1N\sum_{s=t-N+1}^tp_s(z). The target is p_t, the conditional law of round t given its preceding history. Prediction of a later round requires an additional bound connecting its law to p_t.
A catalogue here is a list of candidate windows fixed before their observations are inspected. Each window receives part of the allowed failure probability. The next bound adds sampling error to a separate bound on how its average law differs from its current target. Simultaneous coverage across the catalogue then permits selection among those windows after seeing the data.
Proposition 7.2 (Simultaneous inference with full-law drift).
Fix a finite deterministic catalogue \mathcal W of windows and positive budgets \delta_W with \sum_W\delta_W\leq\delta<1. Let D_W\geq0 satisfy, on one event G with \mathbb{P}(G)\geq1-\rho, \|\overline p_W-p_t\|_\infty\leq D_W\qquad\text{for every }W=(t,N)\in\mathcal W. Assume \delta+\rho<1. Define a_W=\sqrt{\frac{\log(8/\delta_W)}{2N}},\qquad \eta_W=a_W+D_W, \mathcal C_W=\{q\in\Delta_3:|q_z-\widehat p_W(z)|\leq\eta_W\text{ for every }z\}, where \Delta_3=\{q\in[0,1]^4:\sum_zq_z=1\}. With probability at least 1-\delta-\rho, every p_t belongs to its corresponding \mathcal C_W. The conclusion holds simultaneously for any window selected from \mathcal W after observing the data.
Proof. For one cell, X_s=\mathbf1\{Z_s=z\}-p_s(z) has conditional mean zero and conditional range of length one. The logarithm of its conditional moment generating function has value and derivative zero at the origin. Its second derivative is a tilted Bernoulli variance, at most 1/4. Consequently \mathbb{E}(e^{\lambda X_s}\mid\mathcal{F}_{s-1})\leq e^{\lambda^2/8} for every real \lambda. Iterated conditional expectation over the window and exponential Markov inequality give \mathbb{P}\left(\left|\widehat p_W(z)-\overline p_W(z)\right|>a\right) \leq 2e^{-2Na^2}. This is the bounded-increment argument underlying Hoeffding’s inequality [38]. A union bound over four cells gives failure probability at most \delta_W at a=a_W. A second union bound covers the entire catalogue. Intersecting that event with G costs at most \rho more. The triangle inequality gives every asserted cell bound. Each true law lies in the simplex. Selection preserves a statement that already holds for every catalogue entry. ◻
The envelopes D_W can be random. Their simultaneous validity on G is a separate premise. Predictability of an envelope, or passage of a change detector, does not establish that premise. The guarantee bounds the probability of any false selected interval. Conditioning on a decision to issue an interval requires a separate argument. For several pairs, include every tested pair, time, and window in the catalogue and its failure budget. Dependence between pairs does not change the union bound.
Corollary 7.3 (Parameter intervals and the cell boundary).
Set \ell_z=\max(0,\widehat p_W(z)-\eta_W) and u_z=\min(1,\widehat p_W(z)+\eta_W). For a positive target law, project \mathcal C_W\cap(0,1)^4 through the two-site inverse. For each balanced contrast \theta_c(q)=\tfrac14\sum_zc_z\log q_z, with two signs of each kind, an enclosure is \frac14\log\frac{\prod_{c_z=1}\ell_z}{\prod_{c_z=-1}u_z} \ \leq\theta_c(p_t)\leq\ \frac14\log\frac{\prod_{c_z=1}u_z}{\prod_{c_z=-1}\ell_z}. A zero lower product gives the corresponding infinite endpoint. If p_t(z)\geq\kappa>\eta_W for every cell, the empirical inverse exists on the same event and |\theta_c(\widehat p_W)-\theta_c(p_t)| \leq\frac12\log\frac{\kappa+\eta_W}{\kappa-\eta_W}.
Proof. Coordinate monotonicity gives the enclosure, with an outer limit at zero lower endpoints. Under the interior premise, each cell ratio lies in [1-r,1+r], where r=\eta_W/\kappa<1. Two positive and two negative signs bound the contrast error by \tfrac12\{\log(1+r)-\log(1-r)\} in either direction. ◻
An empty empirical cell therefore returns a cell uncertainty set and a boundary estimate, as in Section 6.3. It does not supply a finite unregularized coupling. A true boundary law also remains a distributional output outside the finite Ising parameter family. Pseudocounts can define a separate regularized estimate while the same uncertainty set remains in force.
Suppose \|p_s-p_{s-1}\|_\infty\leq L throughout the relevant histories and rounds. The triangle inequality gives D_W=L(N-1)/2 almost surely, so \rho=0. For M windows with equal budgets, write A=\sqrt{\log(8M/\delta)/2}. Then \eta_N=A N^{-1/2}+\frac{L(N-1)}2. For L>0, the continuous minimizer of this radius is N=(A/L)^{2/3}. Choose an available integer window by its actual interval width, including the cell boundary. For L=0, the radius decreases with the window length. These formulas minimize an explicit confidence radius. A minimax rate claim requires a specified loss, drift class, and lower bound.
The construction combines bounded-increment concentration with a finite failure budget. Time-uniform confidence sequences supply a broader sequential framework [39]. A countable deterministic catalogue also works when its positive failure budgets have summable total at most \delta. Catalogues chosen from observations require coverage of every possible selected entry, or an independently justified selection procedure. Observations missing according to their outcomes change the conditional law of the recorded sample. The sampling contract must identify that law and its target population.
Open Problem 5 (Learning drift and acquiring observations).
Infer simultaneous full-law drift envelopes from structural information and observations, then select measurements by their expected reduction in decision loss. An estimated envelope must carry its own failure budget. Unknown change points, incomplete outcomes, and costs of collecting joint resolutions enter that construction. The finite-window result supplies a usable inference rule whenever those envelopes and aligned observations are available.
7.1 Certified evaluation of the confidence interval
Near a cell boundary, numerical rounding can decide whether a parameter interval is finite. An outward enclosure rounds the lower endpoint down and the upper endpoint up, so the exact interval remains inside it. Exact rational inputs permit these enclosures through every arithmetic operation in Proposition 7.2. The required logarithms and square roots reduce to a positive series and an integer square root. The first two lemmas justify these numerical bounds; the corollary applies them to the statistical interval without reducing its coverage.
Lemma 7.4 (Rational logarithm enclosure).
For rational x>0, choose the unique integer k such that x=2^ky with 1\leq y<2. Set z=(y-1)/(y+1). For N\geq1, define S_N(z)=2\sum_{j=0}^{N-1}\frac{z^{2j+1}}{2j+1},\qquad R_N(z)=\frac{2z^{2N+1}}{(2N+1)(1-z^2)}. Then 0\leq z<1/3 and S_N(z)\leq\log y\leq S_N(z)+R_N(z). The same enclosure at z=1/3 bounds \log2. Combining these intervals as k\log2+\log y, with endpoint reversal when k<0, encloses \log x by rationals. If both component widths are at most 2^{-b}/(|k|+1), the resulting width is at most 2^{-b}.
Proof. Integrating the geometric series for 2/(1-z^2) on [0,z] gives \log((1+z)/(1-z))=2\sum_{j\geq0}z^{2j+1}/(2j+1). All terms are nonnegative. Replacing every omitted denominator by 2N+1 and summing the geometric tail proves the enclosure. At z=1/3, the remainder is 3/\{4(2N+1)9^N\}. Thus each positive rational tolerance is reached after finitely many terms. Signed interval multiplication gives width at most |k|R_N(1/3)+R_N(z), which proves the last claim. The case x=1 returns the exact interval [0,0]. ◻
Lemma 7.5 (Scaled square-root enclosure).
For rational q>0, choose integer k so that q=2^{2k}v with 1\leq v<4. Given b\geq1, let g=2^b and m=\left\lfloor\sqrt{\left\lfloor vg^2\right\rfloor}\right\rfloor. The integer square root determines m exactly. Then 2^k\frac{m}{g}\leq\sqrt q\leq2^k\frac{m+1}{g}. If the lower endpoint squares to q, both endpoints can equal it. Otherwise the interval width is 2^{k-b}, at most 2^{-b}\sqrt q. For q=0, return [0,0].
Proof. The definition gives m^2\leq\lfloor vg^2\rfloor<(m+1)^2. Since both squares are integers, m^2\leq vg^2<(m+1)^2. Taking nonnegative roots and multiplying by 2^k/g proves the enclosure. The width bound uses \sqrt q\geq2^k. ◻
Corollary 7.6 (Computed intervals preserve simultaneous coverage).
Assume the hypotheses of Proposition 7.2 and rational failure budgets and drift envelopes. Enclose \log(8/\delta_W) by Lemma 7.4, divide by 2N, and apply Lemma 7.5 to both endpoints. Add D_W exactly to obtain a rational interval [\eta_W^-,\eta_W^+] containing \eta_W. Use \eta_W^+ in each cell endpoint of Corollary 7.3. Evaluate every positive endpoint odds ratio exactly, then enclose its logarithm and divide by four. Take the lower logarithm bound for a lower endpoint and the upper bound for an upper endpoint. For positive target laws, the resulting intervals contain the target parameters simultaneously with probability at least 1-\delta-\rho. Selection of the narrowest computed enclosure preserves this guarantee. Zero lower products retain their infinite endpoints.
Proof. Every arithmetic enclosure is outward. Monotonicity of the square root and logarithm therefore preserves inclusion at each stage. Using \eta_W^+ expands the cell box and its parameter enclosure. The simultaneous statistical event already covers every declared window, including the selected entry. All finite enclosure widths are rational and can be compared exactly. An unbounded interval has infinite width, with ties resolved by catalogue order. ◻
The integer b specifies component precision. The final interval reports its actual rational endpoints and width. Close to the boundary, test \eta_W^+<\kappa before using the interior error bound. If that test fails, retain the cell uncertainty set and its potentially infinite parameter interval. Outward decimal display uses integer floor for a lower endpoint and integer ceiling for an upper endpoint after multiplication by 10^d. For example, -1/3 displays as [-.3334,-.3333] at four decimal places. The calculation accepts rational inputs, exact cell counts, and an immutable catalogue. A resource limit returns an explicit failure. These arithmetic guarantees preserve the specified confidence statement. Sampling validity and drift evidence remain its statistical premises.
8 How Quote Error Changes the Inferred Parameters
A quote model can carry explicit uncertainty before any conversion to parameters. The reference law may be a pricing law or a physical law. A claim about physical uncertainty requires an evidence-based link between that law and the quotes.
For a center (p_i,p_j,\pi), admit coordinate errors (a,b,c)\geq0. The four affine cell probabilities have exact box minima \begin{aligned} \kappa_B=\min\{&\pi-c,\ p_i-\pi-a-c,\ p_j-\pi-b-c,\\ &1-p_i-p_j+\pi-a-b-c\}. \end{aligned} \tag{18} Each minimum follows by choosing the corner that minimizes its affine expression. The whole box is interior exactly when \kappa_B>0. If \kappa_B=0, the box touches the boundary. If \kappa_B<0, it contains an infeasible triple. The finite inverse requires restriction to admissible triples. A projected or regularized result reports its changed set and residual. For example, center (.6,.5,.3) and radii (.01,.02,.01) give cell minima (.29,.28,.17,.16). Center (.2,.3,.02) with common radius .03 crosses the zero-++ facet despite its positive central cells.
In the theorem, \kappa denotes this uniform minimum for a common radius \eta. The segment between admitted triples inherits the same bound. A logarithmic market scoring rule (LMSR) prices contracts through an exponential function of the outstanding inventory [2]. Its liquidity parameter b>0 sets how strongly a trade changes the price. A binary LMSR inventory change of size q changes its logit by q/b and its quote by at most q/(4b). That fact bounds a change in the quote law. Physical probability error requires a separate calibration bound.
Theorem 8.1 (Robustness of the parlay inversion).
Let (p_i, p_j, \pi_{ij}) be the true triple and (\tilde p_i, \tilde p_j, \tilde \pi_{ij}) the quoted triple with |\tilde p_i - p_i| \leq \eta, |\tilde p_j - p_j| \leq \eta, |\tilde \pi_{ij} - \pi_{ij}| \leq \eta. Suppose every triple in the L^\infty \eta-ball around the true triple has all four Fréchet functionals \{\pi,\; p_i - \pi,\; p_j - \pi,\; 1 - p_i - p_j + \pi\} at least \kappa > 0. Then the parameters inferred through the closed form (7) obey \begin{align*} |\tilde h_i - h_i| &\;\leq\; \frac{5\eta}{4\kappa}, \qquad\text{and likewise for } h_j, \tag{19} \\ |\tilde J_{ij} - J_{ij}| &\;\leq\; \frac{2\eta}{\kappa}. \tag{20} \end{align*} All three coordinates are quarter-log contrasts of the same four cells. Varying \pi moves every cell. Signed cancellation in the field derivatives gives their smaller bound.
Proof. The four Fréchet functionals are affine in (p_i, p_j, \pi), so they are at least \kappa along the whole segment between the true and quoted triples, and the mean-value inequality bounds each parameter change by \eta times the sum of the three suprema of its partial derivatives along that segment.
For h: write h_i(p_i, p_j, \pi) = \tfrac{1}{4}\bigl[\ln \pi + \ln(p_i - \pi) - \ln(p_j - \pi) - \ln(1 - p_i - p_j + \pi)\bigr]. Its \pi-derivative is a signed sum of two positive and two negative cell reciprocals, and a difference of two positive pairs is at most the larger pair, each pair being at most 2/\kappa: \begin{gather*} \Bigl|\frac{\partial h_i}{\partial \pi}\Bigr| \;=\; \frac{1}{4}\Bigl|\frac{1}{\pi} - \frac{1}{p_i - \pi} + \frac{1}{p_j - \pi} - \frac{1}{1 - p_i - p_j + \pi}\Bigr| \;\leq\; \frac{1}{2\kappa},\\ \Bigl|\frac{\partial h_i}{\partial p_i}\Bigr| \;=\; \frac{1}{4}\Bigl[\frac{1}{p_i - \pi} + \frac{1}{1 - p_i - p_j + \pi}\Bigr] \;\leq\; \frac{1}{2\kappa},\\ \Bigl|\frac{\partial h_i}{\partial p_j}\Bigr| \;=\; \frac{1}{4}\Bigl|\frac{1}{1 - p_i - p_j + \pi} - \frac{1}{p_j - \pi}\Bigr| \;\leq\; \frac{1}{4\kappa}. \end{gather*} The sum is \eta\,\bigl(\tfrac{1}{2\kappa} + \tfrac{1}{2\kappa} + \tfrac{1}{4\kappa}\bigr) = 5\eta/(4\kappa), which is (19); the argument for h_j is the same with i and j exchanged.
For J: write J(p_i, p_j, \pi) = \tfrac{1}{4}\bigl[\ln \pi + \ln(1 - p_i - p_j + \pi) - \ln(p_i - \pi) - \ln(p_j - \pi)\bigr]. All four reciprocals now enter the \pi-derivative with the same sign: \begin{gather*} \Bigl|\frac{\partial J}{\partial \pi}\Bigr| \;=\; \frac{1}{4}\Bigl[\frac{1}{\pi} + \frac{1}{1 - p_i - p_j + \pi} + \frac{1}{p_i - \pi} + \frac{1}{p_j - \pi}\Bigr] \;\leq\; \frac{1}{\kappa},\\ \Bigl|\frac{\partial J}{\partial p_i}\Bigr| \;=\; \frac{1}{4}\Bigl[\frac{1}{1 - p_i - p_j + \pi} + \frac{1}{p_i - \pi}\Bigr] \;\leq\; \frac{1}{2\kappa}, \qquad\text{and likewise}\qquad \Bigl|\frac{\partial J}{\partial p_j}\Bigr| \;\leq\; \frac{1}{2\kappa}. \end{gather*} The sum is |\tilde J_{ij} - J_{ij}| \;\leq\; \eta\Bigl(\frac{1}{\kappa} + \frac{1}{2\kappa} + \frac{1}{2\kappa}\Bigr) \;=\; \frac{2\eta}{\kappa}. \quad\square ◻
Remark 8.2 (Boundary regime).
The robustness bound degrades at the admissible-region boundary (Theorem 3.1(iii)): when \pi is close to \max(0, p_i + p_j - 1) or \min(p_i, p_j), the Lipschitz constant 1/\kappa diverges, reflecting that |J_{ij}| is going to infinity. Markets quoting boundary parlay prices are informative about the sign of J_{ij} but provide poor magnitude estimates; a declared shrinkage rule can move the quoted triple toward the independence point by a factor of 1 - \alpha for a small \alpha. This stabilizes the inverse by changing the input triple and therefore introduces bias. The resulting regularized estimate is distinct from the unregularized inverse. This is a market-microstructure design choice not pursued further here.
9 Computing Prices from a Joint Model
The preceding results concern information and uncertainty. A specified network law presents a further question: how to compute its contract probabilities within the available resources. Exact summation gives a reference value for small models. Mean field reduces computation by replacing the joint law with independent spins, which can alter joint prices even when individual prices remain accurate.
9.1 Marginals and exact computation for small M
The marginal probability of event e_k occurring is \mathbb{P}(\omega_k = +1) \;=\; \frac{\sum_{\omega: \omega_k = +1} \exp(-H(\omega))}{Z} \;=\; \frac{Z_k^+}{Z}, \tag{21} where Z_k^+ sums over configurations with \omega_k = +1.
Direct enumeration computes the partition function and all marginals in O(M^2 2^M) operations. The feasible size depends on the measured resource budget. Tree inference is exact by sum-product, while general junction-tree cost depends exponentially on graph width [36].
9.2 Mean-field approximation for larger M
Mean field approximates the joint law with a product distribution. Its magnetisation at site i is the mean spin, so a value m_i corresponds to event probability (1+m_i)/2. The approximation chooses each site’s mean consistently with the average influence of its neighbours. Its variational objective minimizes Kullback–Leibler divergence from the product distribution to the specified Ising law. This divergence measures the expected logarithmic probability ratio under the first law. Algorithm selection depends on graph structure, error requirements, and computation cost.
Definition 9.1 (Mean-field equations).
The mean-field approximation uses a product law with magnetisations \hat m_i satisfying \hat m_i \;=\; \tanh\!\left(h_i + \sum_{j \neq i} J_{ij} \hat m_j\right), \qquad i = 1, \ldots, M. \tag{22} The mean-field marginal is \hat p_i = (1 + \hat m_i)/2.
The row sum of coupling magnitudes bounds how much one mean-field update can amplify an input error. When its maximum is below one, repeated updates contract that error and determine a unique fixed point. The theorem separates this convergence guarantee from the variational bound on the normalizing constant.
Theorem 9.2 (Mean-field contraction and variational bound).
Let r_i = \sum_{j \neq i}|J_{ij}|, \qquad r = \max_i r_i, so r \leq d_{\max}J_{\max} when the interaction graph has maximum degree d_{\max} and J_{\max}=\max_{i,j}|J_{ij}|.
- (i)
-
If r<1, the mean-field map has a unique fixed point \hat m. From every m^{(0)}\in[-1,1]^M, its iterates satisfy \|m^{(n)}-\hat m\|_\infty \leq 2r^n.
- (ii)
-
If 0<r<1 and 0<\varepsilon<2, then n_\varepsilon iterations give \ell^\infty error at most \varepsilon, where n_\varepsilon =\left\lceil\frac{\log(2/\varepsilon)}{\log(1/r)}\right\rceil. Each sparse iteration costs O(M(1+d_{\max})) operations. If r=0, one iteration reaches the fixed point.
- (iii)
- For every product law q on \{-1,+1\}^M, define F_{\mathrm{MF}}(q) = \mathbb{E}_q[H]+\mathbb{E}_q[\log q]. Then F_{\mathrm{MF}}(q)\geq-\log Z.
Proof. Define \Psi:[-1,1]^M\to[-1,1]^M by \Psi_i(m)=\tanh(h_i+\sum_jJ_{ij}m_j). Since \tanh is one-Lipschitz, \|\Psi(m)-\Psi(m')\|_\infty \leq r\|m-m'\|_\infty. When r<1, Banach’s theorem gives the unique fixed point and \|m^{(n)}-\hat m\|_\infty\leq r^n\|m^{(0)}-\hat m\|_\infty\leq2r^n. Solving 2r^n\leq\varepsilon gives part (ii); evaluating the sparse matrix-vector product and the M hyperbolic tangents costs O(M(1+d_{\max})) per iteration.
For part (iii), let p(\omega)=e^{-H(\omega)}/Z. Non-negativity of D_{\mathrm{KL}}(q\|p) gives 0\leq D_{\mathrm{KL}}(q\|p) =\mathbb{E}_q[\log q]+\mathbb{E}_q[H]+\log Z =F_{\mathrm{MF}}(q)+\log Z. ◻
Convergence to the fixed point leaves the approximation error relative to the true law unresolved. The next bound controls that error in singleton means. Its proof compares the exact conditional expectation with the mean-field update by a second-order Taylor estimate.
Proposition 9.3 (Mean-field marginal error).
In the regime r<1 of Theorem 9.2, the true magnetisations m_i=\mathbb{E}_{\mathbb{P}}[\omega_i] and the mean-field fixed point \hat m satisfy \|\hat m-m\|_\infty \;\leq\; \frac{2}{3\sqrt{3}}\,\frac{r^2}{1-r}.
Proof. The exact conditional law gives m_i =\mathbb{E}_{\mathbb{P}}\!\left[ \tanh\!\left(h_i+\sum_{j\neq i}J_{ij}\omega_j\right) \right]. Set S_i=\sum_{j\neq i}J_{ij}\omega_j and \bar S_i=\mathbb{E}_{\mathbb{P}}[S_i]=\sum_{j\neq i}J_{ij}m_j. The second derivative of \tanh has supremum norm 4/(3\sqrt3). Taylor’s theorem about \bar S_i, followed by expectation, therefore gives \left|m_i-\tanh(h_i+\bar S_i)\right| \leq \frac{2}{3\sqrt3}\mathop{\mathrm{Var}}(S_i). The random variable S_i lies in an interval of length 2r_i, so \mathop{\mathrm{Var}}(S_i)\leq r_i^2. With \Psi as in the preceding proof, \|m-\hat m\|_\infty \leq \|m-\Psi(m)\|_\infty + \|\Psi(m)-\Psi(\hat m)\|_\infty \leq \frac{2r^2}{3\sqrt3}+r\|m-\hat m\|_\infty. Rearranging proves the result. ◻
The susceptibility below is the derivative of the mean-field means with respect to the fields. It measures how the approximation responds to a small field change. Here \|\cdot\|_{\mathrm{op}} is the Euclidean operator norm, and \lambda_{\max} denotes the largest eigenvalue. The formula concerns this approximation; exact environmental covariances use Proposition 4.3.
Proposition 9.4 (Linearised susceptibility).
Define \chi_{ij} = \partial \hat m_i / \partial h_j \big|_{h = 0}. For \|J\|_{\mathrm{op}} < 1, \chi \;=\; (I - J)^{-1}\big|_{\hat m = 0}, \qquad \|\chi\|_{\mathrm{op}} \;\leq\; \frac{1}{1 - \|J\|_{\mathrm{op}}}. \tag{23} Within the mean-field approximation, \|\chi\|_{\mathrm{op}} = (1 - \lambda_{\max}(J))^{-1}, which diverges as \lambda_{\max}(J) \to 1^-.
Proof. Differentiating (22) at \hat m = 0: \partial \hat m_i / \partial h_j = \delta_{ij} + \sum_k J_{ik} \, \partial \hat m_k / \partial h_j, i.e. \chi = I + J \chi, so \chi = (I - J)^{-1}. The operator-norm bound follows from the Neumann series. For the exact norm: J is symmetric with every eigenvalue \lambda_k \in (-1, 1), so \|(I - J)^{-1}\|_{\mathrm{op}} = \max_k (1 - \lambda_k)^{-1} = (1 - \lambda_{\max}(J))^{-1}. ◻
The iteration bound controls distance to the mean-field fixed point. Proposition 9.3 controls singleton approximation error within the specified Ising law. Estimation accuracy, physical calibration, and trading profit require separate evidence.
For example, take two sites with zero fields and coupling J_{12}=0.5. Mean field converges immediately from zero and returns both singleton probabilities correctly. Its product joint probability is 0.25. The exact joint probability is (1+\tanh(0.5))/4=0.3655292893. Thus exact singleton predictions can accompany a joint pricing error of 0.1155292893.
10 Numerical Evidence and Market Comparisons
The calculations below answer separate empirical questions about the specified models. They compare parameter estimates with known generating parameters, check interval coverage over repeated samples, and record cash losses under explicit trading rules. All numerical examples are synthetic. The identification and market calculations use seed 20260905, NumPy 2.4.3, and SciPy 1.17.1. The accompanying script records input laws, counts, solver residuals, and results. Probability errors and coverage refer to the stated generating law.
10.1 Identification, uncertainty, and allocation
The three-site chain has zero fields and interactions J_{02}=.4, J_{12}=.6, J_{01}=0. Its effective endpoint coupling is .2069563793. The alternative isolated-edge law matches the selected observation interface within 1.2\times10^{-16}. The smallest cell across both laws is .03588, while their full tables differ by up to .11462. Thus the negative witness uses two positive laws.
Exact full-vector likelihood fitting recovers the population endpoint interaction within 1.4\times10^{-11}. Conditional logistic fitting recovers it within 1.2\times10^{-11}. On the same 5000 independently sampled full-vector rounds, both estimates are .0154254. The pair-only estimate on those rounds is .206945 and retains its effective-coupling label. The fixture supplies no parlay quote, so the quote interface returns no estimate while aligned outcomes remain usable.
For the pair law (.45,.20,.20,.15), the true effective coupling is .1308120. Each row below uses 2000 independent training replications. The regularized estimate uses \alpha=.5. Coverage refers to the conservative coupling enclosure in Proposition 6.8. The Wald interval is the point estimate plus or minus 1.96 estimated standard errors. In Table 1, RMSE means root mean squared estimation error across replications. Enclosure coverage is the fraction of replications whose cell-derived coupling enclosure contains the generating coupling. Median width describes that enclosure’s precision.
| Rounds | Enclosure coverage | Wald coverage | Regularized RMSE | Median width |
|---|---|---|---|---|
| 104 | 1.0000 | .9515 | .10836 | .98091 |
| 520 | 1.0000 | .9455 | .04771 | .42056 |
| 2600 | 1.0000 | .9545 | .02073 | .18529 |
| 5000 | 1.0000 | .9520 | .01525 | .13319 |
At 2600 rounds, the coupling enclosure certifies positive sign in .9635 of replications. At 104 rounds, that fraction is .001. The regularized cell predictor’s expected log loss on an independent test draw, integrated exactly under the generating law, falls from 1.30217 to 1.28797 across the first and last rows. The generating law’s entropy is 1.28767. These quantities measure predictive precision under the synthetic law. A real data audit must separately establish its sample size, sampling contract, and decision utility.
The zero-count fixture (0,20,30,50) returns Boundary and a lower-unbounded coupling interval. Its declared \alpha=.5 estimate remains finite. The eight-cell saturated reconstruction has maximum cell error below 10^{-15}. Known-variance allocation returns (100,200,300) and equal Karush–Kuhn–Tucker (KKT) multiplier values for the three coordinates. These multipliers express the equality-constrained optimum of Proposition 5.4. The incompatible three-spin quotes return the analytical maximum-norm residual 17/30.
10.2 An informative window under full-law drift
Set T=32768 and generate independent categorical outcomes with laws p_s=(a_s,b_s,b_s,a_s),\qquad a_s=.4-2\times10^{-6}(T-s),\qquad b_s=.1+2\times10^{-6}(T-s). All cells are positive. The current coupling is J(p_T)=\tfrac12\log4=.693147\ldots. The cell drift bound is L=2\times10^{-6} per round. Declare four windows ending at T, with lengths 512, 2048, 8192, and 32768. Allocate \delta_W=.05/4 to each and use D_W=L(N-1)/2. Python’s random.Random with seed 20260905 supplies one uniform variate per round, assigned in cell order (++,+-,-+,--).
| Rounds | Cell counts | Cell radius | Coupling estimate | Coupling enclosure |
|---|---|---|---|---|
| 512 | (207,48,44,213) | .079947 | .759655 | [.332634,1.993488] |
| 2048 | (799,208,206,835) | .041765 | .686345 | [.458085,1.002706] |
| 8192 | (3236,882,909,3165) | .028050 | .636866 | [.485383,.819845] |
| 32768 | (12084,4345,4361,11978) | .042697 | .508307 | [.307144,.757180] |
Selecting the narrowest computed coupling enclosure chooses 8192 rounds and retains a positive lower endpoint. All four enclosures contain the current coupling in this realization. The longest window reduces sampling error but increases pooling bias, which its drift term covers. This realization demonstrates an informative output. Proposition 7.2 supplies the coverage statement under the stated law, rather than a coverage frequency from one draw. The interval estimates the current effective coupling. Direct-edge identification and physical calibration retain their separate observation requirements.
10.3 Environmental and computational checks
Small cycle environments have fields linearly spaced from -.2 to .2 and equal edge coupling J. Exact enumeration gives the following covariance row sums and their influence bounds.
| Sites | J | \alpha | \chi_{\mathrm{env}} | \alpha/(1-\alpha) |
|---|---|---|---|---|
| 3 | .12 | .23885 | .25787 | .31381 |
| 5 | .18 | .35616 | .42350 | .55319 |
| 6 | .12 | .23885 | .26661 | .31381 |
Enumeration checks these examples. The coupling argument in Proposition 4.3 establishes the general constant.
A six-site path uses the field vector and successive edge couplings h=(.15,-.1,.2,-.15,.1,-.05),\qquad J_E=(.22,-.18,.20,.16,-.21). The model determines all reference inputs. The reconstruction uses its exact singleton and tree-edge marginals. Each resulting law prices the same six singleton and fifteen pair statistics. Exact enumeration and known-tree reconstruction agree within 3\times10^{-16}. The independent law with exact singleton marginals has maximum pair-moment error .209898. Mean field has maximum singleton-mean error .013955 and pair-moment error .211517. Its fixed-point iteration takes 25 steps to a maximum coordinate change below 10^{-12}. Here r=.4, so the contraction guarantee holds despite the pair error. The 64-state reference makes these errors directly observable. It does not establish a universal dimension cutoff or speed ranking.
10.4 Market mechanisms and loss
An accurate joint distribution does not by itself determine a market maker’s monetary loss. That loss also depends on how requests change prices, what shares are issued, and what premiums are collected. This subsection compares those execution choices on common synthetic inputs. An automated market maker is the trading rule that quotes prices and issues the requested contracts. We use binary indicators X_A,X_B\in\{0,1\} here, with joint cell order 00,01,10,11. This is the same outcome space as the spin representation, with the zero outcome corresponding to spin -1. A primitive is one executed sub-trade within a request.
ParlayMarket uses trade-target gradients to update a shared Ising law [3]. Its learning and expected-loss bounds impose additional local curvature, gradient-noise, and sensitivity conditions. The Automated Parlay Market Maker (APMM) maintains separate LMSR books for subsets of events [37]. Each book records its own contract inventory and prices. APMM propagates inventory updates from smaller subsets into their containing books. Its mixed YES/NO contract universe is larger than ParlayMarket’s basic positive-conjunction universe. APMM’s loss calculation includes charged premiums, issued holdings, and premium-free propagation. Neither mechanism’s learning conditions establish physical calibration of observed prices.
A price update determines a monetary loss only after execution specifies the shares issued and the premium collected. The following experiment records both quantities through settlement. Each arrival supplies one Bernoulli target for A or B. The executed contracts are A, its complement, B, and its complement, with unit settlement payouts. Both mechanisms natively accept this common interface. An APMM pair request supplies a complete joint target and causes several sub-trades. A scalar ParlayMarket request supplies different information. The first comparison executes singleton requests and reports pair-book prices as diagnostics. A second comparison below supplies complete pair information through explicit adapters.
An arrival is one trader request with a stated probability target and purchase budget. The maker receives a premium when it issues shares and pays their specified unit amount when the outcome settles. All runs use liquidity b=10, initial maker cash K=100, a budget of 100 per arrival, and zero fees. Every mechanism receives the same prior, target sequence, and declared settlement law within each fixture. Cash covers every possible issued payout after each accepted request. A request exceeding its trader budget or this cash constraint leaves prices, cash, and holdings unchanged.
For a singleton quote q and target \tau, define the share difference and minimum buy-only bundle by d=b(\mathop{\mathrm{logit}}\tau-\mathop{\mathrm{logit}}q),\qquad h_1=\max(d,0),\qquad h_0=\max(-d,0). The trader receives h_0,h_1 shares and pays c=b\log\bigl((1-q)e^{h_0/b}+qe^{h_1/b}\bigr). This cost difference is specified in ParlayMarket [3, Appendix E.4] and APMM [37, Eq. 8]. Adding equal shares to both outcomes adds the same amount to premium and payout, so the minimum choice preserves net loss. Let T contain the 0/1 singleton and pair products of ParlayMarket’s exponential family. Let \phi denote the coefficients of this 0/1 exponential family. In the update below, P_\phi denotes the replay’s pricing law, rather than the physical outcome law P used earlier. The ParlayMarket replay (PM replay) takes one step of size \eta=.2 using the pre-trade law: \phi'=\phi-\eta\frac{q-\tau}{q(1-q)}\operatorname{Cov}_{P_\phi}(T(X),X_i). Exact enumeration supplies the expectation in its trade-target update [3, Eq. 13 and Appendix E.4]. A purchase gauge chooses how many equal shares to add to both outcomes without changing their net payoff difference. The minimum gauge adds none beyond the required nonnegative bundle. The PM replay implements Appendix E.4 with exact inference, this gauge, and unpaid internal shadow resets. A shadow reset changes the internal pricing state without a corresponding external trade. These resets issue zero external shares. Its reported losses refer to this declared replay realization.
APMM adds the real bundle to the selected singleton block. The resulting change propagates into containing books, with zero extra premium and zero extra issued holdings [37, Sections 3.2–3.3]. The exact joint-LMSR reference applies the same real bundle to its four outcome inventories and prices from their normalized exponential weights. For every mechanism, let D_t(x) be all real shares payable at outcome x through arrival t. With collected premiums c_s and fees f_s, define L_t(x)=D_t(x)-\sum_{s\leq t}(c_s+f_s),\qquad R=\max_{t,x}\max(L_t(x),0). The terminal maker cash is K-L_T(x), and R is the reserve required by that particular path. All four outcomes are enumerated. The expected loss uses the explicitly supplied synthetic law.
The independent fixture starts at the uniform law and settles under (.06,.14,.24,.56), ordered 00,01,10,11. It repeats the arrivals (A,.8),(B,.7) thirty times. Every method reduces maximum singleton error from .3 to at most .034747. The correlated fixture starts at (.4,.1,.1,.4) and settles under (.16,.04,.16,.64). It receives (A,.8),(B,.68) once. The repeated-information fixture gives that same pair of targets four times, with the same prior and settlement law. Table 4 summarizes complete ledgers for these paths.
| Fixture | Mechanism | Error | \mathbb{E}L_T | \max_x L_T | R |
|---|---|---|---|---|---|
| Independent | Joint LMSR | .000000 | 2.75028 | 8.06476 | 8.06476 |
| PM replay | .034747 | 18.28019 | 76.81987 | 76.90263 | |
| APMM | .000000 | 2.75028 | 8.06476 | 8.06476 | |
| Correlated | Joint LMSR | .000000 | 1.92745 | 4.70004 | 4.70004 |
| PM replay | .255665 | 2.40579 | 7.23517 | 7.23517 | |
| APMM | .000000 | 2.59022 | 7.77488 | 7.77488 | |
| Repeated | Joint LMSR | .000000 | 1.92745 | 4.70004 | 4.70004 |
| PM replay | .165536 | 6.22167 | 20.90693 | 20.90693 | |
| APMM | .000000 | 2.59022 | 7.77488 | 7.77488 |
After the correlated fixture’s first arrival, APMM’s joint table is (.16,.04,.16,.64), while its untouched singleton B book quotes .5. The resulting book disagreement is .18. After the second arrival, both singleton books equal their supplied targets, but the joint’s maximum marginal disagreement is .138697. Repeated targets preserve that disagreement. ParlayMarket’s quotes come from one law throughout, while its gradient step continues to leave a target discrepancy. Repeated information therefore produces additional real trades and loss in this declared execution rule. These finite results identify effects of the update and execution choices. They supply no parameter-independent ranking or general loss bound for either mechanism.
A separate APMM fixture executes a full pair target (.16,.04,.16,.64) from the uniform prior. Its three sub-trades run bottom-up through A, B, and the pair. Each premium uses the current book after lower-order propagation, then charges only its own block change. Their premiums are 9.162907, 4.462871, and 12.237754. The total issued payouts are (21.400662,7.537718,21.400662,35.263605). Expected loss is 3.854895, and maximum loss is 9.400073. A budget sufficient only for the first sub-trade rejects the full request without changing any block. This checks complete pair accounting within APMM. The following construction matches the information supplied to the two interfaces.
The supplementary scripts market_ledgers.py and test_market_ledgers.py record the complete arrivals, premiums, fees, issued payouts, and prefix reserves. Fourteen checks cover cash conservation, entropy loss, gradient differentiation, propagation, pair routing, fees, and atomic rejection. The estimation results of this paper remain separate from these execution mechanisms.
10.4.1 Matching complete pair information
The previous experiment gives each method the same singleton targets. A pair comparison must also equalize the joint information supplied with each request. The next construction converts one complete pair table into three scalar targets and records all resulting trades. A positive pair law contains three independent probabilities. Write its cells as \tau=(\tau_{00},\tau_{01},\tau_{10},\tau_{11}) and define \alpha=\tau_{10}+\tau_{11},\qquad \beta=\tau_{01}+\tau_{11},\qquad \gamma=\tau_{11}. These are the target prices of A, B, and A\cap B. The inverse is \tau=(1-\alpha-\beta+\gamma,\ \beta-\gamma,\ \alpha-\gamma,\ \gamma). Its positive domain is \max(0,\alpha+\beta-1)<\gamma<\min(\alpha,\beta). Thus a complete pair target and these three scalar targets carry exactly the same information. This equivalence supplies a comparison without assigning extra information to one mechanism.
Each request now carries the entire positive table. The PM adapter presents its three scalar targets in the fixed order A,B,A\cap B. Each step uses the preceding staged pricing law in the gradient and binary cost formula. The declared Appendix E.4 replay buys the charged YES or NO bundle for each contract. For the conjunction, NO pays 1-X_AX_B, rather than the mixed outcome (1-X_A)(1-X_B). APMM performs its native full-pair request through the three subset books. The joint LMSR moves its four-outcome table directly to \tau. The three mechanisms retain these different native portfolios and primitive counts. All payouts are recorded on the same four terminal outcomes.
The adapter stages every primitive before committing the request. One budget covers all premiums and fees, and a persistent trader balance funds that budget. Maker cash must cover aggregate issued payouts at every staged prefix. Failure preserves the prior pricing state, cash, and holdings. Replaying the same command returns its recorded outcome without issuing further shares. A new command carrying the same target is a new trading request. The exact probability codec uses rational arithmetic. Cash accounting treats every finite quoted premium and share quantity as its exact binary rational value. Pricing itself uses binary64 enumeration. This separation permits exact cash conservation while retaining the numerical limits of the price calculation.
The ledgers allow a direct comparison of issued payouts after deducting collected premiums. The next identity isolates the effect of propagating a lower-order price change without an additional premium. Its normalizing factor records the difference between this execution and one trade in a joint book. Let p_0 be a positive initial pair law, and let p_{0,A},p_{0,B} denote its marginals. The APMM singleton books initially equal those marginals. Let \tau_A,\tau_B be the target marginals, and define f_A(x_A)=\frac{\tau_A(x_A)}{p_{0,A}(x_A)},\qquad f_B(x_B)=\frac{\tau_B(x_B)}{p_{0,B}(x_B)},\qquad Z=\sum_xp_0(x)f_A(x_A)f_B(x_B).
Proposition 10.1 (Loss of a full pair request).
Suppose both mechanisms accept the complete target \tau, with common liquidity b and zero fees. Let \ell(x) be the new issued payout minus the premium collected for this request. For exact joint LMSR and the bottom-up APMM realization, \ell_{\rm joint}(x)=b\log\frac{\tau(x)}{p_0(x)},\qquad \ell_{\rm APMM}(x)=\ell_{\rm joint}(x)+b\log Z. If p_0=p_{0,A}p_{0,B}, then Z=1 and the net payouts agree at every outcome.
Proof. A cost-priced move from q to r has net payout b\log(r(x)/q(x)). A uniform buy gauge increases premium and payout equally. The two singleton trades therefore contribute b\log f_A and b\log f_B. Their unpaid propagation changes the pair quote to p_2=p_0f_Af_B/Z. The final pair trade contributes b\log(\tau/p_2). Adding the three contributions gives the identity. For a product prior, the two marginal expectations of f_A and f_B are each one, so Z=1. ◻
The correction can have either sign. Here \kappa denotes the prior covariance, a local use distinct from the positive cell margin in Section 8. Writing u=p_{0,A}(1), v=p_{0,B}(1), and \kappa=p_0(11)-uv gives Z=1+\frac{\kappa(\alpha-u)(\beta-v)}{u(1-u)v(1-v)}. Its sign relative to one therefore depends on the initial covariance and the two marginal changes. For prior (.4,.1,.1,.4) and target (.16,.04,.16,.64), it is b\log(1.1296). For prior (.1,.4,.4,.1) and the same target, it is b\log(.8704). At b=10, these change every terminal net payout by 1.218636 and -1.388024, respectively. Fees subtract their actual charged totals from these net payouts.
Corollary 10.2 (A quote-restoring payout cycle).
Under the proposition’s routing and zero-fee assumptions, take p=(.4,.1,.1,.4) and \tau=(.16,.04,.16,.64). If both requests execute, the cycle p\to\tau\to p restores every APMM book quote and pays b\log\frac{353}{272}>0 net to the trader at every outcome. The direct joint-LMSR cycle has zero net payout.
Proof. Each full target restores coherence between the pair and singleton books. The two normalization factors are 706/625 and 625/544. The two log target-ratio terms cancel, leaving b\log(353/272). The joint-LMSR net payouts telescope without these normalization terms. ◻
At b=10, the cycle pays 2.606660 in every state. The ledger reproduces this payout and retains it as a settlement liability. Finite maker capital eventually stops further requests, while the accrued liability remains payable. The result concerns the stated own-book charging and unpaid propagation rule. A mechanism that changes that rule requires its own cash and payout calculation. Thus coherent final quotes alone do not establish a bounded-loss execution mechanism.
Equal net payouts also permit different gross funding needs because the buy gauges can differ. From the uniform prior, the direct joint request costs 18.325815, while the APMM request costs 25.863533. Their payout difference is the same constant 7.537718 in every state. A common finite budget must therefore be checked against the actual premium, even when net losses agree.
Table 5 supplies the same target (.16,.04,.16,.64) from the uniform prior. All runs use b=10, maker cash 100, trader cash 10{,}000, zero fees, and a 100 budget per complete request. The PM step size remains .2. Each request supplies three independent probabilities, whether processed in one primitive or three.
| Requests | Mechanism | Primitives | Cell error | \mathbb{E}L_T | \max_xL_T |
|---|---|---|---|---|---|
| 1 | Joint LMSR | 1 | .000000 | 3.85490 | 9.40007 |
| PM replay | 3 | .340346 | 5.51235 | 16.27121 | |
| APMM | 3 | .000000 | 3.85490 | 9.40007 | |
| 60 | Joint LMSR | 60 | .000000 | 3.85490 | 9.40007 |
| PM replay | 180 | .076001 | 23.97888 | 84.30032 | |
| APMM | 180 | .000000 | 3.85490 | 9.40007 |
All three methods improve the joint estimate while maintaining funded settlement in these runs. The PM adapter’s terminal Kullback–Leibler divergence from the target table to its quoted table falls from .385490 initially to .296983 after one request and .062447 after sixty. Its direction is D_{\mathrm{KL}}(\tau\|p), with p the terminal pricing law. Its repeated gradient updates continue to produce paid trades. The other two mechanisms reach the full target in their first request. The six PM primitive orders give distinct final laws, with one-request divergence between .296388 and .296993. The reported order is fixed before evaluation.
This experiment compares specified responses to common information and funding. Equal nominal b does not imply equal curvature or equal native loss subsidy. The PM result concerns one ordered gradient pass per request, rather than an optimized allocation of computation across learning rates and iterations. Method-wide performance depends on those choices, information arrival, and funding. The supplementary pair_adapter.py records the full requests, staged trades, and exact cash accounts. Eighteen additional checks cover the codec, pair gradients, funding, replay, conservation, information use, and the loss identity. They include 165 rational pair laws, 40 statewise loss comparisons, and 60 funded random-path requests across the three mechanisms.
For a common exact execution baseline, let all 64 complete-outcome contracts have prior probabilities q_x>0 and liquidity b=10. Use the cost C(v)=b\log\sum_xq_xe^{v_x/b},\qquad C(0)=0. Its gradient is a normalized positive pricing law. Since C(v)\geq v_x+b\log q_x, the loss at any settlement is bounded by b\log(1/\min_xq_x). The inventory vector v records shares outstanding in each complete-outcome contract. A common seeded path buys 100 contracts in quantities .5, 1, or 2. Uniform, independent, and Ising priors use the same path, contract universe, liquidity, and unit payouts. Their largest settlement losses on that path are 4.42924, 4.51896, and 4.60225. Their all-inventory loss bounds are 41.58883, 48.56959, and 55.59163. Expected profit under the specified synthetic law is .32678, .23706, and .15378. These are outputs of this exact cost-function mechanism. They are separate from ParlayMarket and APMM ledger benchmarks. In this example, greater model accuracy does not yield greater maker profit.
A comparison on real markets additionally requires held-out calibration, observed trade behavior, and realized settlement losses. The synthetic results establish finite computational and estimation witnesses, with those empirical questions retained explicitly.
11 Discussion
11.1 Relation to prior Ising-based models of financial markets
The econophysics literature [7, 8, 9, 10, 11] applies Ising-type models to equity markets with J_{ij} modelling trader-imitation strength and \omega_i modelling a trader’s decision. The present model is different: spins are event outcomes (binary resolutions of real-world events), and J_{ij} is a pairwise interaction term in the joint distribution over outcomes, identified from the specified joint probabilities or estimated from aligned resolved outcomes. The trader-interaction reading of J is not claimed.
The Ising structure-learning literature provides complementary results. Ravikumar, Wainwright, and Lafferty [15] prove sample-complexity bounds for recovering the sparsity pattern of J from samples of the full spin vector via \ell_1-regularised logistic regression. Santhanam and Wainwright [24] give the matching information-theoretic lower bound n = \Omega(c^d \log p) for structure recovery, and Bresler [25] gives the first \mathrm{poly}(p)-time algorithm without the incoherence condition. Here n counts full-vector samples, p the sites, and d the maximum degree. The constant c in the cited lower bound depends on its interaction class. Vuffray et al. [16] give a sample-optimal interaction-screening estimator. All of these operate on samples of the full \omega-vector and target the full adjacency under sparsity regularisation. The parlay-MLE of Section 6.1 operates on samples of the pair-marginal of \omega and targets one effective pair coupling with unpenalized estimation; the sample-complexity regime is different (parametric \epsilon^{-2}\log(1/\delta) per pair rather than combinatorial scaling in the full-adjacency support). Combining the two, by using market-derived pair-level moment estimates as side information for the \ell_1 or interaction-screening estimator, is a natural direction.
On what market prices identify probabilistically: Manski [26] showed that under heterogeneous beliefs or risk aversion, a prediction-market price estimates a functional of the cross-trader belief distribution rather than the objective probability. A risk-neutral scoring rule represents the participating beliefs under its assumptions. Physical calibration requires independent resolved outcomes.
11.2 Relation to combinatorial prediction markets
Hanson’s combinatorial information markets posed the joint-outcome pricing problem [1]. Chen and Pennock [13] developed a utility framework for bounded-loss market makers on combinatorial outcome spaces, and Abernethy, Chen, and Wortman Vaughan [14] showed that automated market-making on combinatorial markets admits a convex-optimization formulation. A central design question in that literature is the choice of prior over joint outcomes; the standard LMSR cost function implicitly uses a uniform prior. The Ising parameterisation (3) provides an alternative, namely a prior with pairwise moment constraints, and the MLE of Theorem 6.3 gives a procedure for estimating its effective pair parameters from aligned resolved outcomes. Whether an LMSR with an Ising-parameter prior strictly dominates the uniform-prior LMSR on specific classes of correlated-event markets is an open question.
The Claim as Primitive defines the cost-function pricing interface that consumes these joint probabilities [32]. One-Way Coupling of Prediction Markets to Automated Market Makers defines the incentive-compatible channel that can move an automated-market-maker price from a prediction signal [33].
11.3 Open problem: dynamic identification from lead-lag data
The parlay-based estimator requires joint observations for each pair (i,j). Marginal price paths provide a different type of observation. The static Ising law specifies the distribution of resolved outcomes, but it does not specify a stochastic process for pre-resolution prices or returns. A Granger statistic measures the additional predictive contribution of one variable’s past, conditional on the included past variables. The contemporaneous R^2 below instead measures the fraction of variation explained by a regression within the same round.
The zero-field two-site law gives the exact contemporaneous normalization \begin{aligned} \rho_{ij}:=\mathop{\mathrm{Corr}}(\omega_i,\omega_j)&=\tanh J_{ij}, &J_{ij}&=\operatorname{arctanh}\rho_{ij},\\ R_{\mathrm{contemp}}^2&=\rho_{ij}^2=\tanh^2 J_{ij} =J_{ij}^2+O(J_{ij}^4). \end{aligned} \tag{24} Indeed, Z=4\cosh J_{ij}, symmetry gives zero means and unit variances, and \partial_{J_{ij}}\log Z=\mathbb{E}[\omega_i\omega_j]. The one-regressor contemporaneous identity R^2=\mathop{\mathrm{Corr}}^2 and the Taylor expansion then give (24). This identity is contemporaneous. It does not identify J_{ij} from a lagged Granger statistic. For example, let the vectors \omega_t be independent across time, with each vector drawn from the same two-site law and J_{ij}\neq 0. Equation (24) still holds within each round, while every population lagged-Granger partial-R^2 is zero.
A lead-lag estimator therefore requires an explicit transition law. One possible starting point is the synchronous kinetic Ising model \mathbb{P}(\omega_{t+1}\mid\omega_t) =\prod_j \frac{\exp\!\left\{\omega_{j,t+1}\left(a_j+\sum_i K_{ij}\omega_{i,t}\right)\right\}} {2\cosh\!\left(a_j+\sum_i K_{ij}\omega_{i,t}\right)}. \tag{25} If this transition law holds and the lag vector has full support, its conditional probabilities identify the directed coefficients K_{ij}. Consistent estimation also requires a stationary or explicitly modelled non-stationary sampling law. Relating K_{ij} to the static symmetric couplings J_{ij} requires an additional model. Deriving that relation for marginal-market price paths remains open.
11.4 Limitations
The results are stated under explicit modelling assumptions:
- (L1)
-
Pairwise interactions only. The Ising Hamiltonian (2) omits genuine higher-order log-linear interactions. Event triples with joint dependence but pairwise independence (e.g., bracket-conditional tournaments) are misspecified.
- (L2)
-
Stationarity. Theorems 6.3 and 6.4 assume \theta^{\mathrm{eff}} is constant over the observation window. Section 7 estimates a current conditional pair law with separate sampling and full-law drift bounds. Establishing those drift bounds requires a declared model or independent simultaneous evidence.
- (L3)
-
Symmetric J_{ij} = J_{ji}. The Ising parameterisation is undirected. Asymmetric causal relationships cannot be represented; a directed graphical-model alternative is more expressive and lacks the simple closed-form identifiability of Theorem 3.1.
- (L4)
-
LMSR-to-Ising bridge. The inversion (7) assumes the quoted marginals and the quoted parlay price equal the corresponding Boltzmann probabilities. Section 8 propagates a declared quote-error set. Scoring-rule equilibrium alone supplies no physical calibration certificate. A full model of the price-formation process mapping the Boltzmann distribution to observed order-book prices is not developed here.
- (L5)
-
Admissible-region boundary. When a parlay price lies near the Fréchet-Hoeffding boundary, the closed-form estimator is unstable and the MLE has large variance (Remark 8.2). Section 6.3 supplies boundary output and uncertainty sets.
- (L6)
-
Mean-field convergence condition. Theorem 9.2 requires r<1. This condition can fail in dense networks with strong coupling; exact enumeration and structured inference remain available according to their resource and graph-width requirements.
- (L7)
-
Survivorship bias. Quote availability and outcome availability are recorded separately. Selective listing or missing resolutions can alter either target population. Inclusion in outcome inference depends on aligned observations and the declared sampling law.
- (L8)
-
Causal sufficiency. Two events that co-move through an unobserved common factor produce nonzero \hat J_{ij} with zero direct interaction. Effective coupling records the observed pair law. Direct-interaction interpretation requires the additional network model and observations in Section 5.
- (L9)
-
One-shot pairs. Definition 6.2 counts settled rounds, not contracts, so a pair of event types that resolves once supplies N_p = 1 whatever the open interest, and Theorems 6.3 and 6.4 give no useful precision guarantee from it. Proposition 6.8 still returns a defined uncertainty set. The estimation results apply to pair types that recur, such as per-fixture, per-round or per-day markets. Pooling structurally similar pairs into one type to raise N_p requires the common-law and independence conditions of Assumption 6.1(R2). Those conditions are data-model assumptions.
12 Appendix: Proofs
The main sampling assumptions appear before the likelihood in Section 6. This appendix supplies the exponential-family facts used in its proofs and the matrix identity behind Woolf’s variance. The final example shows how joint estimation of the fields increases the coupling’s attainable variance.
12.1 Notation and regularity
Fix a pair (i, j) and suppress indices: \omega = (\omega_i, \omega_j), \theta = (h_i, h_j, J_{ij}) \in \Theta \subset \mathbb{R}^3. Let \mathbb{P}_\theta denote the two-site Ising distribution with density p_\theta(\omega) \;=\; \exp\!\bigl(h_i \omega_i + h_j \omega_j + J_{ij} \omega_i \omega_j - A(\theta)\bigr), \qquad A(\theta) = \ln\!\sum_\omega \exp(\cdot). The sufficient statistics are T(\omega) = (\omega_i, \omega_j, \omega_i \omega_j) \in \{-1, +1\}^3, and the model is a minimal regular exponential family on the 4-point support \{-1, +1\}^2 with three-dimensional natural parameter space \Theta = \mathbb{R}^3 (full-rank, since the three sufficient statistics are affinely independent on the support) [18, Ch. 1][19, Ch. 2]. Minimality, together with \mathbb{R}^3 being the full domain of finiteness of the log-partition function A, places the model in the canonical regular-exponential-family setting of [18, Thm. 3.6]. The mean-parameter map \theta \mapsto \mu(\theta) := \mathbb{E}_\theta T is consequently a C^\infty diffeomorphism from \mathbb{R}^3 onto the interior of the convex hull of T(\{-1, +1\}^2) \subset \mathbb{R}^3, with Jacobian d\mu / d\theta = I(\theta) = \mathop{\mathrm{Cov}}_\theta(T) \succ 0 everywhere in \Theta. Injectivity of \theta \mapsto \mathbb{P}_\theta on \Theta follows immediately.
Assumption 6.1 states the sampling and target conditions used throughout the fixed-law estimates.
Identifiability also follows from strict convexity of the log-partition function A on the interior of \Theta [18, Thm. 1.13]: distinct natural parameters give distinct mean parameters \mathbb{E}_\theta T.
12.2 Schur complement and the Woolf identity
Lemma 12.1 (Schur complement).
Let I = \begin{pmatrix} A & b \\ b^\top & d \end{pmatrix} \in \mathbb{R}^{3 \times 3} be positive-definite with A \in \mathbb{R}^{2 \times 2} and d > 0. Then: (i) the Schur complement S := d - b^\top A^{-1} b > 0; (ii) [I^{-1}]_{33} = S^{-1}; (iii) S \leq d with equality iff b = 0.
Proof. (i) Standard property of the Schur complement of a principal block of a positive-definite matrix [21, Thm. 7.7.6]. (ii) Block-inverse formula [21, §0.8]. (iii) A^{-1} \succ 0 gives b^\top A^{-1} b \geq 0 with equality iff b = 0. ◻
Writing the cells as p_{s_i s_j} = \tfrac{1}{4}(1 + m_i^* s_i + m_j^* s_j + C_{ij}^* s_i s_j), expanding S = I_{33} - b^\top A^{-1} b in (m_i^*, m_j^*, C_{ij}^*), and clearing denominators verifies the identity S^{-1} \;=\; \frac{1}{16}\sum_{s_i, s_j \in \{\pm 1\}} \frac{1}{p_{s_i s_j}} used in Theorem 6.3(iii). The identity is also forced, with no computation, by the delta-method derivation in that proof: both expressions are the asymptotic variance of the same estimator.
Remark 12.2 (Schur correction: numerical illustration).
For m_i^* = m_j^* = 0.3, C_{ij}^* = 0.2: b = (0.3 - 0.06, 0.3 - 0.06)^\top = (0.24, 0.24)^\top, A = \begin{pmatrix} 0.91 & 0.11 \\ 0.11 & 0.91 \end{pmatrix}, A^{-1} \approx \begin{pmatrix} 1.115 & -0.135 \\ -0.135 & 1.115 \end{pmatrix}, giving b^\top A^{-1} b \approx 0.0576 \cdot 1.96 \approx 0.113.
Marginal variance: 1/((1 - 0.04)N_p) \approx 1.042/N_p. Schur-corrected efficient variance: 1/((0.96 - 0.113)N_p) = 1/(0.847N_p) \approx 1.181/N_p. The correction raises the efficient variance by 13\% in this regime, confirming that the marginal bound overstates the attainable efficiency when (m_i^*, m_j^*) \neq (0,0).
References
[1] R. Hanson. Combinatorial information market design. Information Systems Frontiers, 5(1):107-119, 2003.
[2] R. Hanson. Logarithmic market scoring rules for modular combinatorial information aggregation. Journal of Prediction Markets, 1(1):3-15, 2007.
[3] R. Rana, V. Nadkarni, N. Moshrefi, and P. Viswanath. ParlayMarket: Automated market making for parlay-style joint contracts. arXiv:2603.22596v3, 10 August 2026. https://arxiv.org/abs/2603.22596v3.
[4] B. Woolf. On estimating the relation between blood group and disease. Annals of Human Genetics, 19(4):251-253, 1955.
[5] E. Ising. Beitrag zur Theorie des Ferromagnetismus. Zeitschrift für Physik, 31:253-258, 1925.
[6] S. Haykin. Adaptive Filter Theory. Pearson, 5th edition, 2014.
[7] R. Cont and J.-P. Bouchaud. Herd behavior and aggregate fluctuations in financial markets. Macroeconomic Dynamics, 4(2):170-196, 2000.
[8] J.-P. Bouchaud and R. Cont. A Langevin approach to stock market fluctuations and crashes. European Physical Journal B, 6(4):543-550, 1998.
[9] D. Sornette. Why Stock Markets Crash: Critical Events in Complex Financial Systems. Princeton University Press, 2003.
[10] T. Kaizoji. Speculative bubbles and crashes in stock markets: an interaction model of speculative activity. Physica A, 287(3-4):493-506, 2000.
[11] T. Kaizoji. An interaction model of financial markets from the viewpoint of nonextensive statistical mechanics. Physica A, 370(1):109-113, 2006.
[12] M. Mézard, G. Parisi, and M. A. Virasoro. Spin Glass Theory and Beyond. World Scientific Lecture Notes in Physics, Vol. 9, 1987.
[13] Y. Chen and D. M. Pennock. A utility framework for bounded-loss market makers. In Proceedings of UAI 2007, pp. 49-56, 2007.
[14] J. Abernethy, Y. Chen, and J. Wortman Vaughan. Efficient market making via convex optimization, and a connection to online learning. ACM Transactions on Economics and Computation, 1(2):12:1-12:39, 2013.
[15] P. Ravikumar, M. J. Wainwright, and J. D. Lafferty. High-dimensional Ising model selection using \ell_1-regularized logistic regression. Annals of Statistics, 38(3):1287-1319, 2010.
[16] M. Vuffray, S. Misra, A. Lokhov, and M. Chertkov. Interaction screening: Efficient and sample-optimal learning of Ising models. In Advances in Neural Information Processing Systems (NeurIPS), 2016.
[17] E. L. Lehmann and G. Casella. Theory of Point Estimation. Springer Texts in Statistics, 2nd edition, 1998.
[18] L. D. Brown. Fundamentals of Statistical Exponential Families. IMS Lecture Notes Monograph Series 9, 1986.
[19] B. Efron. Exponential Families in Theory and Practice. Cambridge University Press, 2022.
[20] A. B. Tsybakov. Introduction to Nonparametric Estimation. Springer Series in Statistics, Springer, 2009.
[21] R. A. Horn and C. R. Johnson. Matrix Analysis. Cambridge University Press, 2nd edition, 2013.
[22] J. Besag. Spatial interaction and the statistical analysis of lattice systems. Journal of the Royal Statistical Society, Series B, 36(2):192-236, 1974.
[23] D. H. Ackley, G. E. Hinton, and T. J. Sejnowski. A learning algorithm for Boltzmann machines. Cognitive Science, 9(1):147-169, 1985.
[24] N. P. Santhanam and M. J. Wainwright. Information-theoretic limits of selecting binary graphical models in high dimensions. IEEE Transactions on Information Theory, 58(7):4117-4134, 2012.
[25] G. Bresler. Efficiently learning Ising models on arbitrary graphs. In Proceedings of the 47th ACM Symposium on Theory of Computing (STOC), pp. 771-782, 2015.
[26] C. F. Manski. Interpreting the predictions of prediction markets. Economics Letters, 91(3):425-429, 2006.
[27] T. Plefka. Convergence condition of the TAP equation for the infinite-ranged Ising spin glass model. Journal of Physics A: Mathematical and General, 15(6):1971-1978, 1982.
[28] H. J. Kappen and F. B. Rodrı́guez. Efficient learning in Boltzmann machines using linear response theory. Neural Computation, 10(5):1137-1156, 1998.
[29] Y. M. M. Bishop, S. E. Fienberg, and P. W. Holland. Discrete Multivariate Analysis: Theory and Practice. MIT Press, 1975.
[30] A. Agresti. Categorical Data Analysis. Wiley, 3rd edition, 2013.
[31] T. Richardson and P. Spirtes. Ancestral graph Markov models. Annals of Statistics, 30(4):962-1030, 2002.
[32] R. Lorgat. The Claim as Primitive. Companion paper, 2026.
[33] R. Lorgat. One-Way Coupling of Prediction Markets to Automated Market Makers. Companion paper, 2026.
[34] C. J. Clopper and E. S. Pearson. The use of confidence or fiducial limits illustrated in the case of the binomial. Biometrika, 26(4):404–413, 1934. https://doi.org/10.1093/biomet/26.4.404.
[35] P. Rebeschini and R. van Handel. Comparison theorems for Gibbs measures. Journal of Statistical Physics, 157:234–281, 2014. https://doi.org/10.1007/s10955-014-1087-7.
[36] M. J. Wainwright and M. I. Jordan. Graphical models, exponential families, and variational inference. Foundations and Trends in Machine Learning, 1(1–2):1–305, 2008.
https://doi.org/10.1561/2200000001.
[37] N. Moshrefi, R. Rana, and P. Viswanath. APMM: Automated parlay market maker. arXiv:2607.18299v2, 26 August 2026. https://arxiv.org/abs/2607.18299v2.
[38] W. Hoeffding. Probability inequalities for sums of bounded random variables. Journal of the American Statistical Association, 58(301):13–30, 1963.
[39] S. R. Howard, A. Ramdas, J. McAuliffe, and J. Sekhon. Time-uniform, nonparametric, nonasymptotic confidence sequences. Annals of Statistics, 49(2):1055–1080, 2021.
Accompanying computations.
The archive market-ledgers-supplement.zip contains the market ledgers, pair adapter, synthetic experiments, and drift and certified-arithmetic fixtures. Its README.md gives reproduction commands, dependency versions, comparison rules, and the scope of each computation. The archive’s SHA-256 digest is
a8a6ea2c1212bbb5f70b32ff0d52f964eec86fddc893239967dc4332a478c489.