AI Research and Engineering

Evidence uncertainty and decision guarantees in investment research

Project 4: AI-Driven Investment Research and Explainable Risk Analysis
Author: Theodore Ouyang

An exact explanation of a prediction can still support an incorrect comparison between companies. Feature attribution describes how a model uses its current inputs. Investment research must also establish how those inputs might change after verification, which other models fit the validation data, and whether either source of uncertainty could reverse the conclusion. In the project, I carried evidence corrections through to scores, predictions, rankings, and explanations. The analysis below develops that practice into a decision protocol: represent unresolved evidence jointly, identify comparisons that remain valid across the resulting set, and allocate verification effort according to the decision regret it can remove.

1 Where existing methods leave the decision unresolved

SHAP's local accuracy property decomposes a given prediction into feature contributions. It imposes no requirement that the input facts be correct. Conditional Shapley methods address dependence among features, while sets of explanations across near-optimal models address uncertainty in model selection. These methods establish properties of attribution and its reliability. A further step is needed to determine which correction to the evidence could change the current research decision. See SHAP, Aas et al., and Marx et al..

For a linear model under a fixed marginal-background definition, the attribution to feature j j is ϕ j = β j ( x j − E Q X j ) \phi_j=\beta_j(x_j-\mathbb E_QX_j) . If its unresolved input error satisfies | δ j | ≤ d j |\delta_j|\leq d_j , the corresponding sensitivity to correction is | β j | d j |\beta_j|d_j . Over a box of admissible errors, the exact worst-case change in the prediction is

sup | δ j | ≤ d j | f ( x + δ ) − f ( x ) | = ∑ j | β j | d j . \sup_{|\delta_j|\leq d_j}|f(x+\delta)-f(x)| =\sum_j|\beta_j|d_j.

The triangle inequality gives the upper bound. Choosing the endpoint of each error interval so that all contributions move the prediction in the same direction attains it. A feature with a large attribution may already have been thoroughly verified, leaving a small d j d_j . Conversely, a feature with an attribution close to zero may retain substantial unresolved error. Ranking verification work by absolute SHAP values omits this connection between uncertainty and the consequences of a correction.

Resampling the current records can also miss a common measurement bias. Suppose reported observations satisfy X i = μ + b + ε i X_i=\mu+b+\varepsilon_i , where the ε i \varepsilon_i are independent N ( 0 , σ 2 ) N(0,\sigma^2) errors and b ≠ 0 b\neq0 is an unmodeled bias shared by all records. A conventional normal interval centered on the sample mean covers the true μ \mu with probability

Pr { μ ∈ [ X ¯ − z σ n , X ¯ + z σ n ] } = Φ ( z − b n σ ) − Φ ( − z − b n σ ) ⟶ 0. \Pr\!\left\{\mu\in \left[\bar X-z\frac\sigma{\sqrt n},\bar X+z\frac\sigma{\sqrt n}\right]\right\} =\Phi\!\left(z-\frac{b\sqrt n}{\sigma}\right) -\Phi\!\left(-z-\frac{b\sqrt n}{\sigma}\right) \longrightarrow0.

Here z = Φ − 1 ( 0.975 ) z=\Phi^{-1}(0.975) . A larger sample reduces random error without removing the shared bias. A bootstrap that only resamples the same biased records has no additional information with which to identify b b . Verification must therefore inform the description of input uncertainty before a stability analysis can address the relevant source of error.

2 Joint uncertainty over models and evidence

Fix the prediction target, input definitions, and output scale. To give the model set a statistical basis, consider M M candidate models specified before examining an independent validation set. Assume independent and identically distributed validation observations and loss bounded in [ 0 , 1 ] [0,1] . For failure probability α \alpha , Hoeffding's inequality and a union bound give

Pr { max f ∈ F 0 | L ^ n ( f ) − L ( f ) | ≤ t n } ≥ 1 − α , t n = log ⁡ ( 2 M / α ) 2 n . \Pr\left\{\max_{f\in\mathcal F_0} |\widehat L_n(f)-L(f)|\leq t_n\right\}\geq1-\alpha, \qquad t_n=\sqrt{\frac{\log(2M/\alpha)}{2n}}.

Let f ∗ f^* minimize population loss within this candidate set, and retain

F α = { f ∈ F 0 : L ^ n ( f ) ≤ min g ∈ F 0 L ^ n ( g ) + 2 t n } . \mathcal F_\alpha= \left\{f\in\mathcal F_0:\widehat L_n(f)\leq \min_{g\in\mathcal F_0}\widehat L_n(g)+2t_n\right\}.

On the simultaneous concentration event, L ^ n ( f ∗ ) ≤ L ( f ∗ ) + t n ≤ min g L ^ n ( g ) + 2 t n \widehat L_n(f^*)\leq L(f^*)+t_n\leq\min_g\widehat L_n(g)+2t_n , so f ∗ ∈ F α f^*\in\mathcal F_\alpha . The same bound shows that every retained model has population excess loss at most 4 t n 4t_n . This gives an explicit criterion for retaining multiple models. It is a finite-candidate construction based on uniform convergence; related work on sets of explanations appears in Marx et al.. Repeatedly adapting models to the same validation set requires fresh independent evaluation or a bound that accounts for the selection process.

The evidence set X ( E ) \mathcal X(E) records unresolved questions about entity identity, reporting period, units, measurement definitions, and shared sources. Each element specifies a joint input configuration for all relevant companies. Dependencies are retained rather than represented as unrelated error bars for individual fields. Define Ω ( E ) = F α × X ( E ) \Omega(E)=\mathcal F_\alpha\times\mathcal X(E) . For companies i , k i,k , compute

Δ ― i k = inf ( f , X ) ∈ Ω ( E ) [ f ( x i ) − f ( x k ) ] , Δ ― i k = sup ( f , X ) ∈ Ω ( E ) [ f ( x i ) − f ( x k ) ] . \underline\Delta_{ik} =\inf_{(f,X)\in\Omega(E)}[f(x_i)-f(x_k)], \qquad \overline\Delta_{ik} =\sup_{(f,X)\in\Omega(E)}[f(x_i)-f(x_k)].

For a nonempty set with bounded differences, Δ ― i k > 0 \underline\Delta_{ik}>0 certifies a strict ordering throughout the stated set. If calibration also establishes that the evidence set contains the true inputs with probability at least 1 − β 1-\beta , a union bound gives joint inclusion of the candidate-optimal model and the true inputs with probability at least 1 − α − β 1-\alpha-\beta . The unconditional probability of issuing an incorrect ordering certificate is therefore at most α + β \alpha+\beta , without assuming that the two coverage events are independent. A scenario set specified solely by expert judgment supports a guarantee within that set; statistical coverage requires its own justification. The certificate concerns the defined score or prediction comparison, not future investment returns.

3 Computing comparison bounds with support functions

The joint evidence model also determines whether the calculation is tractable. Let u ∈ U u\in\mathcal U collect shared unresolved quantities, with company inputs x i ( u ) = x ¯ i + B i u x_i(u)=\bar x_i+B_i u . Under affine predictions, the pairwise difference has the form Δ f ( u ) = m f + a f ⊤ u \Delta_f(u)=m_f+a_f^\top u . With support function h U ( a ) = sup u ∈ U a ⊤ u h_{\mathcal U}(a)=\sup_{u\in\mathcal U}a^\top u ,

inf u ∈ U Δ f ( u ) = m f − h U ( − a f ) , sup u ∈ U Δ f ( u ) = m f + h U ( a f ) . \inf_{u\in\mathcal U}\Delta_f(u)=m_f-h_{\mathcal U}(-a_f), \qquad \sup_{u\in\mathcal U}\Delta_f(u)=m_f+h_{\mathcal U}(a_f).

For U = { C z : ‖ z ‖ 2 ≤ r } \mathcal U=\{Cz:\|z\|_2\leq r\} , the Cauchy-Schwarz inequality, with equality attained in the direction of C ⊤ a f C^\top a_f , yields h U ( a f ) = r ‖ C ⊤ a f ‖ 2 h_{\mathcal U}(a_f)=r\|C^\top a_f\|_2 . When U \mathcal U is a box centered at zero, the support function is a weighted sum of absolute values; a nonzero center adds its inner product with the direction. These expressions replace a class of continuous scenario searches with direct calculation or convex optimization. The underlying tools are established in robust optimization. See Ben-Tal and Nemirovski.

Joint modeling can also remove unnecessary conservatism. If two companies evaluated by the same linear model share an identical additive source error, then B i = B k B_i=B_k and a f = ( B i − B k ) ⊤ β f = 0 a_f=(B_i-B_k)^\top\beta_f=0 . The common error cancels exactly in the comparison. Constructing separate intervals for the companies and subtracting them would allow the same source to vary in opposite directions, introducing a worst case that cannot occur under the stated evidence model.

For a nonlinear model, the approximation error must be bounded. Set m f = Δ f ( 0 ) m_f=\Delta_f(0) and a f = ∇ Δ f ( 0 ) a_f=\nabla\Delta_f(0) . If the pairwise difference is twice continuously differentiable on ‖ u ‖ 2 ≤ r \|u\|_2\leq r , with Hessian operator norm bounded by K f K_f , Taylor's remainder gives

| Δ f ( u ) − m f − a f ⊤ u | ≤ κ f , κ f = 1 2 K f r 2 , Δ ― f ≥ m f − r ‖ a f ‖ 2 − κ f . |\Delta_f(u)-m_f-a_f^\top u|\leq\kappa_f, \qquad \kappa_f=\tfrac12K_fr^2, \qquad \underline\Delta_f\geq m_f-r\|a_f\|_2-\kappa_f.

A local linearization certifies the ordering only if the lower bound remains positive after the remainder is deducted. Tree models do not satisfy this smoothness condition; they require bounds based on leaf regions, enumeration, or a suitable global optimization method. An implementation can first use inexpensive bounds to resolve stable comparisons, then concentrate computation on company pairs close to the zero boundary.

4 Choosing verification queries by the regret they can remove

Stability intervals identify comparisons that need more evidence. Resource allocation also requires an account of how that evidence could change the decision. Let A \mathcal A be a finite action set, such as advancing a company to the next diligence stage, retaining a rating, or deferring judgment. For each scenario ω ∈ Ω \omega\in\Omega , define action utilities U ( a , ω ) U(a,\omega) in advance on a common scale. The minimax regret over deterministic actions is

R ∗ ( Ω ) = min a ∈ A sup ω ∈ Ω [ max b ∈ A U ( b , ω ) − U ( a , ω ) ] . R^*(\Omega)=\min_{a\in\mathcal A}\sup_{\omega\in\Omega} \left[\max_{b\in\mathcal A}U(b,\omega)-U(a,\omega)\right].

Hold the action set and utility definition fixed. A verification query q q yields a possible observation o o , restricting the scenarios to a nonempty subset Ω q , o ⊆ Ω \Omega_{q,o}\subseteq\Omega . Every original scenario must remain compatible with at least one possible observation. Define the guaranteed reduction in regret as

V ( q ) = R ∗ ( Ω ) − sup o ∈ O q R ∗ ( Ω q , o ) ≥ 0. V(q)=R^*(\Omega) -\sup_{o\in\mathcal O_q}R^*(\Omega_{q,o})\geq0.

The inequality follows from set inclusion. For any fixed action, worst-case regret cannot increase when the scenario set shrinks; minimizing over actions preserves the inequality. The definition requires no probability distribution over possible query outcomes. If a query can return no useful information, that outcome belongs in its outcome set. Noisy evidence must likewise retain scenarios compatible with the error mechanism. A valid reduction in uncertainty depends on what the observation actually establishes.

With several queries and a budget, the objective can be written as minimizing the remaining regret under the worst observation sequence, subject to total cost at most B B along every execution path. Selecting queries sequentially by V ( q ) / c q V(q)/c_q is a computable candidate policy. Complementarity between queries can make it suboptimal, so it should be compared with small batches of jointly selected queries or tractable discrete optimization. Regret reduction has an established role in information acquisition; this design applies it to evidence claims that determine company inputs and comparisons. See Robust Active Preference Elicitation.

Verification priority then depends on which unresolved claim can improve the current decision within budget. A shared source affecting several companies may be especially consequential: one query can narrow uncertainty in multiple comparisons at once.

5 Checking that corrections propagate through attribution

After deciding what to verify and obtaining a correction, the analysis must establish that the change has propagated through the outputs. Hold the prediction function f f , background Q Q , attribution definition, and output scale fixed. Shapley efficiency gives

f ( x ) = ϕ 0 + ∑ j ϕ j ( f , Q , x ) , ∑ j [ ϕ j ( f , Q , x ′ ) − ϕ j ( f , Q , x ) ] = f ( x ′ ) − f ( x ) . f(x)=\phi_0+\sum_j\phi_j(f,Q,x), \qquad \sum_j[\phi_j(f,Q,x')-\phi_j(f,Q,x)] =f(x')-f(x).

The second identity follows by subtracting the decompositions before and after correction. If approximate attributions have summation errors bounded by η , η ′ \eta,\eta' , respectively, the difference check permits a discrepancy of at most η + η ′ \eta+\eta' . This bound can serve directly as an acceptance criterion for recomputation. See SHAP's local accuracy property.

Each correction therefore carries versions for the affected fields, sources, model, background, and rules. Dependencies identify which scores, predictions, rankings, and explanations need to be recomputed. A change to the model or background requires a separate version comparison, so its effect is not attributed to the evidence correction. Dependency records establish the scope of recomputation; the identity checks the resulting values. Both are needed to verify a batch update.

6 Reproducible quantitative comparisons

The following analytical experiments use public, specified inputs to test the mechanisms derived above. These are constructed experiments, not client performance measurements.

The first experiment sets b = 0.2 , σ = 1 b=0.2,\sigma=1 . Ignoring the common bias, a nominal 95% interval has actual coverage of 48.40% at n = 100 n=100 , falling to approximately 0.000637% at n = 1000 n=1000 . If verification establishes | b | ≤ 0.2 |b|\leq0.2 , increasing the half-width to 0.2 + 1.96 / n 0.2+1.96/\sqrt n gives approximately 97.50% coverage in this example. Accounting for the shared evidence error restores coverage at the cost of a wider interval.

The second experiment examines a shared source. Two companies have a nominal difference of 0.6 and share the same additive source error in [ − 1 , 1 ] [-1,1] . Subtracting independently constructed intervals gives [ − 1.4 , 2.6 ] [-1.4,2.6] , which cannot certify an ordering. The joint source model gives [ 0.6 , 0.6 ] [0.6,0.6] , certifying the comparison. The tighter result removes incompatible combinations of errors; it does not require a smaller error range for the source itself.

The third experiment allows one query of equal cost across eight scenarios formed from two candidate models and four input configurations. The model differences are Δ s = 0.15 s + z 1 + 0.3 s z 2 \Delta_s=0.15s+z_1+0.3sz_2 , where s ∈ { − 1 , 1 } s\in\{-1,1\} , z 1 ∈ { − 0.4 , 0.4 } z_1\in\{-0.4,0.4\} , and z 2 ∈ { − 0.2 , 0.2 } z_2\in\{-0.2,0.2\} . A further shared field has a large attribution but has already been verified. The action is to select company i i or k k , with regret measured in model-score differences.

Verification queryRemaining regret after the worst possible observationGuaranteed reduction in regret
Recheck the resolved field with a large attribution0.610
Verify the smaller unresolved variable z 2 z_2 0.610
Verify z 1 z_1 , which determines the sign of the difference00.61

Enumerating all eight scenarios produces these results. Attribution ranking and decision-based verification can select different queries. In this simple separable example, a reasonable interval-sensitivity baseline also selects z 1 z_1 . Whether general regret optimization justifies its computational cost must be tested with shared sources, multiple comparisons, and sequential budgets. A production comparison should therefore include random verification, attribution ranking, interval sensitivity, and joint decision policies, counting optimization, retrieval, and expert judgment in the total cost.

Adoption should depend on the consequences for research: whether serious judgment errors, remaining regret, and error-correction rates improve under the same total budget; whether every affected conclusion is recomputed when evidence changes; and whether the advantage persists across defensible model sets and evidence ranges. The proposed protocol brings input validity, model uncertainty, and comparison stability into one computable decision model for allocating subsequent verification work.