AI Research and Engineering

Statistical inference and resource allocation in expert data production

Project 1: Scalable Expert Quality Control and Adaptive Review
Author: Theodore Ouyang

Joint design of selective verification, residual-corrected estimation, and reviewer capacity

Review policies change the distribution of errors we observe. Directing more expert time toward difficult submissions can raise the error rate in the reviewed sample and assign harder work to stronger experts. Uncorrected quality statistics may therefore penalize an effective allocation policy. The technical question is how to improve delivered work while preserving quality evidence that can support the next round of decisions.

In this project, I co-developed operating procedures for assessing contributor reliability, assigning tiered review, conducting random expert audits, and updating evaluation criteria. The analysis below formalizes those production practices as a testable research design. The derivations specify the conditions under which the design is valid, and reproducible numerical experiments quantify the gains under those conditions.

1 The observation mechanism creates its own failure modes

Let X X denote task features and Y ∈ { 0 , 1 } Y\in\{0,1\} indicate an error in the original submission. Under fixed acceptance criteria, the conditional error probability is q ( x ) = P ( Y = 1 ∣ X = x ) q(x)=P(Y=1\mid X=x) . Let S = 1 S=1 indicate selection for an independent reference audit. Conditional on X X , selection is randomized with probability π ( x ) > 0 \pi(x)>0 , independently of the reference label that has yet to be observed.

Using the error fraction among verified submissions to assess the full population targets:

E [ Y ∣ S = 1 ] = E [ π ( X ) q ( X ) ] E [ π ( X ) ] , E [ Y ∣ S = 1 ] − E [ Y ] = Cov ⁡ ( π ( X ) , q ( X ) ) E [ π ( X ) ] . \begin{aligned} E[Y\mid S=1] &=\frac{E[\pi(X)q(X)]}{E[\pi(X)]},\\ E[Y\mid S=1]-E[Y] &=\frac{\operatorname{Cov}(\pi(X),q(X))}{E[\pi(X)]}. \end{aligned}

When review priorities identify high-risk tasks, the covariance is typically positive. A higher error detection rate in the reviewed sample cannot then be read directly as a deterioration in overall quality. This identity describes the distribution of a selected item. In a finite batch, the ratio of observed errors to selected items also has a random denominator, so the two quantities should not be conflated.

Expert rankings are affected by the same mechanism. Let a j h a_{jh} be the probability that expert j j is correct on task stratum h h , and let p j h t p_{jht} describe the mix of tasks assigned to that expert in period t t . Even with unchanged ability, observed accuracy can change:

A j t = ∑ h p j h t a j h , A j , t + 1 − A j t = ∑ h ( p j h , t + 1 − p j h t ) a j h . A_{jt}=\sum_h p_{jht}a_{jh}, \qquad A_{j,t+1}-A_{jt} =\sum_h(p_{jh,t+1}-p_{jht})a_{jh}.

Comparisons therefore require a fixed target task mix w h w_h , reporting A j std = ∑ h w h a j h A_j^{\mathrm{std}}=\sum_h w_ha_{jh} , with overlapping assignments or shared anchor cases to support estimation within each stratum. A beta-binomial posterior can account for unequal sample sizes in comparable audits. Pooling different difficulty levels, criteria versions, and selection mechanisms into a single binomial sample leaves the preceding bias unresolved.

These distinctions also guide method selection. Dawid-Skene models represent differences in reviewer error patterns. DeCCaF already incorporates task conditions, error costs, and reviewer capacity into deferral, and discusses how assigning difficult cases can distort evaluations of expert performance. The design here addresses the interface between these methods and population quality inference in ongoing production. Dawid-Skene, DeCCaF

2 Correct predictions with audits to recover the population target

Fix a batch of N N original submissions, its acceptance criteria, and its reference labels Y 1 , … , Y N Y_1,\ldots,Y_N . The target is the finite-batch error rate μ = N − 1 ∑ i Y i \mu=N^{-1}\sum_iY_i . Let f i ∈ [ 0 , 1 ] f_i\in[0,1] be a fixed, inexpensive risk prediction and π i ∈ [ π min , 1 ] \pi_i\in[\pi_{\mathrm{min}},1] the audit probability, with independent draws S i ∼ Bernoulli ( π i ) S_i\sim\mathrm{Bernoulli}(\pi_i) . Both predictions and probabilities are set before the current batch's reference labels are revealed.

Use a prediction-assisted, residual-corrected estimator:

μ ^ = 1 N ∑ i = 1 N [ f i + S i π i ( Y i − f i ) ] . \widehat\mu =\frac1N\sum_{i=1}^N \left[f_i+\frac{S_i}{\pi_i}(Y_i-f_i)\right].

Its validity follows directly. Write e i = Y i − f i e_i=Y_i-f_i and condition on the fixed batch, predictions, and sampling probabilities:

μ ^ − μ = 1 N ∑ i ( S i π i − 1 ) e i , E S [ μ ^ − μ ] = 0. \widehat\mu-\mu =\frac1N\sum_i\left(\frac{S_i}{\pi_i}-1\right)e_i, \qquad E_S[\widehat\mu-\mu]=0.

Inverse-probability weighting restores representation altered by selective verification. Predictions reduce the residual that expert audits must resolve. This design-based unbiasedness does not require f i = q ( X i ) f_i=q(X_i) : inaccurate predictions can reduce efficiency, but they do not introduce the same selection bias under the specified sampling design. The estimator follows established active statistical inference methods. The project design specifies the object being audited, the criteria version, and the records needed for subsequent decisions. Active Statistical Inference: mean estimation and sampling design

Original and revised submissions must be retained separately. Estimating μ raw \mu^{\mathrm{raw}} on originals and μ final \mu^{\mathrm{final}} on final submissions distinguishes defects generated in production from defects corrected by review. When criteria change, the same anchor cases can be adjudicated under both versions; submissions from different periods can then be compared under a common version. This prevents stricter standards from being mistaken for declining ability.

3 Allocate audit effort according to residuals and cost

Independent Bernoulli sampling eliminates cross-covariance terms, giving:

Var S ⁡ ( μ ^ ) = 1 N 2 ∑ i 1 − π i π i e i 2 . \operatorname{Var}_S(\widehat\mu) =\frac1{N^2}\sum_i\frac{1-\pi_i}{\pi_i}e_i^2.

This expression identifies the allocation objective: audit effort should address what the predictions have not resolved. For a binary error label with conditional error probability q i q_i ,

m i := E [ ( Y i − f i ) 2 ∣ X i ] = q i ( 1 − q i ) + ( q i − f i ) 2 . m_i:=E[(Y_i-f_i)^2\mid X_i] =q_i(1-q_i)+(q_i-f_i)^2.

The two terms are conditional variance and squared prediction error. A high error probability does not necessarily imply high inferential value. If a category is almost certain to contain an error and the predictor correctly recognizes that fact, another audit may provide less information than one on a moderately risky task whose outcome is uncertain. The high-risk task may still merit correction, so the value of measurement and the value of remediation require separate models.

Let c i > 0 c_i>0 be the cost of an independent reference audit and B B the expected audit-cost budget. First consider the optimal design when the true residual second moments m i > 0 m_i>0 are known:

min π ∑ i m i π i , s.t. ∑ i c i π i ≤ B , π min ≤ π i ≤ 1. \begin{aligned} \min_{\boldsymbol\pi}\quad &\sum_i\frac{m_i}{\pi_i},\\ \text{s.t.}\quad &\sum_i c_i\pi_i\le B,\qquad \pi_{\mathrm{min}}\le\pi_i\le1. \end{aligned}

The omitted constant ∑ i m i \sum_i m_i does not depend on the allocation. For probabilities strictly between their lower and upper bounds, the Lagrangian first-order condition yields:

− m i π i 2 + λ c i = 0 ⟹ π i ∗ = clip [ π min , 1 ] ⁡ m i λ c i . -\frac{m_i}{\pi_i^2}+\lambda c_i=0 \quad\Longrightarrow\quad \pi_i^* =\operatorname{clip}_{[\pi_{\mathrm{min}},1]} \sqrt{\frac{m_i}{\lambda c_i}}.

Here λ > 0 \lambda>0 is determined by the budget, and feasibility requires B ≥ π min ∑ i c i B\ge\pi_{\mathrm{min}}\sum_i c_i . If the budget covers a full audit, every item is audited. The square-root rule increases sampling for larger residuals while reducing the relative frequency of expensive audits. The positive probability floor preserves observation across the task space.

Implementation substitutes regularized estimates m ^ i > 0 \widehat m_i>0 obtained from earlier independent audits or a pilot batch. True m i m_i defines the optimality analysis; estimated values determine the allocation. Residuals can be estimated by task stratum, criteria version, and cost stratum, without fitting unstable parameters for individual submissions. A schedule that requires a fixed number of audits in every batch instead needs a fixed-size sampling design and its inclusion probabilities. The budget above is an expectation, not a hard capacity guarantee for each batch.

4 Derive the gain and identify when it disappears

Compare uniform audit probabilities with cost-aware allocation using the same predictor. Let C = ∑ i c i C=\sum_i c_i , so that the uniform probability is p = B / C p=B/C . If the true residual second moments are available and no optimal probability reaches a clipping boundary,

V uniform = 1 N 2 [ C ∑ i m i B − ∑ i m i ] , V optimal = 1 N 2 [ ( ∑ i m i c i ) 2 B − ∑ i m i ] . \begin{aligned} V_{\mathrm{uniform}} &=\frac1{N^2}\left[ \frac{C\sum_i m_i}{B}-\sum_i m_i\right],\\ V_{\mathrm{optimal}} &=\frac1{N^2}\left[ \frac{(\sum_i\sqrt{m_ic_i})^2}{B}-\sum_i m_i\right]. \end{aligned}

Here V V is design variance averaged over the conditional label distribution. If a common sampling probability is used within each fixed task stratum, substituting that stratum's true mean squared residual for m i m_i also gives an exact finite-batch result. The Cauchy-Schwarz inequality implies:

V uniform − V optimal = C ∑ i m i − ( ∑ i m i c i ) 2 N 2 B ≥ 0. V_{\mathrm{uniform}}-V_{\mathrm{optimal}} =\frac{C\sum_i m_i-(\sum_i\sqrt{m_ic_i})^2}{N^2B} \ge0.

Equality holds if and only if m i / c i m_i/c_i is constant across items. Gains therefore depend on identifiable differences in residual second moment per unit audit cost. With homogeneous residuals and costs, the optimal allocation reduces to uniform sampling. When clipping is active, constrained optimization can still be compared with a feasible uniform policy, but the gain must be calculated from the clipped solution's objective value.

The same analysis exposes the cost of misspecification. If the allocation uses inaccurate estimates m ^ i \widehat m_i and neither allocation is clipped, the excess variance relative to the true optimum is:

V ( π ^ ) − V ( π ∗ ) = 1 N 2 B [ ( ∑ i m ^ i c i ) ( ∑ i m i c i m ^ i ) − ( ∑ i m i c i ) 2 ] ≥ 0. V(\widehat{\boldsymbol\pi})-V(\boldsymbol\pi^*) =\frac1{N^2B} \left[ \left(\sum_i\sqrt{\widehat m_ic_i}\right) \left(\sum_i m_i\sqrt{\frac{c_i}{\widehat m_i}}\right) -\left(\sum_i\sqrt{m_ic_i}\right)^2 \right]\ge0.

Severely underestimating m ^ i \widehat m_i in part of the task space increases the penalty through the denominator. A positive sampling floor limits the consequences of misspecification; independent calibration batches determine whether active allocation is worthwhile. Mixing active and uniform probabilities as π i ( ρ ) = ( 1 − ρ ) π ^ i + ρ p \pi_i^{(\rho)}=(1-\rho)\widehat\pi_i+\rho p preserves the same cost budget. The mixing coefficient should be selected through independent audit comparisons. Adding a fixed amount of random sampling does not itself guarantee an improvement over uniform sampling. Robust active statistical inference studies budget-preserving paths and optimization under residual misspecification. Robust Sampling for Active Statistical Inference

5 Allocate review by expected loss reduction

For task i i under action a a , let r i a r_{ia} be the expected error loss after the action, d i a d_{ia} its cost, and t i a j t_{iaj} the time required from expert j j . Actions may include additional review, expert adjudication, obtaining further evidence, or revision. Binary choices z i a ∈ { 0 , 1 } z_{ia}\in\{0,1\} satisfy:

min z F ( z ) = ∑ i , a z i a ( r i a + η d i a ) , ∑ a z i a = 1 , ∑ i , a z i a t i a j ≤ H j ∀ j . \begin{aligned} \min_z\quad F(z) &=\sum_{i,a}z_{ia}(r_{ia}+\eta d_{ia}),\\ \sum_a z_{ia}&=1,\\ \sum_{i,a}z_{ia}t_{iaj}&\le H_j\qquad\forall j. \end{aligned}

Here H j H_j is available reviewer time, and η ≥ 0 \eta\ge0 converts resource cost to the scale of error loss. DeCCaF provides a basis for cost- and capacity-constrained deferral. This design defines an action as the complete handling procedure and uses independent reference audits to evaluate residual risk after that procedure. DeCCaF: capacity constraints and global assignment

When review costs are equal and each task uses one slot, the appropriate ranking is by gain, Δ i = r i 0 − r i 1 \Delta_i=r_{i0}-r_{i1} . Missing evidence may make a high-risk task difficult to repair through one more review, while a clear calculation error in a lower-risk task may be readily corrected. For tasks i , k i,k , moving the review slot from i i to k k changes total loss by exactly:

Δ F = ( r i 0 + r k 1 ) − ( r i 1 + r k 0 ) = Δ i − Δ k . \Delta F=(r_{i0}+r_{k1})-(r_{i1}+r_{k0}) =\Delta_i-\Delta_k.

Whenever Δ k > Δ i \Delta_k>\Delta_i , the exchange strictly reduces loss, regardless of the tasks' initial risk ranking. Under more complex capacity constraints, a batch optimizer can apply the same gain-based objective while excluding infeasible assignments, such as sending every difficult case to one expert.

The precision required of risk estimates can also be stated explicitly. Suppose | r ^ i a − r i a | ≤ δ \lvert\widehat r_{ia}-r_{ia}\rvert\le\delta for every feasible action, costs and capacities are known, and z ^ \widehat z exactly minimizes the predicted objective. Relative to the true optimal feasible assignment z ∗ z^* ,

F ( z ^ ) ≤ F ^ ( z ^ ) + N δ ≤ F ^ ( z ∗ ) + N δ ≤ F ( z ∗ ) + 2 N δ . F(\widehat z) \le\widehat F(\widehat z)+N\delta \le\widehat F(z^*)+N\delta \le F(z^*)+2N\delta.

This bound identifies post-action loss as the calibration target. An approximate optimizer adds its objective-error bound. If action losses cannot be estimated on comparable tasks, optimization cannot supply the missing evidence. Production validation should therefore prioritize randomized pilots among permitted actions, calibrating risk against task conditions, criteria versions, and outcomes of the complete procedure.

The two budgets can also enter a single batch-level planning problem. Let τ 2 \tau^2 be the target variance of the quality estimate, B total B_{\mathrm{total}} the total cost budget, and V m ^ ( π ) = N − 2 ∑ i ( 1 − π i ) m ^ i / π i V_{\widehat m}(\boldsymbol\pi)=N^{-2}\sum_i(1-\pi_i)\widehat m_i/\pi_i the planning variance computed from the fixed residual model. The joint objective is:

min z , π ∑ i , a z i a r ^ i a , s.t. ∑ i , a z i a d i a + ∑ i c i π i ≤ B total , V m ^ ( π ) ≤ τ 2 , ∑ a z i a = 1 , z i a ∈ { 0 , 1 } , ∑ i , a z i a t i a j ≤ H j review ∀ j , π min ≤ π i ≤ 1. \begin{aligned} \min_{z,\boldsymbol\pi}\quad &\sum_{i,a}z_{ia}\widehat r_{ia},\\ \text{s.t.}\quad &\sum_{i,a}z_{ia}d_{ia}+\sum_i c_i\pi_i\le B_{\mathrm{total}},\\ &V_{\widehat m}(\boldsymbol\pi)\le\tau^2,\\ &\sum_a z_{ia}=1,\quad z_{ia}\in\{0,1\},\\ &\sum_{i,a}z_{ia}t_{iaj}\le H_j^{\mathrm{review}} \qquad\forall j,\\ &\pi_{\mathrm{min}}\le\pi_i\le1. \end{aligned}

Here H j review H_j^{\mathrm{review}} is hard capacity allocated specifically to review actions; statistical audits are scheduled separately. This mixed-integer program with a convex audit constraint imposes a model-based precision target for subsequent decisions while seeking to reduce current defects. The audit target is the frozen original submission, so m i m_i does not depend on the batch's revision actions. Estimating final-submission quality instead requires an action-dependent residual model. Audit cost remains an expectation, and the formulation makes no hard guarantee about realized audit hours in each batch. Fixed quotas require the corresponding sampling design. Independent audits must check residual calibration to establish whether the variance target is met. An infeasible program makes the conflict between quality requirements and available resources explicit.

6 Reproducible experiments at equal cost

The following synthetic experiments test the mechanisms; they do not report client outcomes. A fixed batch contains 10,000 submissions in strata of 6,000, 3,000, and 1,000 items, with 60, 300, and 350 errors, respectively. The overall error rate is 7.10%. Audit costs are 1, 1.5, and 3 units per item. Every policy has an expected audit-cost budget of 2,000 units and a minimum sampling probability of 3%.

Risk predictions for the three strata are 2%, 8%, and 30%. Their batch average is 6.60%, differing from the true error rate by 0.50 percentage points. Cost-aware allocation uses the true within-stratum mean squared residuals, representing ideal residual calibration. The misspecified policy reverses the residual ordering across strata. With predictions, budget, and original submissions held fixed, 200,000 independent sampling repetitions check the analytical variances.

Audit policySampling probabilities across the three strataStandard deviation of the corrected estimate, percentage pointsVariance change relative to uniform sampling
Uniform sampling with the same residual correction14.81% / 14.81% / 14.81%0.568Baseline
Prioritize predicted error probability3.00% / 11.56% / 43.33%0.65733.77% increase
Residual- and cost-aware allocation with ideal calibration7.89% / 19.37% / 21.84%0.51517.79% reduction
Misspecified residual ordering23.00% / 11.77% / 3.00%0.983199.66% increase
Misspecified allocation with 70% weight on uniform sampling17.27% / 13.90% / 11.27%0.61517.07% increase

Under ideal residual calibration, standard deviation falls by 9.33%. This measures greater precision in quality estimation, not a reduction in submission errors. For the high-risk-first policy, the uncorrected detection rate averages approximately 19.60%, while the corrected estimate averages approximately 7.1015%. For all five policies, Monte Carlo variances differ from their analytical values by less than 0.5%, confirming agreement between the sampling implementation and the estimator.

A separate, exactly calculable allocation experiment isolates remediation gains. Two groups contain 100 items each, with initial error risks of 0.40 and 0.15. Additional review reduces those risks to 0.30 and 0.01, respectively. Review costs are equal and capacity is 100 items. Ranking by initial risk leaves 100 ( 0.30 ) + 100 ( 0.15 ) = 45 100(0.30)+100(0.15)=45 expected errors; ranking by expected loss reduction leaves 100 ( 0.40 ) + 100 ( 0.01 ) = 41 100(0.40)+100(0.01)=41 . Expected error loss falls by 8.89% at the same review volume. This example treats action risks as known to test the allocation mechanism; production gains also depend on the accuracy of estimated action risks.

With homogeneous residuals and costs, optimal audit probabilities become uniform. With equal remediation gains, exchanging review slots yields no improvement. These two zero-gain conditions, together with the misspecification experiment, determine when more complex policies merit adoption.

7 Establish whether the design improves production

Production trials should test quality inference and control of delivered work separately. Detecting more errors is not, by itself, evidence of better delivered quality.

First, use frozen batches with complete independent adjudication to compare uniform sampling, the existing high-risk review policy, cost-aware residual allocation, and calibrated mixtures. Keep the predictor fixed and include pilot audits in the budget. Compare bias, mean squared error, and interval coverage computed for the actual sampling design. Examine task strata and criteria versions separately to check that aggregate improvements do not conceal missed defects in particular groups.

Next, compare existing tiered rules with gain-based action allocation at equal total expert hours. Measure serious defects in final submissions and total cost per accepted submission. Cost includes routine review, random audits, adjudication, revision, and resources occupied while work waits. If shared reviewer capacity creates interference between tasks, randomize policies by batch or matched time window and estimate uncertainty at that level.

Adoption can then follow an operational criterion: reduce cost per accepted submission while remaining within a prespecified tolerance for serious defects, or exceed a prespecified minimum meaningful reduction in residual serious errors at equal total cost. Audit precision must remain within its required range. Ablation studies should distinguish gains from the inference method, gains from review actions, and gains from their combination.

The contributor records, tiered review procedures, and random audits I established in production provide an operational basis for this design. The further research contribution is to distinguish three quantities: a task's initial risk, the information value of an independent reference audit, and the loss that additional handling can remove. Each budget then has a defined optimization objective, and quality evidence remains interpretable after the review policy changes.