Task validity in financial synthetic data
Project 2: Financial Task Design and Synthetic Data Engineering
Author: Theodore Ouyang
I led the design of a task space in which experts could consistently produce financial and business reasoning cases for research teams to validate. The playground, or task environment, specified the available information, business relationships, and capabilities to be tested. Golden cases, or reference cases, supplied the reference judgments and the evidence supporting them. As synthetic expansion changed each case, its reference solution and preference review were updated accordingly. The company subsequently adopted Batch Zero, the pilot process I created, for other new tasks.
The difficulty was that a generator could approximate the statistical distribution of real data increasingly well while still producing large numbers of cases that violated business requirements, retained outdated answers, or gave predictors access to future information. The analysis below develops the project's design principles through mathematical analysis and reproducible tests. All numerical results come from experiments with publicly specified assumptions that isolate particular failure mechanisms and their corrections.
1 Why better distributional fit may leave task validity unchanged
CTGAN addresses mixed data types, multimodality, and class imbalance. TabDDPM applies Gaussian and multinomial diffusion to numerical and categorical variables, respectively. Both provide useful tools for learning distributions. Financial task design also has to preserve accounting identities, chronological order, the conditions under which a rule applies, and the relationship between inputs and reference judgments. A single measure of overall distributional similarity cannot replace these requirements. CTGAN, TabDDPM
Consider a failure mode that admits an exact analysis. Valid inputs satisfy , where has full row rank. The valid distribution is supported on the affine subspace and has finite second moments. Let the columns of form an orthonormal basis for its normal space. Suppose a generator adds a small amount of noise in these normal directions: , where is independent of . Write the generated distribution as . For every ,
The first line uses a coupling constructed from the same . The second follows because is invertible and a continuous Gaussian random vector equals zero with probability zero. For the third, take . Thus, Wasserstein distance can converge to zero while the acceptance rate under exact business constraints remains zero. This is a statement about what the optimization criterion guarantees, rather than a performance finding about a particular generator.
The counterexample determines how to divide responsibilities in the production specification. Copula models, CTGAN, and TabDDPM learn statistical relationships. Task design specifies which relationships must hold exactly, what numerical tolerances are acceptable, and which changes in the inputs require a different reference judgment. Explicit transformations and rejection sampling provide this additional constraint enforcement in existing generation systems. SDV constraint-augmented generation
2 Enforce business constraints while preserving the correct conditional distribution
Rejecting invalid candidates after generation is straightforward, but it can become impractical as several constraints tighten together. Suppose a candidate has standardized constraint residuals that are independent standard normal variables, and each must lie in . The acceptance probability and the expected number of attempts needed for one valid input are
Each additional constraint multiplies the expected number of candidates required. Setting each tolerance to 0.05 times its residual standard deviation gives the following analytical results.
Generating through the degrees of freedom allowed by the constraints avoids these repeated rejections. In the linear Gaussian case, the construction also preserves the correct conditional distribution. Let the initial candidate be , with . Use a covariance-weighted projection:
The first-order Lagrangian condition is . Substituting the constraint and eliminating gives the second line. Since , every draw satisfies . Its moments are
An affine transformation preserves Gaussianity, and these moments give the usual regular conditional Gaussian law of . The construction establishes both feasibility and compatibility with the specified probability model. An arbitrary edit to one field may satisfy the equality without preserving that conditional distribution. Majumdar and Majumdar on Gaussian distributions under linear conditioning
The accompanying program generates 200,000 samples in 12 dimensions under 4 linear constraints. The maximum absolute constraint residual is . The relative Frobenius error of the empirical covariance against the theoretical conditional covariance is 0.662%. Every draw from this construction is feasible for the specified equalities. The figure 395,442.59 in the table is a ratio of candidate counts; runtime also depends on matrix factorization, the cost of each draw, and the remaining acceptance checks.
Method selection in the project followed the same principle. Fields determined by basic quantities were computed from their definitions. Bounds and date orderings that admitted transformations were handled in the generation space. Dependencies that had to be learned were assigned to an appropriate generator. For nonlinear constraints, a parameterization also needs to be checked for coverage of the feasible region and for any required density correction. A Gaussian copula's conditioning formula in latent Gaussian space does not, by itself, enforce nonlinear constraints in the original business variables.
3 Use decision margins to determine when reference judgments must be recomputed
Expanding a financial case often preserves its narrative while changing decisive numbers, evidence, or conditions. Retaining the original reference answer can then produce systematically incorrect supervision, even when every input is valid. FinQA organizes questions, supporting evidence, and executable reasoning programs together. Synthetic expansion requires that connection between an answer and its justification to be updated when the inputs change. FinQA
Consider a specified rule for choosing between two alternatives. Let be the difference between their evaluations, so the reference decision is given by . If is Lipschitz on a feasible neighborhood, a sufficient condition for safely retaining the original decision is
Both and must be feasible, with the same rule and available information. The condition makes the effect of a numerical edit testable: the perturbation must be smaller than the decision margin permits. Cases near the boundary may be especially useful for research, but their labels require particular care when inputs change.
For a closed-form error calculation, take a comparison rule that is exactly linear on the support of the perturbed inputs, , with initial margin . Let the feasible perturbation lie in the allowed subspace, with . The error rate from retaining the original decision is
When the margin-to-noise ratio is 0.25, 1, or 2, the inherited-label error rate is 40.13%, 15.87%, or 2.28%, respectively. In this experiment, where the rule is known and evaluated exactly, recomputing eliminates this source of stale-label error. Open-ended financial judgments still require independent validation of the reference evidence. A simulation with one million draws checks the analytical probabilities.
This is why I connected case expansion rules to reference-solution review. A golden case should identify the conditions supporting its conclusion and the variables that determine it. Expansion can deliberately cover both sides of a decision boundary. Each new input first receives a valid reference judgment, followed by the corresponding model responses and a new preference assessment on the five-level scale. Preference strength also depends on the nature of the errors and the quality of the reasoning; a numerical margin alone cannot determine it. HelpSteer2-Preference
4 Define the information available for each temporal task
SDEs and neural SDEs treat entire trajectories as the objects to be generated. Neural SDEs model distributions over paths, allowing them to represent dependencies beyond single-time marginals. The use of a path generator therefore requires an explicit account of the information on which generation is conditioned. Neural SDEs as Infinite-Dimensional GANs
A Brownian bridge provides a failure mode whose size can be calculated directly. Let , with . Under squared loss, the optimal predictor given the history at time has conditional mean squared error . If the future endpoint is also observed, Gaussian conditioning gives
The second line conditions on the additional observation using the covariance between the two future increments. When is the midpoint between and , the conditional variance is halved, producing an apparent 29.29% reduction in RMSE. The entire gain comes from additional information. Giving the predictor the endpoint would be evaluation leakage if the forecasting task prohibited that observation. If the task explicitly supplies the endpoint and asks for interpolation or analysis of the intervening path, this conditioning is appropriate. Pitman on the Brownian bridge
The task environment must therefore define separately the information used to generate a case, the information visible to the predictor, and the evidence used for scoring. A generator may use a complete path or latent variables to construct a case, and scoring may use subsequently realized ground truth. A predictor acting at time is restricted to the task's permitted . Providing future information to that predictor constitutes leakage. A comparison against a reference predictor with additional future information must disclose the difference in information sets; an informational advantage cannot be attributed to superior model capability. An oracle whose additional information is explicitly stated is a legitimate comparator. Selecting or generating cases conditional on a future endpoint may also change the target evaluation distribution. That conditional target must be specified. Estimating performance under the original target distribution instead requires adequate support coverage, identifiable distribution weights, and an appropriate correction. The generator must also respect the domains of its state variables, whether these are prices, log prices, or other quantities. For neural SDEs, conditional path sampling has to match the underlying dynamics; a Gaussian bridge formula cannot substitute for every process.
5 Test independent execution through Batch Zero and retain the task family as a statistical unit
After verifying constraints, reference answers, and information conditions, the production method still has to work when other experts execute it. I created Batch Zero so that experts with established records of reliable delivery could independently construct tasks, expand inputs, produce reference solutions, and assess preferences. Failed cases and disagreements then informed revisions to the specification. Its subsequent adoption for other tasks is an observed organizational outcome. Further quantitative evaluation should test the quality of independent execution.
Synthetic sample size is easy to overstate. Suppose there are independent golden-case families, each containing variants. A quality measure has common variance , and any two records within a family have correlation . Expanding the covariance sum for the sample mean gives
With 100 families, 20 records per family, and within-family correlation of 0.5, the nominal 2,000 records provide the same precision for the mean as approximately 190.48 independent records. Ignoring the correlation would understate the standard error by a factor of approximately 3.24. This effective sample size describes the stated variance model; inference should still operate at the family level.
Production and evaluation therefore need consistent units. Training and test sets are separated by source family. Pilot tasks cover different scenarios and decision boundaries. Confidence intervals account for the relevant family or higher-level clustering, so similar variants do not inflate the evidence for improvement. Bias shared across families because of a common expert also requires an additional clustering or hierarchical component.
6 Establish whether the complete method justifies scaled production
A realistic baseline retains the existing generator and its normal expert review and rework process. Hold the task mix, generator training budget, expert budget, and final acceptance criteria fixed, then add constraint enforcement, reference recomputation, and checks on information timing in stages. Reviewers should not know which method produced a case, and the complete deliverables should be evaluated across independent families.
Acceptance rate alone can reward avoiding difficult cases. Generation speed alone omits expert repair. Let index predefined task cells, with weights , , and a required number of independent valid families . A coverage and cost measure for production decisions is
Weights and targets are fixed before comparison, together with the rules for counting each family. Repeatedly expanding easy cases then ceases to increase the score once their coverage target has been met. Coverage efficiency is compared only among methods that meet the same acceptance requirements for serious failure rates, reference correctness, and information conditions. Interval estimates and tests of differences operate at the family level. Distributional fidelity, coverage, and the risk of copying existing records are assessed separately because they address different aspects of generation quality. Alaa et al. on synthetic data evaluation
The mechanisms of improvement are specific. Constraint enforcement reduces candidates outside the feasible space. Recomputing references prevents changed inputs from becoming incorrect supervision. Checks on information timing exclude invalid forecasting advantages, while pilot production and evaluation at the family level test whether the method scales. Together, these checks determine whether generated data can support credible evaluation and post-training research. They also provide comparable datasets for subsequent experiments that hold the model and training budget fixed.