Past Experience

01

Artificial Intelligence

Alignerr

ProjectFrontier LLM Capability Boundary Probing via Graduate-Level Physics & Risk-Engineering Benchmarks

  • Built physics problems centered on long-form derivations, invariant-based reasoning, and tight constraint management to prevent shortcut solutions
  • Recombined rigorous physics foundations into mechanics-heavy prompts that demand both physical intuition and careful mathematical derivations
  • Converted risk-engineering research themes into quantitative prompts that require explicit assumptions, nontrivial modeling choices, and end-to-end reasoning
  • Raised difficulty by layering advanced techniques and edge conditions, pushing problems beyond standard textbook patterns
  • Representative example: designed a right-triangle kinematics puzzle with path-invariant vertical time, reducing optimization to horizontal scheduling; proved the fastest path, showed the slowest path has no maximizer but a supremum, and derived harmonic-number asymptotics for a staircase construction approaching the bound
  • Reused the same layered-difficulty design playbook across multiple topics and scenarios, balancing high-compute reasoning with creativity

Alignerr

ProjectEnd-to-End RLHF Preference Data Engineering with Scenario Design, Rubrics & A/B Ranking Signals

  • Owned end-to-end human preference data engineering for RLHF alignment, designing scenarios, datasets, and rubrics to rank multiple model variants via A/B preference tests
  • Crafted evaluation scenarios that probe where strong reasoning diverges from human engineering criteria, turning alignment goals into testable, domain-grounded tasks
  • Built realistic input datasets with natural messiness and confounders, avoiding contrived control-variable setups so model behavior reflects real decision contexts
  • Defined quantifiable rubrics and structured response requirements to convert subjective judgments into consistent, rankable preference signals
  • Authored reusable task templates plus data-generation rules so each scenario can be re-instantiated and rerun at high volume without synthetic-looking artifacts
  • Representative example: designed an ESG rating evaluation that surfaces old-company advantage bias by pairing comparable new vs. established firms and rewarding substantive performance over disclosure tenure; applied the same bias-probing approach across multiple domains and scenarios
  • Produced pairwise rankings with graded preference strength and clear comparative rationales that pinpointed failure modes and decision-critical trade-offs
  • Supported evaluator consistency through lightweight calibration on edge cases and evidence re-check steps when inputs or sources were inconsistent

Micro1

ProjectLLM Contextual Advertising Domain Leadership & A/B-Validated Decision Playbooks

  • Served as the domain expert for contextual advertising in LLM chat experiences, translating ambiguous ad-strategy tradeoffs into clear decision criteria
  • Used a consistent workflow of marketing research, structured debate, interviews with senior domain experts, and A/B validation to raise decision confidence and reduce low-value experiments
  • Applied the same decision framework across consumer categories including travel, food delivery, FMCG, and e-commerce, ensuring consistency rather than one-off judgments
  • Representative example in travel-intent ads: compared a shortlist display of about 20 hotels with a full-inventory display of more than 100 hotels for a booking brand; recommended full inventory to capture clicks and reduce competitive leakage; validated via A/B testing and documented the guideline for future use
  • Codified when scarcity marketing is likely to work by defining criteria such as decision-journey length, market-leader dynamics, and luxury or design-led categories, avoiding misapplication in long-journey decisions like travel planning
  • Narrowed experiment scope by replacing step-by-step settings with 2 representative endpoints, default versus scarcity, and validating the best default through A/B testing
  • Maintained evaluation rigor at scale through rubric-based scoring and written rationales, calibration and adjudication on borderline relevance and usefulness, and validity checks under landing-page variants, geo or session differences, redirects, and page changes; flagged rubric and workflow gaps as tasks evolved

Duke University

ProjectSustainability Event Intelligence Labeling for Equity Research Backtests & AI Workflows

  • Designed a labeled dataset linking firm-level sustainability events to subsequent stock returns, preparing clean inputs for equity research backtests, systematic screening rules, and future ML or AI workflows
  • Reviewed 10-K filings, sustainability reports, proxy statements, and news flow, coding each event’s dimension, direction, and estimated materiality, then merging labels with return and fundamentals data in Python