Practical AI for real work. Clear checks. Human ownership.

I help local and midsize teams put AI into daily operations without buying a black box. I also take contract and full-time remote roles in AI data, evaluation, and quality systems.

Put AI into daily work without buying a black box.

I help you find where AI can help, run a small pilot you can see pass or fail, and set simple rules for when a person must decide. Built for clinics, manufacturers, schools, professional firms, and other midsize teams that want AI in operations, customer service, documents, or scheduling without a research lab on staff.

AI readiness check

Where AI can help, where it should not, and what data you already have. You finish with a decision you can defend, not a shopping list.

  • Where it helps
  • Where it should not
  • What data exists

Pilot build

One workflow live in a few weeks, with a plain pass or fail test agreed before the build starts. If it fails, you learn that cheaply.

  • One workflow
  • Live in weeks
  • Plain pass or fail

Quality and oversight

How staff review AI output, when a human decides, and how you measure whether it is working after the launch week.

  • Review process
  • Human decision points
  • Ongoing measurement

✦ What you get: a clear plan, a working pilot or checklist, and ownership that stays with your people.

Open to contract and full-time remote roles.

AI data operations, evaluation design, human quality systems, and related program lead work.

Open to

Contract and full-time remote. I am most useful where data and evaluation work needs to become a repeatable program rather than a one-off push.

  • AI data operations
  • Evaluation design
  • Human quality systems
  • Program lead work

What I bring

Fifteen years designing technical education programs, then building the data and evaluation systems behind frontier model work.

  • Defined standards, calibration materials, and quality controls for programs involving 3,000+ contributors.
  • Ran client pilots from design through implementation and delivery.
  • Absorbed day-to-day contributor ops and delivery coordination so field engineers and PMs could stay focused on internal work.
  • Hands-on review of training data and AI-generated code
  • 15 years designing STEM and technical education programs
  • Founder, Learning Voyage LLC

✦ A useful first message names the role, the remote policy, and the timeline.

An evaluator can be useful without being entitled to decide.

I am preparing Bounded Review, an open research project on when automated review can support a decision and when a person must own it. It starts from three boundaries: the record sets the ceiling, uncertainty is routed rather than hidden, and authority is granted lane by lane with a named human owner. That work sits behind the business pilots and the evaluation roles above. It is not the front door.

✦ Independent open research, in preparation -- not a certification body, formal standard, or claim that a particular automated reviewer is accurate.

Selected Project Experience

Selected work has included managing data and evaluation projects supporting a frontier AI lab through an external AI data partner.

These engagements are complete; project details remain confidential.

I Learned the System From the Inside Out

I entered frontier-model data work in February 2024, creating difficult mathematics and physics training data. I moved into reviewing contributors, then auditing the decisions of other reviewers -- and from there into designing the standards, calibration processes, and workflows that govern the work.

In February 2025 my scope expanded into project development, program management, client collaboration, and quality leadership. I defined standards, calibration materials, and quality controls for programs involving 3,000+ contributors, ran client pilots from design through implementation and delivery, and absorbed day-to-day contributor ops and delivery coordination so field engineers and PMs could stay focused on internal work. Selected work has included managing data and evaluation projects supporting a frontier AI lab through an external AI data partner.

That progression is the point: each layer -- contributor, reviewer, reviewer of reviewers, systems designer, program lead -- changed what I could see. Learning Voyage is my independent practice and the home for this work as I extend it into specialized-data partnerships and proprietary-data systems.

your navigator

Owen Onderdonk

AI Data & Evaluation Program Lead · Founder, Learning Voyage LLC

  • Defined standards, calibration materials, and quality controls for programs involving 3,000+ contributors.
  • Ran client pilots from design through implementation and delivery.
  • Absorbed day-to-day contributor ops and delivery coordination so field engineers and PMs could stay focused on internal work.
  • Hands-on evaluation of training data and AI-generated code for frontier models
  • 15 years designing STEM and technical education programs
  • Bachelor's in Physics · Master's degree, Randolph College
Connect on LinkedIn →
3,000+ contributors on programs where I defined standards, calibration materials, and quality controls
pilots from design through implementation and delivery
15 yrs building STEM & technical education programs

The Layer Between a Hard Question and Dependable Evidence.

I work where research intent, domain expertise, data production, human judgment, and operational delivery have to become one coherent system. Each area is labeled so established work is never confused with direction or development.

Research data operations Established

Translate capability goals and failure modes into task definitions, contributor qualifications, calibration sets, quality controls, review workflows, and acceptance criteria.

Evaluation systems Established

Design evaluation logic, reviewer protocols, functionality tests, and claim-verification workflows that make decisions inspectable rather than merely plausible.

Human quality systems Established

Build calibration, second-level review, adjudication, and escalation for work that depends on expert judgment. Disagreement is a diagnostic signal, not automatically noise.

Data partnerships Current direction

Provider qualification, provenance review, pilot design, and acceptance criteria for specialized-data sourcing -- a trust and translation layer between model teams and data owners.

Proprietary data for AI Current direction

Help make internal knowledge and operational data usable as governed, testable inputs for retrieval, fine-tuning, evaluation, and agent workflows.

Agent environments & RL In development

Long-horizon tasks built from real or measured data, with hidden variation, explicit baselines, and reproducible grading.

The Problems I Want to Work On Next

These are research and operating questions, not a list of packaged services. I treat utility, provenance, rights, and measurement as separate claims that need separate evidence.

  1. What evidence should support a claim that a dataset is audited or fit for purpose?
  2. How should a model team specify data utility before acquisition rather than discovering its requirements after delivery?
  3. How should technical utility, provenance confidence, licensing scope, and maintainability be evaluated without collapsing them into one quality score?
  4. How can scarce domain expertise become reusable, interactive training signal through agent environments and reinforcement-learning tasks?
  5. Where should model-assisted quality review stop and calibrated human adjudication begin?
  6. How can organizations turn proprietary knowledge into governed, testable infrastructure for retrieval, fine-tuning, evaluation, and agent workflows?
  7. What open specifications and audit records would make private or paid data markets more transparent without requiring the underlying data to be public?

Start a Conversation

Businesses: email with the workflow you want to improve. Recruiters: email or message on LinkedIn with the role and whether it is contract or full-time remote. Either way, a useful first message names the objective and the part that is currently uncertain.

✦ Learning Voyage LLC is my independent practice and the publishing home for this work.