Methodology

Estimate priorities.
Measure learning value.

A candidate score helps decide what to test. A controlled experiment measures whether the selected batch helped this model, on this task, under a defined protocol.

01 / Establish the context

Usefulness depends on the decision.

Define the model and checkpoint, task, label conventions, critical conditions, metrics, acceptable regressions, and budget before evaluating candidates.

Record permissions and source identities. Keep training, development, calibration, and final evaluation partitions separate. Source-group separation matters: related samples should not leak between partitions.

Use development data to investigate model weaknesses. Use separate calibration data to choose operating thresholds when appropriate. Reserve independent real evaluation for the final test.

A usefulness estimate is conditional on a model and task. It is not a permanent value attached to a sample.

02 / Establish the evidence

Valid files are only the beginning.

Technical checks cover required metadata, hashes, lineage, permissions, partition consistency, exact duplicates, and label structure.

Semantic review asks whether the requested scenario is actually present, visible objects are plausible, and labels agree with the output. For driving detection, two-dimensional boxes must be checked against the generated RGB frames.

Current semantic validation uses a human review rubric, recording reviewer, timestamp, rubric, and evidence. Missing evidence remains pending or unknown.

A prompt saying “night” does not verify a night scene. Source-scene labels cannot simply be assumed correct after generation.

03 / Prioritize a hypothesis

Choose candidates worth testing.

The current estimator transparently combines reviewed quality, candidate errors under the current model, development-set weaknesses, and redundancy signals. Redundancy checks identify exact duplicate content; semantic novelty is not yet established.

Selection follows declared batch-count and estimated-cost limits, source-group caps, and optional coverage constraints. Exploration can reserve room for less familiar candidates.

If eligible data cannot satisfy the constraints, the appropriate result is insufficient evidence or an incomplete selection.

Scores estimate priority. They do not establish exact causal contributions or prove that an individual example improves the model.

04 / Compare under a frozen protocol

Learning claims require a controlled test.

Register the experiment before final evaluation. Freeze the base model, training inputs, settings, seeds, metrics, evaluation snapshot, and decision thresholds.

Compare legitimate baselines.

  • Real-only adaptation
  • Real plus random synthetic selection
  • Real plus inexpensive quality filtering
  • Real plus model-targeted synthetic selection

A count-matched comparison tests whether selection changes learning outcomes under a matched training protocol. A separate equal-total-cost comparison is needed to support budget-efficiency claims.

Evaluate on independent real data.

Report target-condition outcomes, clean-condition regressions, per-seed results, and uncertainty. Artificial fixtures verify software execution; they do not demonstrate real learning gains.

05 / Account for the complete investment

Count cost. Show uncertainty.

Include generation failures, annotation, review, and training costs where relevant. Compare recorded costs under explicit limits rather than presenting a sample-count match as a budget-efficiency result.

Retain per-seed outcomes and uncertainty estimates. Report limitations, failed checks, and missing evidence alongside positive results.

“More learning from every data dollar” is our objective. It is not a quantified performance claim.

06 / Preserve a useful fallback

The result should guide an action.

A controlled comparison can justify investigating another batch, keeping a simpler selection policy, requesting more review, collecting missing evaluation data, or pausing expansion.

  • Evidence fallback: a validity and coverage audit when labels or training access are insufficient for a learning claim.
  • Selection fallback: retain random or inexpensive quality filtering when targeted selection has not shown an advantage.
  • Execution fallback: alternative providers are a future direction; automatic provider failover is not currently implemented.

The intended long-term loop uses outcomes to guide the next allocation. The current workflow requires deliberate experiment and review decisions.

An early-stage product

What exists today.

A working local pipeline supports task contracts, manifest and permission checks, lineage, review decisions, prediction validation, calibration, candidate scoring, selection baselines, experiment registration, outcome comparison, uncertainty estimates, and cost checks.

Detection, classification, and extraction evaluation paths exist. This infrastructure portability does not prove a selection policy transfers successfully between tasks.

What remains to be demonstrated.

Actual Cosmos generation, reviewed natural-data batches, controlled real-data training comparisons, complete delivery economics, and cross-model selection-policy validation are pending.

A hosted dashboard, managed cloud execution, physics-aware validation, and autonomous allocation remain future directions.

Discuss a pilot