Experiment cockpit
Variants, evidence, votes, and dissent share one surface.
Shipwright / collaborative A/B testing
Shipwright puts Variant A and Variant B in one decision room. It captures both in the browser, gives AI judges and human reviewers the same evidence, keeps dissent visible, and promotes a winner only when the gates pass.
Example decision
Which direction should reach the agent?
Variants, evidence, votes, and dissent share one surface.
The recommendation lands first; deeper evidence follows.
The product loop
Most review starts after someone has already committed to a direction. Shipwright moves product judgment earlier, while a competing version is still cheap to compare and human feedback can still change the work.
Define the question, hypotheses, audience, expected signal, and Variant A/B allocation.
Render each direction and collect Playwright screenshots and interface signals as evidence.
AI models and human reviewers choose, explain why, and leave minority views visible.
Only a winner with enough evidence becomes a worker prompt. Missing proof stops promotion.
AI + human collaboration
Model judges return a selection, confidence, rationale, and the evidence they used. Human review adds product context, must-fixes, acceptance, or dissent. Shipwright keeps both instead of flattening the decision into one unexplained score.
The comparison surface exposes the decision trail without requiring raw logs.
Variant B is faster to scan. Keep that advantage in the implementation checklist.
Choose A, but retain the runner-up’s lighter first impression.
Surface Grill / human review
Surface Grill anchors must-fixes, accepted feedback, and dissent to normalized coordinates on a captured variant. Each review creates a digest-bound revision packet for the next pass; the review surface itself never edits the source project.
Decision gates
Shipwright separates “this direction won” from “there is enough proof to promote it.” A missing screenshot or incomplete handoff keeps the result in needs-more-signal—even when the weighted score favors one variant.
Promotion checklist
The pre-merge handoff
Once the gates pass, Shipwright turns the winning direction and the strongest dissent into a bounded work order. The next agent sees what to change, what not to lose, and how the result will be verified before merge.
Decision → implementation
The selected variant, runner-up risk, review notes, source boundary, acceptance criteria, and next verification command travel together. Review happens before commitment to the direction—and remains available when the patch comes back for approval.
Proof Story / after release
Once the selected change is on main and its public checks pass, Shipwright builds a short release story from that exact receipt. Every scene stays bound to verified screenshots, source revision, checks, and public URL; stale or missing evidence blocks export.
Release → understanding
The first 15-second Proof Story uses a fresh capture from Shipwright’s production origin. It ships with the approved claims, source and deployment receipts, capture hash, poster, share copy, release note, and a machine-readable provenance manifest—without a model, music, telemetry, simulated interaction, or automatic posting.