Question layer
Hooks, skills, and explicit decision scaffolds.
Test whether surfacing a missing requirement or tradeoff changes the next action and reduces wrong-work.
Crux Research
Interaction frameworks already know how to pause. The harder problem is deciding whether a question is necessary, which question changes the plan, and when asking is worth the interruption.
A staged program
Each stage has a different evidence bar. We label the current product, the next public experiment, and the longer-term engine separately so a roadmap claim never masquerades as a shipped feature.
Question layer
Test whether surfacing a missing requirement or tradeoff changes the next action and reduces wrong-work.
Calibrated asking
Publish protocols, adverse results, and limits. Measure question behavior before claiming improvements to task outcomes.
Mechanistic control
Keep enforcement outside the model and test containment under manipulation, drift, and steering.
Research rules
Claims boundary
The commercial link
With explicit consent, preview workflows can teach us which questions arrive too early, too late, or not at all. Customer data remains customer data; reusable learning starts with aggregate, privacy-preserving measurements.
Record the evidence available before the agent proposed its action.
Measure the plan, artifact, or decision before and after the answer.
Compare information gain and avoided rework against delay and user cost.
Refine the skill or model only after the outcome can be replayed and audited.
Follow the work
The first public releases will focus on integrations, evaluation methods, and reproducible baselines. Sensitive model-control work remains private until review and filing gates are complete.