Crux Research

How should an agent know when to ask?

Interaction frameworks already know how to pause. The harder problem is deciding whether a question is necessary, which question changes the plan, and when asking is worth the interruption.

A staged program

Product first. Research honestly.

Each stage has a different evidence bar. We label the current product, the next public experiment, and the longer-term engine separately so a roadmap claim never masquerades as a shipped feature.

Now · Private preview

Question layer

Hooks, skills, and explicit decision scaffolds.

Test whether surfacing a missing requirement or tradeoff changes the next action and reduces wrong-work.

Next · Open-model preview

Calibrated asking

Read and perturb ask-related signals under controlled conditions.

Publish protocols, adverse results, and limits. Measure question behavior before claiming improvements to task outcomes.

Roadmap · Experimental

Mechanistic control

Internal signals inform bounded ask, act, or halt decisions.

Keep enforcement outside the model and test containment under manipulation, drift, and steering.

Research rules

Make failure legible.

  • Pre-register the decision a result will unlock
  • Separate question frequency from task success
  • Report null, adverse, and over-asking results
  • Keep synthetic evidence distinct from customer traction
  • Use external benchmarks without implying independent replication
  • Publish only after intellectual-property and privacy review

Claims boundary

What this page does not claim

  • That Crux currently understands user intent from model activations
  • That mechanistic steering improves end-to-end task success
  • That an experimental governor is production-ready
  • That stronger control automatically improves safety or utility
  • That private research artifacts are ready for public release

The commercial link

Every deployment is an evaluation opportunity.

With explicit consent, preview workflows can teach us which questions arrive too early, too late, or not at all. Customer data remains customer data; reusable learning starts with aggregate, privacy-preserving measurements.

Was there a material gap?

Record the evidence available before the agent proposed its action.

Did the question change anything?

Measure the plan, artifact, or decision before and after the answer.

Was the interruption worth it?

Compare information gain and avoided rework against delay and user cost.

What should the next policy do?

Refine the skill or model only after the outcome can be replayed and audited.

Follow the work

Open protocols before grand claims.

The first public releases will focus on integrations, evaluation methods, and reproducible baselines. Sensitive model-control work remains private until review and filing gates are complete.

See project status