
Deployment Evaluation
Know how your humanoid is really doing
A robot that finishes the cycle is not the same as a robot that is working. Deployment Evaluation compares what the humanoid does on the station against the human baseline, step by step.
The evaluation loop
How an evaluation runs
- 1
Observe the station
The deployed humanoid is observed on the line from standard video, the same way an operator working that station would be.
- 2
Compare against the human baseline
Each cycle is matched step by step against the human reference drawn from the operator data we already structured for that station.
- 3
Surface slow and failed steps
Every step is scored for success and for timing, so a step that is drifting slower or quietly failing shows up as a finding rather than as downtime.
- 4
Feed the edge cases back
The deviations and edge cases the station produces become new training data, so the next policy version has seen them.
This is the part of the loop worth being explicit about: the human baseline a robot is scored against comes from the same operator data we structure for training, and the edge cases the line produces go back into it. The two products are the same data seen from opposite ends.
What you get
- Step-level success and failure detection against a human reference
- Cycle-time and slowdown tracking per step
- Slow and failing steps surfaced before they turn into downtime
- Edge cases captured on the line, structured for retraining
- Success criteria usable as a reward signal
- Works from standard video, with no extra sensors on the station
Put a deployment under the microscope
Tell us about the station and the robot working it. We will show you what an evaluation reports and how the findings come back.
Talk to our team

