Factory video in, training data out

    The most expensive part of training a humanoid is labeling the data. Khenda does it automatically: upload your operator video, get back structured datasets your models can learn from.

    The pipeline

    How it works

    1

    You upload video

    Standard footage from the cameras you already have. No depth sensors, no motion capture suits, no special rig.

    2

    We find the actions

    The pipeline detects and timestamps every discrete action: grasps, placements, tool uses, inspections, across multiple workers and angles.

    3

    We extract 3D motion

    From plain 2D video we infer joint angles, hand trajectories, and contact points, the spatial data robot policies learn from.

    4

    You get a dataset

    Structured, labeled, and packaged for Vision-Language-Action training. Ready to drop into your pipeline.

    What's in the dataset

    Action recognition and temporal segmentation
    3D kinematics and pose estimation from 2D video
    Object interaction and contact events
    VLA-ready output compatible with leading frameworks
    Automatic quality filtering, no manual annotation
    Scales to thousands of hours of footage

    See it run on your footage

    Send us a sample of factory video and we'll turn it into a structured, training-ready dataset so you can see the output format and quality firsthand.

    Book a Demo