Skip to content

    Humanoid Training Data

    Manipulation data from inside a working factory

    Either we capture it for you across a contracted network of production plants, or you send us footage you already hold. Both routes return the same structured, labelled, consented datasets.

    The network is the product

    Processing footage is a solved problem for a competent team. Getting inside a production plant that is willing, prepared and legally clear to share what happens on its floor is not. That is the part we have already built.

    Contracted and ready

    A large and growing network of production factories already signed up for data sharing, so a request turns into a capture run rather than a search for a pilot site.

    Hundreds of stations

    Multiple industrial sectors, hundreds of distinct work stations and process types, so a brief can be matched to the work rather than the other way round.

    New tasks come online quickly

    New plants and new task types are added inside the framework that already exists. There is nothing to renegotiate each time the brief changes.

    Variety inside one task

    The same task recorded across different plants, operators and part mixes, which is what makes a policy generalise instead of memorise.

    Sourced collection

    A collection engagement from your side of the table looks like this.

    1. 01

      You define the brief

      The task, the target skills, the conditions that matter and what counts as a good example. The brief is the contract for the run.

    2. 02

      We select the stations

      We match the brief to suitable stations across the network, usually across several plants, so the resulting data carries variety rather than one plant's habits.

    3. 03

      Capture runs on your terms

      On our capture kits, or on cameras and rigs you supply so the data arrives through the same optics your robot will use.

    4. 04

      You receive structured datasets

      Labelled, quality-scored and delivered in your format, with provenance and consent documentation attached.

    5. 05

      You iterate on the brief

      Review what came back, change what you are asking for, and run again. Nothing about the arrangement has to be renegotiated to do that.

    Structuring your footage

    If you already hold operator video, you do not need a collection programme. Send what you have and receive structured datasets back.

    No new hardware

    Works with the footage and cameras you already have. Nothing gets installed on a line for this.

    No workflow change

    Your teams keep working the way they work. The footage moves, the operation does not.

    The same outputs

    Identical fields and identical delivery formats to a sourced collection run, so the two can be mixed in one training set.

    Quality handled for you

    Unusable footage is filtered out and every delivered segment carries a quality score, so you are not training on the parts that should have been thrown away.

    What you receive

    The same set of structured outputs, whichever route the footage came in through.

    Temporal action segments
    Every discrete action in the footage, with precise start and end boundaries.
    Task and action labels
    A task label and an action type on every segment, so a segment is usable without watching the video.
    Object and hand masks
    Segmentation masks for the hands and for the objects being handled.
    Interaction and contact events
    Where and when contact is made with an object, and when it is released.
    Hand and wrist trajectories
    Continuous hand and wrist paths through each action.
    Spatial motion data
    Derived from standard 2D video. No depth sensors and no motion capture suits are required on the line.
    Per-segment quality score
    A quality score on every segment, with unusable footage filtered out automatically before delivery.
    Privacy flags and redaction
    Sensitive segments are flagged and redaction is applied before a dataset leaves us.

    Delivery formats

    HDF5LeRobotRLDSyour schema

    Datasets are delivered in HDF5, LeRobot or RLDS. If your team works to its own schema, we deliver to that instead.

    Quality and privacy handling

    Quality scoring on every segment

    Each delivered segment carries a quality score, so you can filter to the threshold your training run needs.

    Unusable footage removed

    Occluded, truncated and unrecoverable material is filtered out automatically rather than shipped and left for you to find.

    Privacy flags travel with the data

    Segments needing downstream care are flagged, so your own review can see what was treated and why.

    Redaction before delivery

    Faces and identifying features are handled before a dataset leaves us, not after it lands.

    Read the full data governance position

    Why factories take part

    Participating plants are not doing anyone a favour. Every one of them receives a process improvement report and an ergonomic analysis report at no cost, produced from the same operator footage, plus access to our commercially deployed process analysis platform. Operators are compensated for their participation, and capture is designed not to disturb the line.

    More on our process intelligence platform for manufacturers(opens in a new tab)

    Scope a pilot on a sample

    Send us a sample of footage, or a task you want captured. We will come back with what we can deliver, in what format, and what the consent position looks like.

    Talk to our team