Physical AI

Egocentric & industrial video, ready to ship.

Real operators, real worksites, sourced through Apna's 300K+ business network. Captured, annotated and delivered end to end.

1,575hrs
OTS-ready footage captured
2,193
annotated video clips
100K+
factory & commercial sites
15.1TB
raw footage delivered
Pixel-art illustration of a worker wearing a head-mounted POV capture camera, repairing an engine part by hand at an arctic industrial workbench, with an Arctic Engine diagnostic screen in the background
Built by Apna and backed by
Tiger Global Sequoia Lightspeed Insight Partners GSV Owl Ventures Greenoaks Tiger Global Sequoia Lightspeed Insight Partners GSV Owl Ventures Greenoaks
The Foundation

A human network built to power physical AI

Apna’s workforce network spans 20+ streams, including LLMs, voice, physical and robotics, across white- and grey-collar India. This is the sourcing engine behind every hour in the catalog below.

115K+
field operators available
87K+
data annotators
28K+
QC reviewers
300K+
business / worksite partners

Non-residential only. Sourced and role-mapped across 27 environment categories, including factory, retail, F&B, healthcare and logistics. First sample batch ships within 2–4 weeks; full-capacity ramp takes 4–8 weeks.

Owned & ready

The OTS data catalog

Captured, timestamped and stored today across 13 sites, pulled from the internal “1,575 Hours” capture log. This is what ships first.

1,575hrs
total captured runtime
2,193
discrete video clips
15,060GB
raw storage footprint
37
distinct task types logged
SiteHoursClipsTasks
Press & Inspection Plant48964316
General Fabrication Facility40553921
Sorting & Logistics Center2152671
Final Assembly Line1941991
PCB Assembly Plant86612
Packaging Facility671051
Commercial Kitchen Site6523010
Secondary Sorting Depot22441
Food Prep & Packaging Site18604
Garment Stitching Unit7291
Cutting Station441
Kitchen & Housekeeping Site163
General Cleaning Site163
Split

90% industrial shop-floor (press, drilling, sorting, PCB assembly) · 10% commercial/domestic (kitchens, cleaning).

Reference landscape

Benchmarked against 25 third-party datasets, sized and priced separately from this catalog.

Task coverage

37 distinct task types logged, from power-press operation to pre-delivery inspection to food prep.

Market reference

Global commercial dataset landscape

A benchmark of 25 comparable third-party datasets, used to size and price the market. Shown separately from Arctic Engines’ own captured footage above.

25
reference datasets mapped
1,977
workers represented
27,414GB
total referenced footprint
Commercial

10 datasets

Residential

15 datasets

Note: one dataset (continuous factory-line monitoring, not episodic capture) makes up ~97% of this set's hours. The other 20 episodic datasets total 4,499.5 hrs across 20,491 clips.

How we deliver

Equipment, taxonomy & the delivery loop

One vendor, end to end — no point solutions to stitch together.

Available now

Mono video capture

iPhone · Android
GoPro Hero 13 Black
DJI Osmo Action 5 Pro
Insta360 GO 3S
Head / chest mounts

Built-in IMU (~200Hz) · exports MP4 + per-episode metadata JSON

2–4 week lead

Multimodal / stereo capture

Intel RealSense D435i
ZED 2i stereo
Unitree SV1-25
Stereo head-mounted rig

HW-synced, 200–400Hz IMU · exports MP4, MCAP, protobuf, depth, JSON

01

Environment Scoping

02

Sourcing & Profile Finalization

03

Video Capture + QC

04

Video Annotation + QC

05

Payroll

Retail & Consumer GoodsFashionRepair ServicesFood & BeverageConstruction & HardwareFood ProcessingFactoryHospitalityAutomotive & TransportIndustrial ManufacturingHealthcare & PharmacyCleaning & SanitationLaboratory / ScientificAgriculture & Farming+ 13 more

Recording confidence, by capture stream

StreamWhat it isDelivery confidence
Household + phoneSmartphone, no extra kit● High confidence
Household + cheap camAction cam at home● High confidence
Household + RGB-DDepth sensor at home◐ AE + partner
Operator + phoneWorker on factory floor● High confidence
Operator + pro rigWorker, head-mount GoPro● High confidence
Operator + multimodalWorker, sensor-fusion rig◐ AE + partner
Tele-op / robot-POVHuman controls robot◐ AE + partner
Proof of pipeline · frame-level

Every frame, synchronized: depth, pose, hands and objects together.

egocentric_trial is our dense, per-frame multimodal pipeline. Built to a partner’s frame-level spec and run end to end on real DJI POV capture, not a mockup.

21s excerpt from a 6.5-minute DJI POV trial capture: live YOLOv8 object boxes, hand/body keypoints and a zero-shot action classifier (left), per-frame depth from Depth-Anything-V2 (right). Every signal below is running simultaneously on this footage.

8-stage capture-to-MCAP pipeline
01

Ingest

Load the raw video.

02

Extract frames

Pull out frames to analyze.

03

Depth estimation

Work out how far away everything in the scene is.

04

Body pose

Track the body position of people in frame.

05

Hand pose

Track exact hand and finger positions.

06

Motion proxy

Estimate how the camera itself is moving.

07

Object + action annotation

Detect objects and label what action is happening.

08

Package it up

Bundle every signal into one synchronized MCAP file.

6.5min
raw DJI POV trial capture
1.57GB
synchronized MCAP output
8
signal channels per container
70+21
body / hand keypoints per frame
Depth (Depth Anything V2)70-pt body pose21-pt hand poseObject detection (YOLOv8)Action classification (CLIP zero-shot)

MCAP channels: rgb_frames · depth_maps · hand_keypoints · body_keypoints · imu_proxy · object_detections · action_segments · audio_transcript. Source: native 2688×1512 @25fps DJI capture, sampled at 4fps for the model stack.

Automated pipeline · temporal segments

Start/end timestamps + natural-language step descriptions — no frame labeling.

Built for a robotics-data client who wanted no per-frame labels, just a (start, end, description) triple per action step, generated automatically from raw egocentric footage.

Below is the pipeline’s actual output on 97 seconds of egocentric POV: a mechanic swaps a scooter battery in nine actions, time-segmented in natural language. Play the clip to watch the matching step highlight.

01
0.0s – 14.0s
A mechanic places a battery into the scooter's compartment and connects the positive terminal cable with a screwdriver.
02
14.0s – 19.5s
The mechanic connects the negative terminal cable to the battery.
03
19.5s – 29.0s
The mechanic disconnects the negative terminal cable from the battery using a screwdriver.
04
29.0s – 32.0s
The mechanic removes the old battery from the compartment and sets it on the floor.
05
32.0s – 44.1s
The mechanic picks up a new battery and prepares it by inserting the square terminal nuts.
06
44.1s – 53.7s
The mechanic places the new battery into the compartment and positions the terminal cables.
07
53.7s – 68.6s
The mechanic attaches the positive (red) terminal cable to the new battery and tightens the screw.
08
68.6s – 88.2s
The mechanic attaches the negative (black) terminal cable to the new battery and tightens the screw.
09
88.2s – 97.4s
The mechanic places the battery hold-down strap over the battery and tightens its screw.
Pipeline — chosen path (native-video VLM)
01

Ingest

Copy/probe source video, write metadata.json.

02

Segment (heuristic path)

Scene-cut + motion-pause detection for a local, free fallback — superseded by native segmentation below.

03

Segment + describe (VLM)

A vision-language model watches the full clip, proposes where each action starts and ends, and drafts a description per step in one pass.

04

Annotate + optimize

A lighter annotation pass on top of the VLM draft tightens boundaries and wording for accuracy, cost and speed.

Segmentation: native-video VLM pipeline · descriptions: VLM draft + human-reviewed annotation pass

Open the live demo →
Live products

See it running, not just described

Two live products already deployed on top of the pipelines above.

apnaEarn — home & field recording

Gig-worker app for sourcing egocentric footage at scale: pick a task, follow the brief, record, get paid per approved clip.

Making chaiMaking chai
Washing dishesWashing dishes
Folding laundryFolding laundry
Cooking sabziCooking sabzi

Egocentric Action Capture — real-time demo

Record a task live in-browser, see hand tracking overlay in real time, get an AI-generated action breakdown.

Point-of-view action camera

Hand tracking runs live in your browser (MediaPipe). Action segmentation runs via a vision-language model. Object/instance segmentation and tracking: coming next.

Sample footage

Media by site

Thumbnails and clip previews are cleared per-engagement. Tiles below map to the OTS catalog sites above.

Press & Inspection Plant
16 tasks · 643 clips
General Fabrication Facility
21 tasks · 539 clips
Sorting & Logistics Center
1 task · 267 clips
Final Assembly Line
1 task · 199 clips
PCB Assembly Plant
2 tasks · 61 clips
Packaging Facility
1 task · 105 clips
Commercial Kitchen Site
10 tasks · 230 clips
Secondary Sorting Depot
1 task · 44 clips
Food Prep & Packaging Site
4 tasks · 60 clips
Garment Stitching Unit
1 task · 29 clips
Cutting Station
1 task · 4 clips
Kitchen & Housekeeping Site
3 tasks · 6 clips
General Cleaning Site
3 tasks · 6 clips

Awaiting cleared sample media for this share. Full clip index (2,193 files) available on request under NDA.

Next steps

Start with a paid pilot.

Scope the brief → prove it on a batch with per-batch QC reporting → scale on the numbers. No ramp before the pilot clears.