Real operators, real worksites, sourced through Apna's 300K+ business network. Captured, annotated and delivered end to end.
Apna’s workforce network spans 20+ streams, including LLMs, voice, physical and robotics, across white- and grey-collar India. This is the sourcing engine behind every hour in the catalog below.
Non-residential only. Sourced and role-mapped across 27 environment categories, including factory, retail, F&B, healthcare and logistics. First sample batch ships within 2–4 weeks; full-capacity ramp takes 4–8 weeks.
Captured, timestamped and stored today across 13 sites, pulled from the internal “1,575 Hours” capture log. This is what ships first.
| Site | Hours | Clips | Tasks |
|---|---|---|---|
| Press & Inspection Plant | 489 | 643 | 16 |
| General Fabrication Facility | 405 | 539 | 21 |
| Sorting & Logistics Center | 215 | 267 | 1 |
| Final Assembly Line | 194 | 199 | 1 |
| PCB Assembly Plant | 86 | 61 | 2 |
| Packaging Facility | 67 | 105 | 1 |
| Commercial Kitchen Site | 65 | 230 | 10 |
| Secondary Sorting Depot | 22 | 44 | 1 |
| Food Prep & Packaging Site | 18 | 60 | 4 |
| Garment Stitching Unit | 7 | 29 | 1 |
| Cutting Station | 4 | 4 | 1 |
| Kitchen & Housekeeping Site | 1 | 6 | 3 |
| General Cleaning Site | 1 | 6 | 3 |
90% industrial shop-floor (press, drilling, sorting, PCB assembly) · 10% commercial/domestic (kitchens, cleaning).
Benchmarked against 25 third-party datasets, sized and priced separately from this catalog.
37 distinct task types logged, from power-press operation to pre-delivery inspection to food prep.
A benchmark of 25 comparable third-party datasets, used to size and price the market. Shown separately from Arctic Engines’ own captured footage above.
10 datasets
15 datasets
Note: one dataset (continuous factory-line monitoring, not episodic capture) makes up ~97% of this set's hours. The other 20 episodic datasets total 4,499.5 hrs across 20,491 clips.
One vendor, end to end — no point solutions to stitch together.
Built-in IMU (~200Hz) · exports MP4 + per-episode metadata JSON
HW-synced, 200–400Hz IMU · exports MP4, MCAP, protobuf, depth, JSON
Environment Scoping
Sourcing & Profile Finalization
Video Capture + QC
Video Annotation + QC
Payroll
| Stream | What it is | Delivery confidence |
|---|---|---|
| Household + phone | Smartphone, no extra kit | ● High confidence |
| Household + cheap cam | Action cam at home | ● High confidence |
| Household + RGB-D | Depth sensor at home | ◐ AE + partner |
| Operator + phone | Worker on factory floor | ● High confidence |
| Operator + pro rig | Worker, head-mount GoPro | ● High confidence |
| Operator + multimodal | Worker, sensor-fusion rig | ◐ AE + partner |
| Tele-op / robot-POV | Human controls robot | ◐ AE + partner |
egocentric_trial is our dense, per-frame multimodal pipeline. Built to a partner’s frame-level spec and run end to end on real DJI POV capture, not a mockup.
21s excerpt from a 6.5-minute DJI POV trial capture: live YOLOv8 object boxes, hand/body keypoints and a zero-shot action classifier (left), per-frame depth from Depth-Anything-V2 (right). Every signal below is running simultaneously on this footage.
Ingest
Load the raw video.
Extract frames
Pull out frames to analyze.
Depth estimation
Work out how far away everything in the scene is.
Body pose
Track the body position of people in frame.
Hand pose
Track exact hand and finger positions.
Motion proxy
Estimate how the camera itself is moving.
Object + action annotation
Detect objects and label what action is happening.
Package it up
Bundle every signal into one synchronized MCAP file.
MCAP channels: rgb_frames · depth_maps · hand_keypoints · body_keypoints · imu_proxy · object_detections · action_segments · audio_transcript. Source: native 2688×1512 @25fps DJI capture, sampled at 4fps for the model stack.
Built for a robotics-data client who wanted no per-frame labels, just a (start, end, description) triple per action step, generated automatically from raw egocentric footage.
Below is the pipeline’s actual output on 97 seconds of egocentric POV: a mechanic swaps a scooter battery in nine actions, time-segmented in natural language. Play the clip to watch the matching step highlight.
Ingest
Copy/probe source video, write metadata.json.
Segment (heuristic path)
Scene-cut + motion-pause detection for a local, free fallback — superseded by native segmentation below.
Segment + describe (VLM)
A vision-language model watches the full clip, proposes where each action starts and ends, and drafts a description per step in one pass.
Annotate + optimize
A lighter annotation pass on top of the VLM draft tightens boundaries and wording for accuracy, cost and speed.
Segmentation: native-video VLM pipeline · descriptions: VLM draft + human-reviewed annotation pass
Open the live demo →Two live products already deployed on top of the pipelines above.
Gig-worker app for sourcing egocentric footage at scale: pick a task, follow the brief, record, get paid per approved clip.




Record a task live in-browser, see hand tracking overlay in real time, get an AI-generated action breakdown.
Hand tracking runs live in your browser (MediaPipe). Action segmentation runs via a vision-language model. Object/instance segmentation and tracking: coming next.
Thumbnails and clip previews are cleared per-engagement. Tiles below map to the OTS catalog sites above.
Awaiting cleared sample media for this share. Full clip index (2,193 files) available on request under NDA.
Scope the brief → prove it on a batch with per-batch QC reporting → scale on the numbers. No ramp before the pilot clears.