Directed multimodal capture
It happened.
Nobody filmed it
Second Orbit produces multimodal training and evaluation data for teams building models that work in the physical world. We design the situation, capture it with an expert in the loop, structure the ground truth, and measure what the model learned
Our team has delivered multimodal collection, annotation and evaluation programs one week from brief to locked spec.
What the asks have in common
Not rare
Unrecorded
A volume vendor will sell you a thousand hours of someone cooking dinner competently. None of it reaches the forty seconds where it goes wrong — because nobody was running two cameras and an annotation schema when it did.
FOUR STAGES, ONE LOOP
A labeling vendor can’t design the situation.
A capture shop can’t tell you whether it worked.
These belong together because each one decides the next.
What the written protocol specifies, who approves it, and what is deliberately left open.
How a protocol gets written → 02Capture the evidenceThe capture configuration, the devices, and the two modes an expert-led session runs in.
See the capture configuration → 03Structure the ground truthThe six annotation families, the review passes, and how the rubric is agreed.
The six annotation families → 04Measure what the model learnedHow an evaluation is built on the same ground truth, and what it returns.
How an evaluation is built →ADVISORY
Some work doesn’t fit a standard data program. AI strategy, coaching and enablement, use-case prioritization, custom applications, and collection scope that falls outside the usual shape. We take a small number of these each year.
Advisory →Capture
Environments, not stock categories.
Every environment below is a real working setting with a domain expert on the line.
Filter by what your model has to survive.
To view more click here
The specification
You write it.
We do not have a house format.
Select what your model actually needs. The sheet on the right is the delivery specification that would be produced — it is the same document our capture leads work from.
Illustrative. Volumes, rates and schema keys are agreed per programme.
How they are produced
Directed is not scripted
The fault is arranged in advance. The diagnosis is genuine. There is no second take — which is the whole point, and the reason this cannot be scraped.
CAPTURE CONFIGURATION
Every program is a designed situation with a checkable outcome. What gets designed into it depends entirely on what your model is failing at. An induced fault in one program, a wayfinding decision in the next, a handoff between two people in the one after that.
Head-mounted, fixed wide, overhead, close-up — as many as the task needs, hardware-synced. Frame offset recorded, not estimated.
Not an actor. The person diagnosing is the person who does this for a living.
Whatever your model is failing at, we arrange for it to occur — an induced fault, a substitution, an interruption, an ambiguous instruction, a handoff between two people. What happens next is not arranged.
Action boundaries, object tracks and the verbal channel, aligned to the same clock.
Consent, release, capture conditions, labeler identity, schema version.
Location class, heading and dwell where the task is spatial. Device, lighting, camera placement and sync offset on every clip. Attached at capture, not reconstructed afterward.
How a program runs
Six steps, and you are in every one of them.
You describe one behavior your model gets wrong. We come back with a written protocol — the objective, the environment, the participants, the devices, the coverage, and the conditions we intend to induce. You approve it before anyone is booked.
We find people who do this for a living, in the setting where they do it — a technician at their own bench, a cook in a working kitchen. Consent and release are agreed before anything is arranged, documented per contributor and recorded against every clip.
Directed is not scripted. The condition is arranged in advance — an induced fault, a substitution, an interruption, a handoff between two people. The participant is briefed on the situation, never on the outcome. What happens next is genuine, and there is no second take.
As many synchronized views as the task needs — head-mounted, fixed wide, overhead, close-up. We have run six on a single work area, all on one clock, with frame offset recorded rather than estimated. A domain expert watches the feed and corrects in real time.
Your annotation schema, not ours. Six annotation families, from pixel-level spatial labels to cross-modal alignment, with boundaries marked to the frame. Multi-pass review, gold-standard validation, and inter-annotator agreement measured per task family.
We build the evaluation on the same ground truth we collected — answer-blind pipelines, human gold standards, rubric scoring. Or you run your own and tell us what it showed. Either way, what the model still gets wrong becomes the scope for the next round.
What you receive
Media, labels, transcript, manifest
Delivered to your schema, not ours. Every clip carries its own provenance record, and every label is traceable to the person who made it.
CAPTURE APPARATUS
Every view at original capture rate, no re-encode above your ceiling. Frame offset recorded per clip, so the views stay on one clock.
Action segments, object tracks, and the failure interval marked — to the frame, against your annotation schema. Tracks held through occlusion and re-entry.
Diarized, timestamped and aligned to frame. The participant’s narration and the moment of correction are tagged, not left for you to find.
One record per clip: device, sync offset, capture conditions, consent and license terms, annotator identity and schema version. Attached at capture.
Held-out by scenario, not by random frame — so the split means something. The held-out set is agreed with you before capture, not carved out afterward.
Multi-pass review against the rubric agreed with you in the pilot, gold-standard validation, and inter-annotator agreement. Sessions that miss it don’t ship.
Evaluation case study
We built the evaluation
before anyone had built the dataset
A frontier AI research team wanted to know whether today’s vision-language models can act as real-time instructors — watching someone work through a physical task and coaching them the way a human expert would. No off-the-shelf benchmark measures that. It needed purpose-built collection and a purpose-built evaluation.
Delivered by this team at The Blue Dot Labs. Client anonymized at their request.
Operations
US capture where context matters.
Global operations where scale does.
An office and capture studio in Sunnyvale, California. On-shore capture when the contract requires it. Global operations when throughput does.
QUALITY SYSTEMS
Multi-pass QA against a rubric agreed before collection · gold-standard validation · inter-annotator checks · human-in-the-loop review at every stage · a written rejection standard, so what disqualifies a session isn’t decided after the fact.
CREATORS
Our capture network is paid at rates we’re comfortable publishing.
Provenance
Every frame has a paper trail.
Directed capture means we know exactly who is in the frame, what they agreed to, and who labelled what. That is the part scraped data cannot give you.
About
We’ve done this before
Second company, same team. Between us, two decades each in the two disciplines this work actually runs on: video and geospatial data at consumer scale, and global operations delivered to a written standard.
Our first company together was The Blue Dot Labs, which we grew to multi-million-dollar revenue in about thirty months. This is the spinoff of our data and services business.
This is the Second Orbit
Start here
Send us the forty seconds you cannot find.
Describe one behavior your model gets wrong. We come back with a capture plan, an annotation schema and a sample — before any volume is discussed.
PRIMARY · SECONDARY · PROCUREMENT
WHAT WE DON’T DO
We don’t do text-only RLHF.
We don’t sell scraped or resold footage.
We don’t run unsupervised collection at volume and call it directed.
CAPABILITIES
Advisory Design the situation Capture the evidence Structure the ground truth Measure what the model learnedFOR CREATORS
Join the networkSecond Orbit · Directed multimodal training data
hello@secondorbit.aiCapabilities
Design the situation
The research question becomes a written protocol. Nothing is left to what happens to occur — and nothing is decided without you.
WHAT THE PROTOCOL SPECIFIES
The behavior your model gets wrong, written as something that can be filmed and checked.
A real working setting, not a set. Location class, lighting and constraints named in advance.
People who do this for a living. Consent and release agreed before anything is arranged.
Head-mounted, fixed wide, overhead, close-up — chosen as a variable rather than a default.
How many views the task needs, and what each one is there to see.
The fault, substitution, interruption, ambiguous instruction or handoff we arrange. What happens next is not arranged.
The protocol is the deliverable of this stage. You approve it before anyone is booked, and it comes back revised after the pilot.
Capabilities
Capture the evidence
Expert-led sessions, reactive and proactive. Real devices, chosen as a variable rather than a default. We have run six synchronized views on a single work area.
THE CAPTURE CONFIGURATION
Head-mounted, fixed wide, overhead, close-up — as many as the task needs, hardware-synced. Frame offset recorded, not estimated.
The person who does this for a living watches the first-person feed and corrects in real time. Not an actor.
Whatever your model is failing at, we arrange for it to occur — an induced fault, a substitution, an interruption, an ambiguous instruction, a handoff between two people. What happens next is not arranged.
Action boundaries labeled to the frame rather than the second, against your taxonomy.
One record per clip: consent, release, capture conditions, annotator identity and schema version.
Location class, heading and dwell where the task is spatial. Device, lighting, camera placement and sync offset on every clip. Attached at capture, not reconstructed afterward.
Capabilities
Structure the ground truth
Six annotation families, from pixel-level spatial labels to cross-modal alignment. Your schema, exported — not ours.
THE SIX FAMILIES
Masks and boxes at pixel level, on the objects that matter.
Action boundaries to the frame, against your taxonomy.
Held through occlusion and re-entry.
Time-aligned, with the correction moment tagged.
Speech, action and object tied to the same clock.
Scored against the rubric agreed in the pilot.
QUALITY CONTROL
Multi-pass review against the rubric agreed in the pilot, gold-standard validation, and inter-annotator agreement measured per task family. Sessions that don’t clear the rubric don’t ship — and the rubric is written down, in your language, before capture starts.
Capabilities
Measure what the model learned
Answer-blind evaluation pipelines, human gold standards, rubric scoring across accuracy and response quality. Standalone, or alongside a collection program.
HOW THE EVALUATION IS BUILT
The judge never sees the reference answer. What comes back is a score you can defend to a reviewer.
Built by the same domain experts who ran the capture, on the same ground truth.
Accuracy and response quality scored separately, against the rubric agreed in the pilot.
Not by random frame — so the split means something.
Capabilities
Advisory
Second Orbit primarily produces multimodal training and evaluation data for physical-world AI. Alongside that, we take on a small amount of advisory work each year — much of it with organizations we already know, and much of it starting as a question about what to collect.
deciding which use cases are worth building, and in what order.
working with leadership teams on how AI changes what their organization does and how it operates.
assessing candidate applications against feasibility, data availability and value.
building the thing, where building it is the right answer.
putting AI into a process that already exists rather than around it.
collection or annotation work that doesn’t fit the usual shape.
We advise a small number of organizations each year, so we’re selective about fit. If you think there’s one, write to us.
Capabilities
Every program starts with a pilot, sized to the question.
The research question becomes a written protocol. Nothing is left to what happens to occur — and nothing is decided without you.
WHAT COMES BACK
We’ve run 10–15 sessions in a week, and three rounds of three with fast feedback between rounds.
One environment, one condition, all views, fully labeled.
What we changed after the first round, and why.
Head-mounted, fixed wide, overhead, close-up — chosen as a variable rather than a default.
In your language, written down before capture starts.
Whether to scale — and if not, what would have to change first.
Legal
Consent, privacy and provenance
How we handle information about visitors to this site, and the terms that apply to using it.
Data captured inside a program is covered separately.
All capture is first-party. Nothing is scraped, resold, or licensed in from a third party. Consent is obtained at the point of capture rather than sought retrospectively, and is documented per contributor — named, scoped, and specific to the use the footage is collected for. The license terms are recorded against each asset, withdrawal terms are recorded alongside the release, and age verification is carried out where required.
PII detection and redaction workflows run across every delivery. Treatment is configurable per program for faces, screens, documents and identifying signage, and is agreed in the protocol before capture rather than negotiated after it. Anything we cannot resolve is flagged rather than quietly delivered, and the redaction flags ship with the footage so you can see what was treated and what was not.
Secure storage, access controls scoped per program, and encrypted delivery.
Every delivery ships with a provenance record, per asset: the consent and release, the capture conditions, the direction record, and the annotator identity against the schema version. Together these form a chain of custody per clip — from the release signed before capture through to the person who reviewed the final label.
QUESTIONS
Contact hello@secondorbit.ai for consent scope, processing terms or security posture.
Legal
Privacy and terms
How we handle information about visitors to this site, and the terms that apply to using it.
Data captured inside a program is covered separately.
To reply to you, and to run the site. We don’t sell it, and we don’t add you to a mailing list you didn’t ask for.
To reply to you, and to run the site. We don’t sell it, and we don’t add you to a mailing list you didn’t ask for.
Our scheduling and email providers, because a calendar invite and a reply have to come from somewhere. They are listed by name below.
[TO CONFIRM — a retention period, and what happens to an enquiry that doesn’t become a project.]
Write to hello@secondorbit.ai and we will tell you what we hold about you, correct it, or delete it.
[TO CONFIRM — site terms, acceptable use, and the limits of what this page promises.]
[TO CONFIRM — jurisdiction. The Second Orbit, Inc.]
QUESTIONS
Contact hello@secondorbit.ai for consent scope, processing terms or security posture.
Join-the-network
Join us
We work with creators across the US, India and the Philippines on real-world capture. If that’s you, we’d like to hear from you.
WHAT THE WORK IS
Kitchens, workshops, warehouses, streets — wherever the task actually happens.
We look for people who do the task for a living, not performers.
Our capture network is paid at rates we’re comfortable publishing.
GET IN TOUCH
Or write to hello@secondorbit.ai We read everything, and we reply to everyone we can work with.
Error 404
Page Not Found