TOUCHSTONE v3 · SHEET 02 · REV 0.2 · SPECIFIED, PARTLY VALIDATED, UNBUILT
Touchstone
the force layer for physical AImeasured · certified · portable
Robots learned to see from a billion photographs. Nothing, ever, recorded how hard a hand pressed.
Touchstone is an instrument, not a dataset vendor. A person works normally; a sleeve on the forearm reads the muscles that move the hand; a handheld standard tells the sleeve what a newton is. What comes out is not footage. It is force in newtons on a known object at a known instant, with a stated interval around it — and because it records the job the object experienced rather than the motion one particular hand made, the same demonstration can be recompiled onto a body that looks nothing like yours.
This sheet is the long version. Everything freely shareable is here, including the parts that went wrong. Where a mechanism is deliberately held back, the sheet says so rather than glossing over it.
- measured
- newtons at the contact, not inferred from pixels
- bare-handed
- nothing on the palm or the fingertips
- portable
- one demonstration compiles onto many robot bodies
- certified
- every number traceable back to a physical reference
- honest
- and every number carries how far it can be trusted
Force is invisible.
That is the entire problem.
A frame of someone holding a cup and a frame of someone crushing it are the same picture, right up until it is too late. Every manipulation failure that embarrasses robots is a force failure: the crushed egg, the dropped glass, the connector that never seats. None of it is visible in video, so none of it is in the training data.
A tactile glove does not close the gap either. It gives a pressure pattern in relative units — rich spatial coverage, no newtons, and no stiffness channel at all. The missing thing is not resolution. It is the unit.
A correction we made to our own earlier position, and keep making in public: we used to say gloves corrupt the demonstration by construction. That was too absolute. Glove data can train force-sensitive policies, and at least one glove was engineered specifically to stay thin and preserve dexterity. Glove interference is a graded concern, not a disqualifier. Our real advantage over gloves is the stiffness channel and the calibration — not the bare fact of leaving the palm uncovered.
And underneath it, a second problem
A human hand has roughly twenty-three kinematic degrees of freedom and five deformable pads. A parallel jaw has two rigid jaws on one axis. Copying joint angles across that difference is close to meaningless, which is why retargeting research keeps fragmenting into one bespoke solution per robot.
The two problems have one answer. Stop recording the hand. Record what the hand did to the object, and how the arm was regulated while doing it. Force on the object does not care what applied it. Stiffness is a command any torque-controlled robot can render. Contact events are the subgoals of manipulation for any body at all. All three are measurable from the muscle side — and two of them only from the muscle side. That single fact is the entire strategic reason to build a wearable rather than anything else.
Three physical things, and nothing on the hand.
A forearm sleeve. A camera the customer already owns. A handheld calibration artifact. The palm, the fingers and the fingertips stay bare. Robots never wear any of it — the sleeve's signals are privileged training information, not a runtime dependency.

- The sleeve
- Measurement equipment in the shape of performance apparel. Skin-contact sensing on the volar side, rigid mass on the dorsal side, one small status light and no display.
- The camera
- Customer-supplied and egocentric, exactly as the open capture rigs already do it. The Record has to be a strict superset of what a customer already collects, or nobody adopts it.
- The artifact
- A handheld object of known, stable, traceable properties. It is what makes every number on the sheet mean something in particular, and it is the reason this is metrology rather than a good estimate.
Fourteen ways to sense an arm, collapsed to five.
There are about fourteen credible ways to read what a forearm is doing. They are not fourteen separate choices — they are seven kinds of information, several of which more than one method can measure. Grouped that way, most of the fourteen either merge or fall out, and five blocks are left.
The selection rule mattered more than the list: kill a candidate only when another is a strict superset of its information. A naive keep-whatever-is-easiest rule would have killed ultrasound in favour of the cheap pressure ring — which would have been wrong, because pressure cannot see stiffness at all.
And they arrive in a deliberate order
The cheapest, most-proven configuration ships first and answers the most commercially important question. Every later stage adds exactly one genuinely new kind of information, and each one sits behind an experiment that can stop it.
- v0electrodes · accelerometers · IMU · FMG · artifact · cameraforce, contact events, kinematics, uncertainty→ K1
- v1A-mode ultrasound, dry silicone couplingfatigue-robust muscle deformation→ K2
- v2the micro-tapper → tendon tensiondirect-physics tension; enables L1→ K3
- v3ultrasound → imaging + elastographymeasured, not inferred, stiffness→ K4
- v4electrodes → high-density gridmotor-unit-level neural drive→ K5
Where each block sits on the arm, how the emitting ones are time-multiplexed so they do not step on each other, and what any of it costs per line are not published. Naming what kind of sensing exists is fine; handing over the layout is not.
Why a squeezable object is a metrology instrument.
In metrology this is called a transfer standard: a physical object of known, stable, traceable properties, used to tie an instrument's readings to real units. Every serious measuring instrument has one. No robot-data-capture product ships one. That absence is the opening.
It does a second thing that matters more. The laboratory method for measuring limb impedance is perturbation: displace the limb suddenly and measure the restoring force. The instrument that does this costs tens of thousands of dollars. An object that can change its own stiffness while being gripped runs the same experiment in the user's hand — which is what makes the stiffness channel possible at all outside a lab.
What the object is made of, how it changes stiffness, and the range it can cover are held back. The purpose, the metrology framing and the fact that it changes stiffness at all are said freely — those are the parts worth being generous about.
The thing that is actually being sold.
Everything upstream exists to produce this. Everything downstream exists to consume it. If this is right and the hardware is mediocre, the product still works. If this is wrong, no amount of sensing rescues it.
Kinematics
200 Hzhand + object pose, arm dynamics
commodity — everyone has this
Wrench
200 Hzforces + torques induced ON the object, in the object's frame
object-frame wrench validated as an embodiment-agnostic space by CHORD (NVIDIA, 2026)
Impedance
50 Hztask-space stiffness ellipsoid over time
tele-impedance and variable-impedance control, a mature literature
Contact events
±1 mstouch · load · lift · hold · slip · replace · release
Johansson & Flanagan — contact events are the subgoals of manipulation
Uncertainty
per samplea stated interval on every force and stiffness sample
conformal prediction — proven elsewhere, unclaimed here
Each field is independently validated in the literature. The combination exists nowhere. One well-known dataset has wrench but no impedance, no events and no uncertainty. Another has dense tactile patterns but no SI units and no stiffness channel. The forearm band that came closest has force and none of the rest.
demonstration:
meta:
subject_id, session_id, sleeve_serial, artifact_serial
enrollment_ref, verification_ref, config, clock
stream kinematics @ 200 Hz → SE(3) hand/object pose + arm state
stream wrench @ 200 Hz → R^3 force, R^3 torque, object-fixed frame
stream impedance @ 50 Hz → R^3x3 stiffness (SPD), principal ellipsoid
events contacts async → touch | load | lift | hold | slip | replace | release
stream quality @ 1 Hz → donning_index, channel_health, drift_estimateOne design choice worth remembering
Wrench is recorded in the object's frame, not the hand's. Hand-frame forces are morphology-specific — a five-fingered grasp and a two-jaw grasp produce completely different hand-frame numbers for the same outcome. Object-frame wrench does not care who or what applied it. That is what embodiment-agnostic means concretely, and it is what makes the Record directly consumable by cross-embodiment methods.
And one commercial one
The schema will be published openly and permissively licensed. The value is not in the format — it is in being the only party who can fill it correctly. An open format recruits the ecosystem that makes the data valuable; a closed one guarantees competing alone against better-funded labs.
Seven layers, scored on themselves.
Most of these use well-known techniques from other fields, and say so. Three are, as far as we have been able to find, unclaimed — and all three only work if you also own the hardware, which is what makes them hard to copy rather than merely hard to think of.
- L1
Physics-anchored fusion
flag plantedOur channels are not peers. One of them is grounded in direct physics; the others are statistical estimators. So the grounded one continuously corrects the rest, and calibration does not go stale during a shift.
The loop itself, and the window it exploits, are held back pre-term-sheet.
- L2
Conformal force intervals
flag plantedTurns a promise of quality into a number you can audit: an interval with a coverage level, from a method with distribution-free guarantees and no priors to argue about.
Validated against public data before any of our own hardware existed — see the evidence below.
- L3
Hybrid physics-ML estimator
Constrain the estimator with musculoskeletal forward dynamics so its outputs are physically realisable rather than merely plausible.
Good engineering, mature literature, not a breakthrough. This layer is not where the moat lives.
- L4
Missing-modality robustness
One software stack serves every hardware tier, because customers will own different configurations and nobody wants a model per SKU.
Commercially essential rather than novel. Without it the staged roadmap multiplies engineering cost by five.
- L5
The Record — the IR
flag plantedThe format itself is a layer. Sleeve streams are the source language, robot actions the target, and this sits in between.
To be published openly. The value is not the format; it is being the only party who can fill it correctly.
- L6
Uncertainty-weighted learning
Error bars only earn their keep if they improve training. Downweight the uncertain samples, or hand the interval width to the model as a feature.
Every existing method weights by demonstrator suboptimality. None weights by sensor-derived measurement uncertainty, because no dataset has ever carried it.
- L7
Privileged distillation
Train with the sleeve's rich channels as privileged information; deploy on a robot that has a camera and nothing else.
One sleeve-collected dataset upgrades an entire fleet of force-blind robots, permanently, with no hardware change on the robots.
What we actually know, including where we were wrong.
Before any hardware existed, the uncertainty claim was tested on a public dataset — twenty subjects, 256-channel forearm recordings, per-finger force ground truth, seventy-four thousand windows. Every prediction was written down and frozen before a model was fitted. Then the write-up was handed to an external reviewer, who was right on every substantive point, and six claims were corrected or withdrawn.
This is the only part of this sheet with real evidence behind it. It is also the part most companies would not publish.
The premise this company was previously built on was that a population-calibrated interval would be too narrow for a new user and would silently under-cover. It does not. Coverage held at nominal; the method absorbed the shift by widening intervals 1.77× instead. Conformal prediction behaved correctly under a shift it was not guaranteed to survive. Any deck still asserting otherwise — including ours, for a while — is wrong.
Something narrower and better evidenced. Per-user calibration is worthless on average — it moves mean coverage by −0.5 points — and it removes the tail entirely: under population calibration 20% of subjects fall below 85% coverage; under per-user calibration none do. One user in five gets materially wrong error bars without it.
We tried to kill our own product
If a smarter model could fix re-donning with zero burden on the user, the per-session ritual would not be justified and the calibration artifact would lose most of its rationale. So we pre-registered the falsification: if the field's standard fix — randomised channel masking during training — lifted cross-session coverage to 83.3% or better, the ritual was unjustified.
It reached 57.1%. It moved coverage the wrong way by 2.4 points, while buying an accuracy gain that does not approach usable. The artifact survived the hardest test available on public data — which is worth more precisely because the test was set up to go the other way.
Read honestly: channel masking is documented to help cross-day classification. It did not help cross-day regression here. That is a real, pre-registered negative result on the method the field would reach for first — on one dataset, with one model class.
And then we caught ourselves doing it
Coverage was being computed from hundreds of overlapping windows treated as independent. Windows are not exchangeable units. Trials are. Recomputed properly:
| calibration | window-level | certifiable at 90% |
|---|---|---|
| 4 trials · 100 s | 86.6% | 0 of 20 subjects |
| 9 trials · 225 s | 89.0% | 20 of 20 · intervals 2.15× wider |
At four calibration trials the honest radius is infinite for every single subject — a 90% guarantee mathematically does not exist at that budget — while the comfortable-looking window-level number reads 86.6%. That gap, a figure that looks acceptable against a guarantee that is not there, is the exact failure mode Touchstone exists to sell against. We shipped it in our own write-up for a day. The corrected spec is longer and more expensive, and it is the one we quote.
The label on every claim
Every factual claim this company makes, internally and publicly, carries a label for how solid it actually is. The rule that gives it teeth: no hypothesis appears in customer-facing material until there is an experiment behind it and the label has changed. A promise to sell calibrated numbers with stated error is worthless from a company that will not hold itself to the same standard.
- DDemonstrated
- independently demonstrated in peer-reviewed work
- CCompany-reported
- claimed by a company, not independently verified
- RResearch prototype
- shown once in a lab, not productised
- IInference
- our own reasoning from evidence, not itself measured
- HHypothesis
- speculative — must be gated by an experiment
Five ways to find out we are wrong, cheaply.
Each is a specific experiment with a written pass/fail line decided before it runs, and a stated plan for what we do if it fails. A failed gate is not a setback. It is a cheap way of learning not to spend money on something, before the money is spent.
- K1
the channel-value experiment
not runthe one that can end the premium tier
Do the impedance and contact-event channels measurably improve force-sensitive policy performance beyond wrench alone?
- requires
- v0 hardware only
- if it fails
- Ship v0 only. Abandon the premium tier. Do not reframe the claim.
Third-party evidence that force helps at all is strong and multiplying — around 22–23% success-rate improvements from force-grounded policies in 2026 work. All of that validates wrench beating kinematics-only. None of it touches K1's question, which is whether contact schedule, impedance and uncertainty weighting beat wrench alone. No third party has run that comparison.
- K2
fatigue robustness
not runDoes ultrasound fusion hold force accuracy across a full four-hour capture shift, where the electrode channel alone degrades?
- requires
- v1 hardware
- if it fails
- Drop the ultrasound block. Rely on anchoring and more frequent re-verification.
- K3
wrist tensiometry
not runhighest risk in the programme — and, now, highest value
Does direct-physics tendon tension sensing work at the wrist, with usable signal-to-noise, across people?
- requires
- v2 hardware — cheap relative to its information value
- if it fails
- L1 degrades to conventional fusion. The highest-novelty hardware claim is withdrawn.
The whole published literature for this measurement is lower-limb; the wrist is unclaimed territory. Since 2026-08-20 its payoff is measured rather than asserted: without it, re-donning the sleeve costs the user minutes of labelled calibration every session. K3 decides between two different products — one where putting the sleeve back on is invisible, one where it costs four minutes every time.
- K4
measured versus inferred stiffness
not runDoes direct elastography beat the cheaper inferred estimate, judged against the artifact's own perturbation ground truth?
- requires
- v3 hardware
- if it fails
- Keep the cheaper inferred impedance. Do not build v3.
- K5
high-density grid value
not runDo motor-unit features improve per-finger resolution enough to justify a tenfold increase in electrode channel count?
- requires
- v4 hardware
- if it fails
- Stay at low-density electrodes permanently.
The numeric threshold behind each of these is written down and decided in advance. It is not published here, because in the wrong room it is a shortcut to somebody else's validation design.
An empty band, and two curves moving the right way.
Below about $130 you can buy relative, uncalibrated signal. Above about $4,900 you can buy professional mocap and haptic hardware that was never designed for force capture. Between the two, essentially nothing carries a public list price. Every number below is checkable by anyone.
The size of it, bottom-up
Roughly $277.5M a year is already being spent on manipulation data — twenty frontier labs at about $6M each, seventy mid-stage robotics companies at about $1.8M, and several hundred academic and industrial teams at about $70k. Cross-checked a second way: at the March 2026 benchmark that implies about 2.35M data-hours a year industry-wide, or around twenty large capture operations across the whole industry — which is roughly what exists.
| 2029 | 2031 | |
|---|---|---|
| Data | $132M | $419M |
| Hardware | $22M | $75M |
| Recurring calibration | $11M | $45M |
| Total | $165M | $540M |
The reframe that matters
The companies raising serious money to build wearable capture networks are building the distribution infrastructure for exactly this — and none of them captures force. We are not competing with them for the same buyer. We are the missing physical channel inside collection networks that are already being deployed by the thousand.
And the honest weak point
The open capture rigs are $371 and open-source. The closest forearm band is likely under $150 of parts. Our premium cannot be defended by having more sensors — only by channels that do not exist at any price: calibrated newtons, stiffness, and a stated interval. If gate K1 does not show those improve policy performance, the premium is indefensible, and the honest response is to ship v0 alone rather than to reframe the claim.
There is a clock on this, too. The labs that proved the physics still exist, all now see the same newly-paying market, and several have working prototypes we do not. “Nobody built the product” can mean there were structural reasons — and there were — or it can mean somebody has not got round to it yet. No framing changes which one it turns out to be.
Said plainly.
- System specificationcomplete — five blocks, seven layers, five gates, costed at three volumesR
- Uncertainty validationseven arms run on public data, externally reviewed, six of our own claims correctedR
- v0 hardwarenothing built. Parts not ordered. This is the binding constraint on everything below—
- Gate K1untouched — there is no robot, no policy and no task anywhere in the programme yetH
- Gate K3untouchedH
- Customer interviews3 of a target 15 — the cheapest high-value work available, and the most consistently avoidedD
- The Record schema, publishednot started—
Calibration does not remove error. It converts an unknown error into a stated one. That is the whole promise, and it is the one no competitor currently makes.
Back toThe atlas — the whole system on one sheet →physcriptthe force layer for physical AI