Video calls are excellent at transporting rectangles of light and sound. They are poor at creating presence.
The problem is not resolution. It is that a call remains an application competing with every other application on a device designed for divided attention. The camera is a feature. The person becomes a window among notifications, tabs, and controls.
I wanted a dedicated object.
An always-available window between two places would not ask me to launch a meeting, arrange my face inside a frame, and perform continuous attention. It would have a stable physical location. I could glance toward it, approach it, leave it quiet, or use it for a deliberate shared act.
Presence is partly a device because attention has geometry.
Connected Reality
Virtual reality places me inside a constructed world. Physical telepresence asks for the opposite movement: bring a remote or imagined object into the room where I already live.
The object might begin as a mobile display or projection surface. A more ambitious version could change shape, create a touchable interface, move an instrument into reach, or let a remote person manipulate a corresponding object at my location.
Imagine asking for a piano keyboard. A local machine unfolds or projects a playable surface in front of me. Keys have physical response. Sound is produced locally. A remote teacher can indicate hand position or play a phrase through a paired instrument.
That is not a video of a piano. It is a temporary device whose state can be shared.
The ultimate claytronics vision is matter that can reassemble into arbitrary forms. I begin with one useful physical degree of freedom at a time and let each device earn the next.
Latency Changes the Meaning of Touch
Latency changes what remote touch can safely mean.
Even across Earth, a round trip takes time. Through congested networks it takes more. For distant robots, continuous joystick control becomes unstable or exhausting. The remote system needs local autonomy: hold balance, avoid collisions, respect force limits, and complete a small action without waiting for every correction.
The machine does not need to predict my fingers before I decide to move them.
An earlier version of this idea imagined continuously modeling the user so aggressively that the remote device could anticipate intent. That would turn presence into surveillance. I prefer explicit rules:
- infer only what is needed for the current action;
- perform inference locally when possible;
- retain no behavioral model without informed consent;
- show when autonomy is active;
- provide an immediate stop;
- never convert a prediction into irreversible physical force without a safety check.
Prediction can hide latency. It must not erase agency.
The Semantic Safety Layer
When communication changes physical state, malformed or hostile messages can hurt people.
A physical-telepresence protocol therefore needs more than encryption and packet delivery. It needs an action model.
The receiving device should know the bounded operations it is willing to perform: move this joint within this envelope, apply no more than this force, heat no surface above this temperature, remain outside this protected volume, stop when sensing disagrees with expectation.
High-level intent can be checked before it becomes actuation. Low-level controllers must still enforce limits even if the high-level model is wrong.
No “strong AI” can guarantee universal safety by understanding every possible human action. Practical safety comes from constrained mechanics, independent limits, authentication, local sensing, fail-safe states, and a protocol whose verbs are smaller than “do whatever the remote person wants.”
The physical world deserves typed interfaces.
A Shared Device Economy
Sophisticated telepresence hardware will not begin cheap. That suggests shared access.
A neighborhood or building could host mobile devices that users reserve for a particular task. The device arrives, authenticates both locations, performs a bounded session, and returns to charge or maintenance. Specialized attachments could be shared rather than purchased by every household.
The model resembles a library of physical capabilities more than a disposable gadget market.
Shared ownership also raises hard questions:
- Who cleans and repairs the device?
- Which sensors operate between sessions?
- Where are recordings stored?
- How does a user inspect the machine before inviting it inside?
- What prevents harassment or physical misuse?
- Can a person use the essential function without surrendering unrelated personal data?
The answers determine whether the shared system feels like public infrastructure or a stranger’s surveillance robot.
Build for Long Life
A physical platform should not become waste because its software fashion changed.
The durable module should separate long-lived capabilities from replaceable ones:
- power and structural frame;
- motors and force-limited joints;
- display, light, sound, and basic sensing;
- local safety controllers;
- replaceable compute and communication modules;
- software-defined behaviors installed through explicit permissions.
No module will contain every future sensor and emitter for decades. That was an attractive but unrealistic dream. Modularity gives a better route to longevity. A useful chassis can accept new compute, radios, tools, and surfaces without discarding the whole machine.
Reconfigurable computation matters here because the device may change roles. It does not make electronic engineering obsolete. Motors, optics, thermal paths, batteries, connectors, and electromagnetic compatibility remain stubbornly physical.
Software can coordinate them only after engineering makes each boundary trustworthy.
Dedicated Attention Without Demanding Attention
An always-on presence device must avoid becoming an always-on obligation.
The window needs social states:
- available for ambient presence;
- available for a deliberate call;
- quiet but open to a gentle signal;
- closed;
- physically covered or powered down.
Both sides should see the state. A dedicated device should make consent clearer than an app, not more ambiguous.
The goal is not continuous eye contact. It is continuity of place.
A grandparent might keep the window near a kitchen table. Two collaborators might leave paired work surfaces open while building separate parts. A musician might share a physical controller rather than compressed audio alone. A remote worker might manipulate one instrument for five minutes, then let the room become a room again.
Back to the Real World
Screens remove gesture, posture, shared objects, and peripheral awareness, then attempt to add them back with more pixels.
Physical presence begins elsewhere. It asks what action two places should share and builds a device for that action.
Sometimes the answer really is a camera and microphone. Sometimes it is a movable pointer, a paired sketch surface, a force-limited gripper, a shared keyboard, or a quiet window that occupies one dependable corner.
I still love the wild destination: remote reality entering the room as something I can touch. The path to it is not one universal morphing robot and not a model that watches me continuously.
It is a sequence of purpose-built devices. Each one gives a remote relationship one new physical verb, keeps that verb safe, and earns a place in the room.