Portals: Persistent, Editable 4D Spatial World Models on Edge Devices

James A. Tunick1,2, Ryan Brant1, Jacob D. Pennock1, Justin Kasowski1
1H3M, Inc. 2The IMC Lab, New York
CVPR 2026 Workshop on 4D World Models: Bridging Generation and Reconstruction

Persistent, editable 4D spatial world models for mobile devices, built with edge-device rendering constraints as a first-class requirement.

Selected footage captured on device (iPhone 14 Pro, iPad Pro, Apple Vision Pro). Download MP4.

Abstract

We present Portals, a deployed systems architecture that bridges 4D world-model research and persistent spatial experiences on phones, smart glasses, and augmented-reality headsets. The defining shift in spatial computing is not from 3D to 4D, but from stateless scenes to stateful worlds—scenes that persist, compound, and stay editable across sessions, devices, and users. Delivering such worlds on constrained hardware is the central systems problem: they must render in real time, survive revisits, and remain authorable through voice, gesture, and no-code tools. Built on 3D Gaussian Splatting and informed by 4D-GS and Generalizable Human Gaussians, Portals has been deployed on iOS, with web-based viewers (including Apple Vision Pro) for reconstructed environments, volumetric humans, and holographic spatial media. We contribute: (1) an edge-device runtime built around LOD-adaptive Gaussian splatting (SPAG) and a shared spatial-media compute substrate that fuses depth, stencil, audio, and ML-pose channels, driving 360+ source-agnostic VFX effects at 60 fps on iPhone 14 Pro (2.7–4.1× speedup); (2) a persistent geospatial scene-state architecture with layered world metadata, reloadable scene payloads, and anchor-guided re-alignment across sessions; (3) a creator-facing composition pipeline that bridges reconstruction and generation through VFX composition, voice-driven semantic actions (with on-device intent parsing and cloud fallback for ambiguous utterances), and no-code authoring; and (4) benchmark axes for evaluating 4D world models under deployment constraints such as mobile rendering efficiency, persistent scenes, and editable world state. Prior clinical deployment of volumetric AR at Memorial Sloan Kettering established the real-time rendering primitives underlying this work.

BibTeX

@inproceedings{portals2026,
  title={Portals: Persistent, Editable 4D Spatial World Models on Edge Devices},
  author={Tunick, James A. and Brant, Ryan and Pennock, Jacob D. and Kasowski, Justin},
  booktitle={CVPR 2026 Workshop on 4D World Models: Bridging Generation and Reconstruction},
  year={2026}
}

Visual Highlights

Real-time volumetric performance capture with generative fire and sparkle VFX on live performers
Real-time volumetric performance capture — live performers driving generative fire, sparkle, and light VFX on device.
Portals deployments across mobile and immersive hardware
Representative deployment surfaces spanning mobile AR, spatial capture, and immersive display.
Portals AI-assisted spatial authoring workflow
AI-assisted spatial composition and scene control inside the shipped creator workflow.
Browser-based splat and web delivery view from Portals
Browser-facing WebGL and WebGPU delivery path for persistent spatial media on the web.

Device Captures

All imagery captured on physical devices — iPhone 14 Pro, iPad Pro, and Apple Vision Pro.

City of holograms — a generative holographic city anchored in real space, captured live on iPhone
City of holograms — a generative holographic city anchored in real space, captured live on device.
Voice composer with AR VFX and voice command overlay
Voice-driven AR composer with real-time VFX. Voice command "add a pink sphere" parsed locally in <1ms.
Apple Vision Pro running Portals at IMC with NYC skyline
Apple Vision Pro rendering interactive 3D worlds with passthrough at The IMC Lab, NYC.
Portals mobile composer UI
Full AR authoring UI on iPhone 15 Pro with voice input, VFX palette, and scene manipulation.
On-device compute-shader library: dozens of galaxy, particle, fractal, cellular-automata and fluid simulations
Compute shaders on-device — galaxy simulations, particle systems, cellular automata, and anatomical models, shown across the full effect library.
iPad running real-time voxelized body tracking on XR LED stage
iPad at 60fps with voxelized body tracking on an LED wall XR production stage.
Hologram mode with procedural environment VFX
Hologram mode with procedural environment VFX and real-time voice parametric control.
Neural style transfer avatar
Real-time neural style transfer applied to volumetric avatar capture.
Depth-to-VFX pipeline in gallery installation
Sparse depth maps driving interactive projections in a gallery installation.
Metavido multi-channel encoding: depth, human stencil, and metadata alongside RGB
Shared spatial-media substrate — a single capture fused into ML depth, human-segmentation stencil, and pose/metadata channels driving downstream VFX.
On-device face-mesh scanning reading ARKit blendshape coefficients
Real-time face-mesh scanning — the on-device ML expression channel reading live ARKit blendshape coefficients.
ML facial-expression channel detecting a smile at 0.76 confidence
Contextual affect awareness — the same ML channel flags a detected smile (confidence 0.76), driving expression-reactive effects.
Holographic city with geolocated routes and scene anchors
Persistent geospatial world-state — GPS/ARKit-anchored scenes with routes and pins re-aligned across sessions.

Paper Figures

Portals system architecture figure from the paper

System architecture showing the persistent world stack, shared spatial-media substrate, and deployment surfaces across mobile, headset, and web clients.

Portals content pipeline from the paper

Content pipeline across capture, reconstruction, editing, and deployment. In the shipped system, asset ingestion centers on phone capture, photogrammetry, and text or reference-image generation, while voice, direct manipulation, and no-code controls act as authoring surfaces over the same persistent scene graph.

Portals model and representation diagrams from the paper

Model and representation diagrams summarizing how environments, volumetric humans, and spatial media fit into the same editable 4D world framework.

Portals level-of-detail and performance figure from the paper

Level-of-detail and performance figure highlighting mobile-first runtime constraints and the efficiency tradeoffs required for edge deployment.

Systems comparison figure from the paper

Comparison of Portals against prior systems along persistence, editability, and deployment-relevant world-model axes.

Supplemental Materials

Extended capability context on the deployed Portals / XRAI platform — the data-driven, ML/AI, and HCI dimensions the system embodies, beyond the camera-ready paper.

Perception & Contextual Awareness

what the system sees, measures, and anticipates — on device

🌀 Each capability stands on its own. Combined, they form the intelligent substratecontextual awareness greater than the sum of its parts that learns, self-corrects, and improves with every use.perceive → understand → act → learn ↻
🗺️ Scene & spatial understanding
CapabilityMost-valuable use caseMobile / HeadsetWeb
📏Metric depth sensingTrue-scale placement & measurement🟢🟡
✂️People occlusion & mattingPeople walk in front of portals; clean cutouts🟢🟡
🏷️Scene mesh + semantic labelingContent reacts to floors, walls, objects🟢🟡
🟦Surface / plane detectionSnap & anchor content to real surfaces🟢🟡
🏠Room-layout understandingInstant scene scaffolds from a room scan🟡🟡
🪟Glass / window de-occlusionPortals stay visible through glass🟡🟡
📍World-locked geo-anchoringCity-scale, GPS-locked persistent portals🟢🟡
🔎Object & image recognitionMarker & product-triggered experiences🟢🟢
🔤Text, code & saliency readingRead signage; attention-aware overlays🟢🟡
🧠 Contextual awareness & prediction 🔒 capability level
CapabilityMost-valuable use caseMobile / HeadsetWeb
🧭User-intent inferenceThe right tool surfaces before you ask🟢🟢
🔮Needs prediction & anticipationAnticipates your next step🟡🟢
📊Behavior modelingPersonalizes to how you move & create🟡🟡
😊Affect & expression sensingEmpathetic response to mood & expression🟢🟢
📈Trajectory predictionPredicts where the ball or person is heading🟡🟡
🕸️Relationship mappingMaps how people, objects & spaces relate to each other🟡🟢
🤖 Multimodal intelligence — Jarvis loop 🔒 capability level
CapabilityMost-valuable use caseMobile / HeadsetWeb
🗣️Voice + vision understandingAsk about what you're looking at, hands-free🟢🟢
👉Open-vocabulary “point & ask”Point at anything, get an answer or action🟡🟢
☝️Raycast pointing — near & far selectionPoint a finger to select & act on objects at any distance🟢🟡
🔁Perceive → understand → act loopSees, understands, and acts in one loop🟡🟢
🏃 Motion, body & trajectory
CapabilityMost-valuable use caseMobile / HeadsetWeb
🎯Trajectory & projectile detectionBall-flight & shot tracing from a phone — no radar🟡🟡
Real-world speed from videoSwing & body speed from a single camera🟡🟡
🦴Full-body 3D pose & skeletonForm & technique coaching; live avatar drive🟡🟡
Per-body-part keypointsFine motion (wrist/club) metrics; gesture control🟡🟢
🚶Gait & motion-form analysisSession-over-session tempo & plane deltas🟡🟡
👥Multi-object trackingPeople & props tracked for interaction🟡🟢

🟢 demonstrated 🟡 in progress 🔒 capability level

Status reflects current maturity per surface. Contextual and agentic capabilities are shown at capability level.

XRAI Platform Capabilities

the interchange format that makes a world portable

🧬 XRAI is the DNA of a Portal — one portable document carries a scene's structure, intent, provenance & edit history and renders on any engine. Store the generation process, not just the output.Portals · XRAI · jARvis
{
  "xrai_version": "1.0",
  "id": "…0001",
  "author": {"type": "human", "id": "example"},
  "origin": {"app": "example", "version": "1.0", "scene": "minimal"},
  "scene": {
    "anchors": [],
    "entities": [
      {"id": "cube_1", "type": "object.primitive",
       "transform": {"position": [0, 0.25, -1.5], "rotation": [0,0,0,1], "scale": [0.2,0.2,0.2]},
       "material": {"color": "cyan", "preset": "neon"}}
    ],
    "relations": [], "events": []
  }
}
📦 File format & interchange
CapabilityWhat it's forMobile / HeadsetWeb
📄Portable scene manifest, geometry out-of-lineTiny revisit traffic — structure travels, payloads stream on demand🟢🟢
💾Persistent + editable world stateSave, re-open, mutate & re-version worlds across sessions, devices & users🟢🟢
🧾Edit-history + per-object provenanceReplay creator intent; track source, license, generator, timestamps🟢🟡
🧩Multi-asset typed containerOne manifest routes primitive / splat / hologram / portal nodes to the right renderer🟢🟢
🔀Cross-runtime portabilityOne .xrai decodes on many renderers via an open schema🟢🟢
Temporal validity windowsContent that appears only during an event, or on a schedule🟢🟡
🧬Generative encoding — "store the rules, not the mesh" 🔒A compact seed / rules expand into a full world🟡🟡
🧊 3D / 4D / nD formats
FormatWhat it's forMobile / HeadsetWeb
Gaussian splats (.splat / .ply / SPZ)Photoreal captured environments; web-light delivery🟢🟡
📦glTF 2.0 / GLB (runtime baseline)Universal asset payload every runtime reads🟢🟢
🍎USD / USDZ interopCompose with the Apple / DCC ecosystem🟢🟡
🎞️Volumetric video (depth + color) 🔒Records a hologram into a standard video texture — streams over ordinary codecs🟢🟡
🧍Mesh + avatar interchange (FBX ↔ GLB, VRM)Photogrammetry meshes, humanoid avatars with expressions🟢🟢
☁️Point-cloud formats (PLY / E57 / LAS)Scan data into the same scene graph🟡🟡
🔗KHR_gaussian_splatting interopRide the emerging cross-platform splat standard🟡🟡
🕸️ Semantic & ontological
CapabilityWhat it's forMobile / HeadsetWeb
🗂️Spatial Scene Graph (SSG)Hierarchy of (geometry, geo-anchor, time-window, LOD) — the core representation🟢🟡
🔵Entity / relation / event / anchor modelId-stable entities and relations that survive edits🟢🟢
Typed relations with runtime meaningreacts-to-audio, tracks, parent-of — behavior, not just hierarchy🟢🟢
🗣️Natural-language → scene-graph operationsVoice edits and direct manipulation hit the same structured ops🟢🟢
🏷️AR scene understanding & semantic labelsFloors, walls, tables, faces, bodies typed with confidence🟡🟡
🧭Faceted ontology (multi-axis)Identity, time, scale, modality, provenance… as first-class axes🟡🟡
🧠 Knowledgebase & knowledge graph
CapabilityWhat it's forMobile / HeadsetWeb
🔌Encode-anything → graphTurn a webpage, repo, paper, or calendar into an XRAI graph🟢🟢
🕸️Knowledge graph → spatial worldNodes become objects/rooms, edges become portals, clusters become districts🟡🟡
🌐3D force-graph visualizationExplore relationships as a live, navigable 3D graph🟡🟢
🔭Semantic-zoom navigationOne continuous zoom from universe → district → artifact🟡🟡
🧷Persistent cross-session memoryWorlds and knowledge compound instead of resetting🟢🟡
✨ Generative & procedural
CapabilityWhat it's forMobile / HeadsetWeb
🖼️Text / image → 3D assetPrompt an object or room-scale asset in seconds🟢🟡
📷AI-assisted reconstructionPublishable splats from a ~20-second phone video🟢🟡
🌱Procedural world generationVoice → procedural environments from seed + rules🟢🟡
🎙️Voice / no-code authoring loopOn-device intent parsing (sub-ms) with cloud fallback; every path emits the same ops🟢🟢
🎨Whole-scene neural style transfer"Make everything look like ___" across the live scene🟡🟡
🧑‍🎤Avatar generationPerson → avatar / volumetric double🟡🟡
🌌 VFX & holograms
CapabilityWhat it's forMobile / HeadsetWeb
🎆360+ source-agnostic real-time VFXConsole-grade effects at 60 fps on a phone🟢🟡
🧱Shared compute substrate 🔒Depth → world / stencil / velocity computed once; each effect adds only render cost🟢🟡
🎚️Multi-channel effect bindingDepth + stencil + audio + pose auto-resolved across effect conventions🟢🟡
🫧Body-driven / stencil-masked VFXSparks, embers, trails that hug a live person🟢🟡
🎥Live + recorded holographic mediaRecord → edit → replay; effects work on live and decoded alike🟢🟡
🛰️Live hologram room (browser ↔ iOS)Share a live hologram across app and browser🟡🟡
⚙️ Render pipelines
CapabilityWhat it's forMobile / HeadsetWeb
⚛️LOD-adaptive Gaussian splatting (SPAG)Panoramic image → editable 3D Gaussians; four-tier LOD (2.7–4.1× speedup)🟢🟡
🎛️Adaptive quality controllerLive FPS + thermal tuning holds the 16.7 ms frame budget🟢🟡
📱Mobile-first / edge rendering60 fps iPhone 14 Pro · 57 fps Galaxy S23 — no cloud required🟢
🔺Mesh / PBR pipelineStandard glTF / URP path for classic assets🟢🟢
🌐WebGL / WebGPU web renderingGraph + scene viewers in the browser🟡🟢
🧩Multi-runtime adaptersOne document, many engines (Three.js + Unity locked; others in progress)🟢🟡
🔬 Splat techniques & pipelines how they rank · ★ our pick
Technique / formatBest atMobileEditableLODWeb-lightOur call
3D Gaussian Splatting (3DGS)Photoreal captured scenes🟡🟡🟡🟡baseline
SPAG — Spherical Pixel-Aligned GaussiansPanorama → editable splats · LOD · O(1) edit🟢🟢🟢🟡 mobile
4D / animated splats (DynGsplat)Moving / temporal splats🟡🟡🟡🟡in test
Human Gaussians (GHG)Feed-forward volumetric avatars🟡🟡🟡🟡roadmap
Latent 3DGS (generative)Fully-synthetic props from a prompt🟡🟡🟡research
Photogrammetry meshStatic hero objects🟢🟢🟢🟢classic
.splat / .plyUniversal splat containers🟢🟡 interchange
SPZ~10× smaller splats for the web🟡🟢 web
KHR_gaussian_splattingEmerging cross-platform standard🟡🟡adopt on ratify
Render pipeline — one .xrai → many enginesSurfaceStatusRank
Unity gsplatiOS / mobile🟢 mobile
Three.jsWeb (reference)🟢 web
PlayCanvas / SuperSplatWeb🟡2
Needle (WebXR)Web XR🟡3
IcosaWeb🟡in test
RealityKitvisionOS (native)🟡planned
UnrealPC / console🟡planned
📊 Data-viz & X-ray observability
CapabilityWhat it's forMobile / HeadsetWeb
🩻Glanceable system X-rayEvery stage as status · timing · why — see a pipeline before you read it🟡🟡
🧾Provenance / lineage per actionCost, model, and source trace behind every generated thing🟡🟡
📈3D data visualizationLarge graphs and datasets rendered in WebGPU🟡🟢
🔍See-through / see-across transparencyInspect how a world was built and how its parts relate🟡🟡
📱 Devices & platforms
PlatformWhat runs thereNativeWeb
📱iPhone (iOS)Flagship: 360+ VFX @ 60 fps, full runtime, voice + gesture authoring🟢🟡
🤖Android (Galaxy S23)Render-verified at 57 fps🟢🟡
🥽Apple Vision ProWeb viewer for environments, volumetric humans & holographic media🟡🟢
🌐Web / browserEncode / decode, graph viewer + editor, voice → scene🟢
👓AR headsets / smart glassesTarget surface; glasses-HUD experiences designed🟡🟡
🧠Edge devices (no-cloud)Stateful worlds that live on the device — the runtime, not a thin client🟢🟡
🎬LED / XR production stageProven in live-event & virtual-production deployments🟢
📎 Grounded examples real shipped fixtures / deployments unless noted

🟢 shipped / demonstrated 🟡 in progress 🔒 capability level our pick

One portable format, every renderer. Statuses are conservative — Web means the encode/decode + viewer surface, not the full mobile VFX stack. 🔒 items are shown at capability level only.

QR code linking to this project page

Share this paper

Scan the code to open on mobile, or share the link.

Status

The paper linked above is the accepted camera-ready submitted to the CVPR 2026 Workshop on 4D World Models (Bridging Generation and Reconstruction), certified by IEEE PDF eXpress on April 10, 2026. The interactive web demo, code release, and expanded project materials will follow after the workshop.