signal / live study

DEPTH IS A GUESS

A flat photograph becomes a spatial claim.

Choose an image, then explicitly ask a small local model to infer relative depth. Pull that reading apart and change viewpoint until the photograph stops behaving like a complete description of the scene. The model does not recover hidden truth; it proposes one structure from visible evidence.

image surface / inferred field

surface → local inference → disturbance → recognition

idle

choose an image below · mobile browsers may offer camera capture

choose an image to create a surface

After inference, move the pointer across the image or tap a viewpoint. Keyboard: arrow keys change viewpoint, Home returns to center.

sourcenone
relative field—
runtimenot loaded
modelDepth Anything V2 Small

Your image is resized and processed in this browser. Warp the Surface does not upload it for inference. Enabling depth downloads open model weights from Hugging Face: about 19 MB on the preferred WebGPU path, with a larger quantized WASM fallback on devices without usable WebGPU.

under the surface

The hidden layer is not depth. It is an assumption about depth.

surface

A photograph arrives as a coherent rectangle and is easy to read as a complete record.

disturbance

A monocular model converts visible cues into one relative depth field, then that field moves the same image evidence apart.

recognition

Changing viewpoint makes the inferred structure useful and visibly provisional at the same time.

carry

Machine perception can reveal structure without becoming ground truth. Prediction is evidence with assumptions attached.

system

The model supplies a field. Warp supplies the experience.

source
local image or explicit mobile camera capture
inference
Depth Anything V2 Small / quantized ONNX
runtime
Transformers.js 4.2 / dedicated browser worker
preferred device
WebGPU / q4f16 / approximately 19.1 MB model weight
fallback
WASM / q8 / approximately 27.3 MB model weight
renderer
Canvas 2D sampled spatial field + one depth-derived surface section
server inference
none
model license
Apache-2.0

This study intentionally avoids a full 3D reconstruction. The point is not to pretend one photograph contains a recoverable scene graph. The smallest useful model estimates a relative field; the browser then turns that prediction into an explorable perceptual contradiction.