Version: v0.4 Draft
Status: speculative design document, not medical advice, not a hypnosis protocol, not a claim of clinical validation
Author: Timur Yusupov, PhD
Working concept: Deep as a reversible gaze-controlled symbolic cognition interface
Attribution / inspiration: This work is my original reflection on the fictional technology described in Лабиринт отражений (Labyrinth of Reflections) by Сергей Лукьяненко (Sergey Lukyanenko). It is not affiliated with or endorsed by the author or rights holders.
Language note: Phrases such as “brain as renderer” and “brain-side reconstruction” are design metaphors in this draft. They describe an interface strategy, not a validated neuroscience claim, unless and until tested.
We are not trying to build a better screen.
We are trying to build a new state of interaction between human imagination and machine structure.
- 1. The Short Version
- 2. What Deep Is
- 3. What Deep Is Not
- 4. The Core Insight
- 5. The Tree Example
- 6. Perceptual Keyframes, Not Hidden Frames
- 7. The Visual Archive
- 8. The Deep Rendering Principle
- 9. Hardware Assumption
- 10. Deep-Lite: Cardboard Mode
- 11. One Camera Is Enough to Open the Door
- 12. Two Cameras Are for Measuring the Room
- 13. Head Tilt as the Movement Bus
- 14. Eyes for Meaning, Head for Motion
- 15. Control Modes
- 16. Assistive Technology Is the Treasure Cave
- 17. Deep-Lite Control Grammar
- 18. Input and Output
- 19. The Eye Interface
- 20. Deep Is Not Photorealism
- 21. Movement in Deep
- 22. The Body State
- 23. The State Machine
- 24. State Definitions
- 25. Return Is Part of the Product
- 26. The Return Script
- 27. Safety Constraints
- 28. The Epistemology Layer
- 29. The Deep Archive Format
- 30. What AI Does in Deep
- 31. First Killer Application: Training
- 32. Second Killer Application: Memory Palaces
- 33. Third Killer Application: Complex Problem Navigation
- 34. Fourth Killer Application: Creative Work
- 35. Why “Soft” Is the Real Product
- 36. The Minimal Prototype: Deep-0
- 37. What Can Be Built Now
- 38. The Research Questions
- 39. The Design Principles
- 40. The Product Stack
- 41. The Social Risk
- 42. The Deep Manifesto
- 43. Working Definition
- 44. Final Line
- References and Grounding Notes
Deep is a proposed human-computer interface mode where:
- the computer does not render a complete virtual world;
- the headset presents low-resolution visual archive markers;
- the brain performs the high-resolution reconstruction;
- brief perceptual keyframes bind object identity;
- simple icons then stand in for rich objects;
- audio remains the normal command and guidance channel;
- eye movement, gaze, focus, blinks, and fixation become the return channel;
- the system is built around one sacred requirement: the user must always be able to return.
In a normal VR system, the machine tries to create reality.
In Deep, the machine gives the brain structured seeds, and the brain grows the world.
Deep is a proposed interface for directed imagination.
It is not simply VR. It is not just hypnosis. It is not a cinematic dream simulator. It is not a chatbot wearing goggles.
It is closer to a new operating mode:
visual archive markers → brain-side reconstruction → gaze-controlled navigation → audio-guided cognition
The computer sends symbolic fragments.
The user’s brain unpacks them.
The eyes steer the system.
The voice layer keeps continuity.
The return protocol protects the user.
Deep is therefore best described as:
A reversible symbolic VR protocol that uses the human brain as the high-resolution renderer, eye movement as the control channel, and audio as the guidance layer.
Or shorter:
An operating system for directed imagination.
Deep is not a proposal to force sleep paralysis.
Deep is not a product that traps a user in an altered state.
Deep is not subliminal manipulation.
Deep is not hidden-frame advertising in goggles.
Deep is not a medical device concept unless and until it is evaluated under the appropriate medical, ethical, and regulatory frameworks.
Deep is not a replacement for consent, context, grounding, or ordinary human judgement.
The safest starting point is not “put people into sleep paralysis.” The safest starting point is:
Voluntary body quieting + high visual absorption + gaze control + reliable return.
Sleep paralysis is a real phenomenon, but it can involve inability to move or speak, chest pressure, hallucinations, fear, and distress. It should not be treated as a casual interaction target. The product target should be absorption, not paralysis.
Most interfaces assume the screen must carry the full representation.
Deep assumes the opposite:
The screen only needs to carry the right trigger at the right moment.
Human visual cognition can construct scenes from hints, fragments, symbols, movement, expectation, and memory.
When this draft says the brain is a renderer, it means that as a design metaphor: the system tries to provide useful cues and let ordinary perception, memory, and imagination fill in detail. It is not a claim that Deep has validated a specific neural mechanism.
A computer does not need to continuously show every leaf on every tree.
It can show:
- a distant marker;
- a growing shape;
- a brief detailed recognition frame;
- a stable symbolic icon.
After the object is bound, the brain carries the rest.
This is the central compression trick.
Imagine a user walking through a green field.
At first, there is only grass.
Far ahead, the user notices a dot.
The dot grows as the user approaches.
At some point the brain begins expecting recognition:
“That is probably something vertical. Maybe a tree.”
At that exact recognition threshold, the system shows a brief detailed image: trunk, branches, shadow, canopy, maybe bark.
Then the detailed image collapses into a simple tree icon.
From that moment onward, the icon is enough.
The user no longer sees “generic tree symbol.”
The user sees that tree.
The brain has been conditioned through a visible, intentional, logged perceptual keyframe.
The runtime sequence is:
dot → growing dot → expectation → perceptual keyframe → tree icon → brain-side continuity
This means Deep does not need constant high-detail display. It needs well-timed identity binding.
6. Perceptual Keyframes, Not Hidden Frames
The phrase “25th frame” may be useful as historical shorthand, but this document retires it after this sentence. It carries baggage from subliminal advertising, hidden persuasion, and cinematic mythology.
Deep should avoid hidden influence.
The better term is:
A perceptual keyframe is a brief, visible, intentional, logged high-information image used at the moment of expected recognition to bind an object’s identity.
It is not secret.
It is not subliminal.
It is not manipulation by stealth.
It is a cognitive anchor.
Object marker + recognition timing + perceptual keyframe = memory-bound object token
After that, a simple icon can represent the object because the user’s brain has already cached its identity.
A visual archive is not a video file.
It is a structured object store for cognitive rendering.
Each object can have:
object:
id: tree_001
category: tree
far_marker: green_dot
approach_marker: vertical_soft_shape
perceptual_keyframe: tree_oak_keyframe_01
token: tree_icon_simple
detail_levels:
- silhouette
- trunk_and_canopy
- bark_and_branch_detail
audio_label: "old oak near the path"
relations:
- beside: path_003
- north_of: field_001
epistemic_status: observed
confidence: 0.92The object is not “drawn” continuously. It is activated.
The archive stores enough to let the brain rebuild the scene without forcing the machine to render all details all the time.
Normal rendering:
machine renders detail → user observes detail
Deep rendering:
machine triggers identity → brain renders detail → machine tracks attention
This is a design principle to be tested, not an established empirical result.
Here, “brain renders detail” means the interface deliberately relies on visible cues, memory, expectation, and imagination instead of forcing the display to carry every detail.
This reverses the usual burden.
The display becomes an index.
The brain becomes the renderer, in the metaphorical sense defined above.
The eyes become the API.
The ideal hardware is not exotic.
It is a VR or mixed-reality headset with:
- high refresh rate;
- low latency;
- inward-facing eye cameras;
- gaze tracking;
- fixation tracking;
- blink detection;
- pupil tracking where safe and meaningful;
- head movement tracking;
- open-ear or high-quality audio;
- optional pass-through cameras;
- physical emergency stop;
- optional breathing or pulse sensor;
- clear session logging.
The headset is not used to create a photorealistic world. It is used to control visual input and track the user’s gaze.
The most important hardware feature is not resolution. It is closed-loop attention tracking.
The first real deployment wedge may not be an expensive headset.
It may be:
smartphone + cardboard/plastic frame + front camera + headphones
This is important because a modern smartphone already contains most of the required prototype hardware:
screen
front camera
speaker/headphones
IMU / gyro / accelerometer
battery
compute
storage
network
touch fallback
Deep does not require photorealistic VR for the first prototype.
Deep-0C only needs to prove the core loop:
visual marker → perceptual keyframe → icon binding → gaze command → safe return
That can be tested with commodity hardware.
Deep-0C means:
Deep-0 = first safe prototype
C = cardboard / commodity camera
The purpose of Deep-0C is not to build perfect eye tracking or perfect immersion.
The purpose is to test whether a person can:
- bind a detailed perceptual keyframe to a simple icon;
- navigate the resulting symbolic world by gaze;
- use head tilt for movement;
- return safely and consistently;
- recall the scene after exit.
The first prototype can be cheap, hackable, and globally reproducible.
This matters.
If Deep begins on luxury headsets only, it becomes another temple object. If it begins on phones and cardboard, it becomes a movement.
For Deep-0C, one camera is enough.
The phone’s front camera can track one eye well enough for coarse command intent:
look left
look right
look up
look down
hold gaze
blink
long blink
double blink
centre fixation
This supports the first control grammar:
select
confirm
cancel
inspect
pause
reset
return
Human eyes are usually coordinated for direction, so monocular tracking can provide a workable attention joystick for a symbolic interface.
The first goal is not precise medical-grade gaze science.
The first goal is:
Can one eye provide enough command bandwidth to navigate a symbolic archive?
If yes, Deep has a cheap first door.
A normal cardboard frame may block the selfie camera or place it at the wrong angle.
So Deep-0C needs a modified cardboard-style frame:
phone screen faces the eyes
front camera has clear view of one eye
small internal eye window or light tunnel
stable phone position
simple calibration grid
headphones or open-ear audio
easy physical removal
The frame does not need to be beautiful.
It needs to be reliable.
One camera can open the door.
Two inward-facing cameras are for proper Deep.
A dedicated headset with two eye cameras can unlock:
binocular gaze
vergence
rough depth focus
better drift correction
better blink confidence
better pupil telemetry
occlusion fallback
higher reliability across users
better support for glasses / face geometry variation
This gives a clean hardware ladder:
Stage 1: Smartphone + cardboard frame + exposed front camera
Stage 2: Smartphone + custom printed frame + better camera angle / light tunnel
Stage 3: Phone + external small eye camera module
Stage 4: Dedicated headset with two inward IR eye cameras
Stage 5: Deep-native headset
The rule is simple:
One eye is enough to open the door.
Two eyes are for measuring the room.
Do not wait for perfect hardware.
Build the symbolic/gaze software now.
Eyes are too valuable to waste on locomotion.
Eyes should be used for meaning:
select object
inspect object
confirm
reject
request detail
pause
exit
Head tilt should be used for motion:
move forward
move back
move left
move right
ascend
descend
slow
stop
A smartphone already has the required inertial sensors.
Deep-0C can use the phone IMU as a crude movement controller:
| Action | Control |
|---|---|
| Move forward | Tilt head slightly forward |
| Move backward | Tilt head slightly back |
| Move left | Roll head left |
| Move right | Roll head right |
| Slow / stop | Return head to neutral |
| Ascend / more abstract | Tilt up or use a dedicated gaze command |
| Descend / more detail | Tilt down or use a dedicated gaze command |
| Inspect | Fixate object |
| Confirm | Hold fixation |
| Return | Look at EXIT or use physical fallback |
Important distinction:
A cardboard phone rig mostly provides orientation, not true body translation.
So “move forward/back/up/down” in Deep-0C means:
tilt-controlled symbolic velocity
Not actual 6DOF physical movement.
That is fine.
Deep does not need realistic walking. It needs controlled relocation of attention.
Neutral head position must always mean stop.
No command should continue forever without confirmation.
Movement should be damped:
small tilt → slow symbolic movement
larger tilt → faster symbolic movement, capped
neutral → stop
erratic → pause or simplify
The user should never feel dragged.
The interface should feel like attention leaning into a space.
This becomes one of Deep’s core control laws:
Eyes are for meaning.
Head is for motion.
Voice is for guidance.
Hands are optional.
Exit is mandatory.
This separation matters.
If eyes control everything, the user cannot inspect without accidentally moving.
If head controls everything, the user loses precision.
If voice controls everything, the interface becomes slow and physically fragile.
The clean split is:
AUDIO → system speaks to the brain
EYES → brain points to meaning
HEAD → user moves through symbolic space
BUTTON → emergency fallback
This creates a practical low-cost control stack for Deep-0C.
Deep should use modes so movement does not conflict with inspection.
head tilt moves through the space
eyes select landmarks
audio guides transitions
head movement is damped or disabled
eyes inspect object details
perceptual keyframes can re-expand
audio explains or asks questions
all movement stops
objects freeze
EXIT dominates the field
room/passthrough or neutral view returns
audio grounds the user
Possible mode transitions:
fixate object → Inspect Mode
look away to horizon → Navigate Mode
look at EXIT → Return Mode
long blink → Pause
physical stop → Abort / Ground
This should be learned while fully awake before any deeper session begins.
Deep should not invent its control grammar from scratch.
There is already a large field of assistive technology for people with severe motor impairments, including ALS, locked-in syndrome, cerebral palsy, spinal cord injury, and other full-body or partial-body disabilities.
That field has spent decades asking the question Deep also asks:
How can a person control a computer when hands and voice are limited or unavailable?
Relevant areas include:
eye-gaze typing
dwell selection
blink-based commands
gaze-controlled wheelchairs
gaze-controlled robotics
eye-tracking AAC systems
head-motion interfaces
electrooculography
brain-computer interfaces
predictive language interfaces
Deep should learn from this work.
The user in Deep may not be disabled, but the interface constraints are similar:
hands unavailable
speech undesirable or too slow
body quiet
eyes available
head orientation available
attention is the main signal
This means disability-access research is not a side note.
It is the control-system foundation.
One of the most important ideas is dwell selection:
look at target → hold gaze → target activates
That maps directly to Deep.
A symbolic archive can use dwell to select, inspect, confirm, or return.
Assistive systems also show that the interface should not require the user to spell everything out.
It should predict likely intent and offer a small number of gaze-selectable choices.
In Deep, after a user focuses on a tree icon, the system might offer:
INSPECT
MOVE TO
COMPARE
REMEMBER
BACK
This is much faster than making the user navigate a giant menu.
Deep should be sparse, predictive, and forgiving.
A first Deep-0C command grammar could look like this:
Gaze:
centre fixation = hold / stabilize
fixate object = select
hold object = inspect / confirm
look left target = previous
look right target = next
look up target = more abstract
look down target = more detail
double blink = reset
long blink = pause
exit fixation 2 sec = return
Head:
neutral = stop
tilt forward = move forward
tilt back = move back
roll left = move left
roll right = move right
small nod = optional confirm
small shake = optional reject
Physical:
remove headset = hard exit
press button/touch = hard exit
Early sessions should use only a small subset.
Do not overload the user with commands.
The interface should grow like a language.
Deep has asymmetric channels.
Best channels:
audio guidance
symbolic visual field
perceptual keyframes
spatial markers
timing and rhythm
Audio remains the clean command layer because speech comprehension can stay natural, familiar, and low-friction.
Best channels:
gaze direction
fixation duration
saccade gestures
blink patterns
pupil changes
head movement
optional breath pattern
Voice is too slow and too physically expensive as a return channel, especially if the body is relaxed or near-immobile.
Eyes are faster.
Eyes can select, confirm, reject, navigate, inspect, and signal overload.
A Deep interface should not ask the user to speak commands while absorbed.
It should provide a fixed spatial control grammar.
Example:
CONTINUE
LEFT INSPECT RIGHT
BACK HOLD DETAIL
EXIT
Possible gaze grammar:
| Eye behaviour | Meaning |
|---|---|
| Fixate object | Select |
| Hold fixation | Confirm |
| Quick left-right-left | No / reject |
| Quick up-down-up | Yes / accept |
| Double blink | Reset / error |
| Long blink | Pause |
| Look to edge | Move / pan |
| Fixate EXIT for 2 seconds | Begin return |
| Erratic gaze | Possible overload, simplify scene |
The exact grammar must be calibrated per user. The point is that the interface should be learned while fully awake before any deeper session begins.
No calibrated eye language, no Deep session.
Photorealism may be a trap.
If the machine renders too much, the user becomes a spectator.
Deep wants the user to become an internal participant.
The right visual style may be:
- sparse;
- symbolic;
- low-detail;
- stable;
- high contrast but not flashing;
- semantically rich;
- low motion;
- predictable in layout;
- responsive to gaze.
Deep should feel less like a game world and more like a living map, memory palace, cockpit, archive, and dream notebook fused into one instrument.
Movement should not imitate walking unless the hardware and user state support it safely.
VR sickness is often linked to conflict between visual motion and vestibular/body signals. If the screen says “we are moving” but the body says “we are still,” nausea, dizziness, disorientation, and discomfort can appear.
Deep should prefer:
gaze target → soft fade → scene re-centre → audio confirms transition
Instead of:
camera rushes forward through a corridor
No rollercoaster.
No sudden full-field motion.
No aggressive artificial locomotion.
Movement in Deep should feel like attention relocating, not a body being dragged.
The body should not be the main interface.
The body should become quiet.
But quiet is not the same as paralyzed.
The target body state for early Deep systems should be:
comfortable
stable
still
relaxed
non-essential
easy to exit from
Not:
trapped
unable to move
unable to speak
panicked
dissociated
The system should never punish movement. If the user moves, the system treats it as telemetry.
movement detected → reduce depth → stabilize → ask continue / pause / return
The purpose is not to defeat the body. The purpose is to stop requiring the body for ordinary interface control.
Deep should be implemented as a state machine, not a vibe.
CALIBRATE
↓
SETTLE
↓
INDUCE
↓
ANCHOR
↓
DEEP
↓
RETURN
↓
GROUND
↓
NORMAL
From every state:
ANY STATE → ABORT → GROUND → NORMAL
The system must always know where the user is in the session.
The system must always know how to leave.
- eye tracking works;
- blink detection works;
- gaze targets are readable;
- exit gesture is learned;
- physical stop is reachable;
- audio is confirmed;
- user is fully awake and oriented.
- reduce visual clutter;
- introduce stable anchor;
- slow interaction;
- establish breathing or attention rhythm if appropriate;
- verify comfort.
- narrow attention;
- present simple visual field;
- introduce first archive markers;
- avoid rapid motion;
- avoid flashing;
- keep audio calm and concrete.
- introduce object markers;
- use perceptual keyframes;
- bind object tokens;
- verify gaze selection;
- verify user can pause and exit.
- user navigates symbolic archive;
- objects expand and collapse;
- gaze controls selection;
- audio provides guidance;
- system monitors overload and distress.
- freeze object movement;
- collapse icons to markers;
- restore room/passthrough;
- widen field;
- shift audio to normal grounding voice;
- ask for simple confirmation.
- user moves fingers;
- user turns head;
- user confirms orientation;
- session summary appears;
- headset can be removed.
- session ends;
- logs are saved;
- user can review objects, actions, discomfort, and recall.
The first commandment of Deep:
The way out is part of the system.
Exit must be easier than entry.
A Deep system should have multiple independent return paths:
| Exit channel | Mechanism |
|---|---|
| Gaze exit | Fixed EXIT glyph, same location every session |
| Blink exit | Long blink or calibrated blink pattern |
| Physical exit | Real button, not hidden in software |
| Voice exit | User can say stop if able |
| Timeout exit | Session returns automatically |
| Watchdog exit | No response or distress triggers return |
| Supervisor exit | External observer can terminate session in early testing |
No single channel should be trusted alone.
The exit system should be boring, repetitive, and reliable.
Return should be concrete, not mystical.
Bad:
Awaken, traveler, from the mirror-realm.
Better:
The scene is now still.
The markers are closing.
You are in the room.
Feel the chair.
Move your fingers.
Look at the exit marker.
The headset is returning to normal view.
Session complete.
Exit is not theatre.
Exit is engineering.
Deep must be designed around conservative constraints from the beginning.
- no rapid full-field flashing;
- no saturated red strobe patterns;
- no hidden flicker tricks;
- no surprise brightness explosions;
- no forced subliminal content;
- no high-contrast repetitive patterns without testing;
- no dense motion fields during induction;
- no unlogged perceptual keyframes.
WCAG guidance around flashing content is relevant here: content flashing more than three times per second can be a seizure risk for susceptible users. Deep should be stricter than ordinary media because it is immersive.
- avoid artificial locomotion early;
- prefer teleport/fade/re-centre;
- maintain stable horizon when possible;
- detect head/eye instability;
- allow immediate pause;
- keep early sessions short;
- log discomfort.
- avoid fear induction;
- avoid forced loss-of-control framing;
- avoid “you cannot move” language;
- avoid identity manipulation;
- avoid false certainty;
- avoid content that blurs fiction and fact without clear labelling.
- always-visible exit;
- physical stop;
- maximum duration;
- no “infinite mode”;
- full session logging;
- explicit consent;
- post-session check-in;
- content provenance.
Deep content can feel more real than text.
That is dangerous.
Therefore every object should visually carry its epistemic status.
Example visual grammar:
| Status | Visual treatment |
|---|---|
| Observed fact | Solid |
| Hypothesis | Translucent |
| Memory | Warm glow |
| Fiction | Distinct border |
| Simulation | Grid overlay |
| Contradiction | Broken line |
| Uncertain | Hollow shape |
| Warning | Red pulse, used sparingly |
| User-created | Hand-drawn edge |
| AI-generated | Machine glyph edge |
The user must know whether an object is fact, memory, theory, fiction, or generated suggestion.
Deep without epistemology becomes persuasion machinery.
Deep with epistemology becomes cognitive instrumentation.
A future Deep Archive might contain:
deep_archive:
metadata:
title: "Field Training Scenario 01"
version: "0.1"
author: "example"
safety_profile: "low_motion_symbolic"
max_session_minutes: 5
spaces:
- id: field_001
type: symbolic_environment
background: low_detail_grass_field
horizon: stable
objects:
- tree_001
- path_001
- gate_001
objects:
- id: tree_001
category: tree
marker_far: dot_green
marker_mid: vertical_shape_green
perceptual_keyframe: oak_tree_keyframe_01
token: tree_icon
detail_levels:
- silhouette
- trunk_canopy
- bark_branch
audio_label: "old oak"
epistemic_status: observed
confidence: 0.92
relations:
- beside: path_001
- id: gate_001
category: gate
marker_far: dot_black
marker_mid: rectangle_dark
perceptual_keyframe: old_gate_keyframe_01
token: gate_icon
epistemic_status: uncertain
confidence: 0.61
relations:
- blocks: path_001
controls:
gaze:
fixation_select_ms: 650
fixation_confirm_ms: 1200
exit_fixation_ms: 2000
blink:
double_blink: reset
long_blink_ms: 1500
long_blink_action: pause
return_protocol:
always_available: true
auto_return_on_no_input_seconds: 20
auto_return_on_distress: trueThis is not a media file.
It is a cognitive scene package.
AI is not the interface.
AI is the scene composer, archivist, narrator, simplifier, and safety assistant.
Given a document, project, memory set, training task, or problem graph, AI can:
extract objects
classify entities
generate symbols
generate keyframes
map relationships
assign confidence
build spaces
write audio guidance
adapt complexity
detect overload
summarise session output
This turns many forms of information into navigable visual-symbolic space.
AI becomes less “answer bot” and more cartographer of thought.
Deep can be used to build mental models faster than slides or flat video.
Example: aviation, medicine, maritime navigation, emergency response, engineering diagnostics.
Instead of reading:
Watch for crosswind drift.
The user sees:
runway marker
wind glyph
aircraft token
drift vector
correction path
The system briefly expands the scene at recognition points, then collapses back to symbols.
The user learns the relationship spatially.
Not memorisation.
Model-building.
Deep can turn memory palaces into programmable archives.
A user can load:
- a legal case;
- a codebase;
- a speech;
- a song structure;
- a language vocabulary set;
- an aircraft checklist;
- a research map;
- a personal project.
Deep converts it into spaces, objects, paths, and glyphs.
The user enters the archive, binds objects, follows relations, and exits with structured recall.
This is not note-taking.
It is spatialised memory construction.
Many problems are not linear.
They are constellations:
actors
events
documents
contradictions
hypotheses
missing evidence
dependencies
risks
timelines
Flat screens struggle with this.
Deep can let the user inhabit the structure.
A fraud investigation becomes a room of objects and threads.
A codebase becomes a machine landscape.
A medical differential becomes competing paths and warning glyphs.
A strategy becomes weather, terrain, constraints, and moving pieces.
The user does not merely read the analysis.
The user enters the analysis.
Deep could be monstrous for creative work.
A song becomes a place.
A music video becomes a moving symbolic stage.
A story becomes a navigable myth-map.
A character becomes an object with memory, voice, costume, conflict, and transformations.
A chorus becomes a pulse in space.
The artist can move by gaze, bind scenes with keyframes, collapse them to icons, and re-expand details.
Not prompt engineering.
Scene gardening.
The hardware is becoming available.
The difficult part is the software.
The soft must:
invite absorption
calibrate gaze
present sparse markers
time perceptual keyframes
avoid overload
keep audio coherent
interpret eye signals
control motion
log state
preserve exit
return reliably
The software does not “force” Deep.
It creates conditions where the user’s brain can enter a useful state voluntarily.
That is the difference between an instrument and a trap.
Deep-0 should not claim altered states.
Deep-0 should test the core interaction.
- Can users bind a detailed object to a simple icon?
- Can they navigate objects by gaze?
- Can they recall more after a symbolic VR session than after a flat screen?
- Can the system safely return the user every time?
- Can the system detect confusion or overload?
Duration: 3–5 minutes
Environment: simple field or room
Objects: 3–7
Motion: minimal
Controls: gaze + blink + physical stop
Audio: simple guidance
Exit: always visible
find object
select object
remember object
compare object
track relationship
return and describe
Deep-0 success is not “the user went into a trance.”
Deep-0 success is:
The user can bind, navigate, remember, and return.
The near-term product should be an awake, ordinary prototype.
It should test the smallest useful loop:
marker → visible keyframe → icon binding → gaze/head command → return → recall check
This can be built now with:
- a phone or commodity headset;
- a simple symbolic scene;
- three to seven objects;
- visible perceptual keyframes;
- coarse gaze or camera-based attention detection;
- head-tilt symbolic movement;
- audio guidance;
- a fixed exit glyph;
- a physical stop path;
- session logs;
- post-session recall questions.
It should not attempt:
- sleep paralysis;
- clinical claims;
- subliminal frames;
- high-risk psychological content;
- long sessions;
- photorealistic locomotion;
- unlogged stimuli;
- medical or therapeutic use.
The long-term Deep vision can remain ambitious.
The first build should be modest, measurable, and easy to exit.
Deep becomes real only if it survives measurement.
Key questions:
- Do perceptual keyframes improve object recall?
- Does icon binding reduce cognitive load?
- Does gaze-only navigation preserve absorption better than hand controllers?
- Does symbolic rendering reduce VR sickness compared to photorealistic movement?
- Can users maintain dynamic object constellations more effectively in Deep than on flat screens?
- What are the safest and most reliable return cues?
- How does individual variation affect depth, comfort, recall, and control?
- Can overload be detected from gaze instability, blink rate, head movement, or task errors?
- What content categories are inappropriate for Deep?
- How should fact, fiction, memory, and hypothesis be visually distinguished?
- The user must always be able to return.
- Do not force paralysis.
- Do not hide stimuli.
- Do not use subliminal manipulation.
- Use the brain as renderer, not victim.
- Use audio for guidance.
- Use eyes for control.
- Use symbols before spectacle.
- Use keyframes for binding, not persuasion.
- Represent uncertainty visually.
- Prefer stillness over motion.
- Prefer attention jumps over artificial locomotion.
- Make exit easier than entry.
- Log what was shown.
- Treat safety as architecture, not disclaimer.
DeepOS
├─ Safety Shell
├─ User Calibration
├─ Gaze Runtime
├─ Blink Runtime
├─ Visual Archive Renderer
├─ Perceptual Keyframe Scheduler
├─ Symbolic Object Engine
├─ Relationship Graph Engine
├─ Audio Guidance Layer
├─ Epistemology Renderer
├─ Overload Detector
├─ Return Protocol
├─ Session Logger
└─ Archive Exporter
The most important module is not the renderer.
It is the Return Protocol.
If Deep works, it will be powerful.
That means it can be abused.
Possible risks:
- persuasion through felt reality;
- false memory reinforcement;
- emotional manipulation;
- over-trust in AI-generated scenes;
- content that feels proven because it was spatially experienced;
- loss of distinction between evidence and simulation;
- surveillance through gaze and attention data.
Therefore Deep must be built with:
consent
audit logs
content provenance
epistemic visual grammar
local-first privacy where possible
clear session boundaries
user ownership of archives
export and delete rights
Attention data is intimate.
Gaze is not just cursor movement. It is behavioural telemetry.
Deep must treat it as sensitive.
We are tired of flat screens pretending to be enough.
We are tired of chatbots answering with confident fog.
We are tired of interfaces that bury thought in tabs, panels, documents, notifications, and tiny rectangles.
The human brain can do more.
It can hold landscapes.
It can track constellations.
It can bind symbols to memory.
It can hear a voice and build a world.
It can move through an archive without touching a keyboard.
Deep is an attempt to build for that brain.
Not to overwhelm it.
Not to trick it.
Not to trap it.
To give it a structured place where imagination becomes an interface.
The machine does not need to draw every leaf.
It only needs to show the right dot, at the right distance, at the right moment.
Then the tree appears where it always appeared:
inside the human mind.
Deep is a reversible gaze-controlled symbolic cognition interface that uses sparse visual markers, visible perceptual keyframes, and audio guidance to let the human brain reconstruct and navigate complex information as an internal spatial model.
Shorter:
Deep is an operating system for directed imagination.
Shortest:
VR supplies the glyphs.
Audio supplies the grammar.
The brain supplies the world.
The eyes supply the API.
There is nothing impossible.
We are building Deep.
But we build it with one law carved into the foundation:
The way out is part of the system.
This document is speculative. The concept should be treated as research/design exploration, not validated neuroscience or medical guidance. The following references ground some safety and feasibility constraints:
-
Cleveland Clinic, “Sleep Paralysis: What It Is, Causes, Symptoms & Treatment.” Notes that sleep paralysis can include inability to move or speak, chest pressure, and hallucinations.
https://my.clevelandclinic.org/health/diseases/21974-sleep-paralysis -
W3C Web Accessibility Initiative, WCAG guidance on flashing content. Notes that flashing content can trigger seizures in susceptible people and recommends strict limits around flashing.
https://www.w3.org/WAI/WCAG21/Understanding/three-flashes.html -
Frontiers in Virtual Reality, “Effects of Linear Visual-Vestibular Conflict on Presence, Perceived Self-Motion, and Cybersickness.” Discusses visual-vestibular mismatch as a contributor to cybersickness in VR.
https://www.frontiersin.org/journals/virtual-reality/articles/10.3389/frvir.2021.582156/full -
Weech, Kenny, Barnett-Cowan, “Presence and Cybersickness in Virtual Reality Are Negatively Related: A Review.” Reviews presence and cybersickness in VR.
https://pmc.ncbi.nlm.nih.gov/articles/PMC6369189/ -
Konkoly et al., “Real-time dialogue between experimenters and dreamers.” Current Biology, 2021. Demonstrates limited two-way communication with lucid dreamers during verified REM sleep using signals such as eye movements and facial muscle contractions.
https://pubmed.ncbi.nlm.nih.gov/33607035/ -
Baird et al., “The cognitive neuroscience of lucid dreaming.” Reviews lucid dreaming and eye-movement signalling research.
https://pmc.ncbi.nlm.nih.gov/articles/PMC6451677/ -
Microsoft Support, “Eye control basics in Windows.” Useful mainstream example of dwell selection, eye-control launchpads, mouse control, scrolling, eye keyboard, and text-to-speech.
https://support.microsoft.com/en-us/windows/eye-control-basics-in-windows -
Nature Communications, 2024, “Enhancing communication in ALS using an LLM-assisted eye-gaze typing system.” Shows how prediction can reduce the burden of eye-gaze typing and improve communication rates.
https://www.nature.com/articles/s41467-024-53873-3 -
Frontiers in Virtual Reality, 2023, “Evaluation of head joystick locomotion control in virtual reality.” Discusses head-orientation/head-lean style locomotion, calibration, and comfort constraints.
https://www.frontiersin.org/journals/virtual-reality/articles/10.3389/frvir.2023.1169654/full -
Review literature on eye-gaze controlled assistive robotics and interfaces for severe motor impairment is directly relevant to Deep’s “eyes as API” control layer.
https://pmc.ncbi.nlm.nih.gov/articles/PMC10909843/