Skip to content

Instantly share code, notes, and snippets.

@Co0olCat
Last active June 13, 2026 07:55
Show Gist options
  • Select an option

  • Save Co0olCat/9323572a4194f64e1573e8118a71f08e to your computer and use it in GitHub Desktop.

Select an option

Save Co0olCat/9323572a4194f64e1573e8118a71f08e to your computer and use it in GitHub Desktop.
Deep Draft: a speculative design for a reversible gaze-controlled symbolic cognition interface, starting with Deep-0C: a smartphone/cardboard prototype where sparse glyphs, perceptual keyframes, gaze commands, head-tilt movement, and reliable return form the first door into directed imagination.

There Is Nothing Impossible: We Are Building Deep

Version: v0.4 Draft

Status: speculative design document, not medical advice, not a hypnosis protocol, not a claim of clinical validation

Author: Timur Yusupov, PhD

Working concept: Deep as a reversible gaze-controlled symbolic cognition interface

Attribution / inspiration: This work is my original reflection on the fictional technology described in Лабиринт отражений (Labyrinth of Reflections) by Сергей Лукьяненко (Sergey Lukyanenko). It is not affiliated with or endorsed by the author or rights holders.

Language note: Phrases such as “brain as renderer” and “brain-side reconstruction” are design metaphors in this draft. They describe an interface strategy, not a validated neuroscience claim, unless and until tested.

We are not trying to build a better screen.
We are trying to build a new state of interaction between human imagination and machine structure.


Table of Contents


1. The Short Version

Deep is a proposed human-computer interface mode where:

  • the computer does not render a complete virtual world;
  • the headset presents low-resolution visual archive markers;
  • the brain performs the high-resolution reconstruction;
  • brief perceptual keyframes bind object identity;
  • simple icons then stand in for rich objects;
  • audio remains the normal command and guidance channel;
  • eye movement, gaze, focus, blinks, and fixation become the return channel;
  • the system is built around one sacred requirement: the user must always be able to return.

In a normal VR system, the machine tries to create reality.

In Deep, the machine gives the brain structured seeds, and the brain grows the world.


2. What Deep Is

Deep is a proposed interface for directed imagination.

It is not simply VR. It is not just hypnosis. It is not a cinematic dream simulator. It is not a chatbot wearing goggles.

It is closer to a new operating mode:

visual archive markers → brain-side reconstruction → gaze-controlled navigation → audio-guided cognition

The computer sends symbolic fragments.
The user’s brain unpacks them.
The eyes steer the system.
The voice layer keeps continuity.
The return protocol protects the user.

Deep is therefore best described as:

A reversible symbolic VR protocol that uses the human brain as the high-resolution renderer, eye movement as the control channel, and audio as the guidance layer.

Or shorter:

An operating system for directed imagination.


3. What Deep Is Not

Deep is not a proposal to force sleep paralysis.

Deep is not a product that traps a user in an altered state.

Deep is not subliminal manipulation.

Deep is not hidden-frame advertising in goggles.

Deep is not a medical device concept unless and until it is evaluated under the appropriate medical, ethical, and regulatory frameworks.

Deep is not a replacement for consent, context, grounding, or ordinary human judgement.

The safest starting point is not “put people into sleep paralysis.” The safest starting point is:

Voluntary body quieting + high visual absorption + gaze control + reliable return.

Sleep paralysis is a real phenomenon, but it can involve inability to move or speak, chest pressure, hallucinations, fear, and distress. It should not be treated as a casual interaction target. The product target should be absorption, not paralysis.


4. The Core Insight

Most interfaces assume the screen must carry the full representation.

Deep assumes the opposite:

The screen only needs to carry the right trigger at the right moment.

Human visual cognition can construct scenes from hints, fragments, symbols, movement, expectation, and memory.

When this draft says the brain is a renderer, it means that as a design metaphor: the system tries to provide useful cues and let ordinary perception, memory, and imagination fill in detail. It is not a claim that Deep has validated a specific neural mechanism.

A computer does not need to continuously show every leaf on every tree.

It can show:

  1. a distant marker;
  2. a growing shape;
  3. a brief detailed recognition frame;
  4. a stable symbolic icon.

After the object is bound, the brain carries the rest.

This is the central compression trick.


5. The Tree Example

Imagine a user walking through a green field.

At first, there is only grass.

Far ahead, the user notices a dot.

The dot grows as the user approaches.

At some point the brain begins expecting recognition:

“That is probably something vertical. Maybe a tree.”

At that exact recognition threshold, the system shows a brief detailed image: trunk, branches, shadow, canopy, maybe bark.

Then the detailed image collapses into a simple tree icon.

From that moment onward, the icon is enough.

The user no longer sees “generic tree symbol.”
The user sees that tree.

The brain has been conditioned through a visible, intentional, logged perceptual keyframe.

The runtime sequence is:

dot → growing dot → expectation → perceptual keyframe → tree icon → brain-side continuity

This means Deep does not need constant high-detail display. It needs well-timed identity binding.


6. Perceptual Keyframes, Not Hidden Frames

The phrase “25th frame” may be useful as historical shorthand, but this document retires it after this sentence. It carries baggage from subliminal advertising, hidden persuasion, and cinematic mythology.

Deep should avoid hidden influence.

The better term is:

Perceptual Keyframe

A perceptual keyframe is a brief, visible, intentional, logged high-information image used at the moment of expected recognition to bind an object’s identity.

It is not secret.

It is not subliminal.

It is not manipulation by stealth.

It is a cognitive anchor.

Object marker + recognition timing + perceptual keyframe = memory-bound object token

After that, a simple icon can represent the object because the user’s brain has already cached its identity.


7. The Visual Archive

A visual archive is not a video file.

It is a structured object store for cognitive rendering.

Each object can have:

object:
  id: tree_001
  category: tree
  far_marker: green_dot
  approach_marker: vertical_soft_shape
  perceptual_keyframe: tree_oak_keyframe_01
  token: tree_icon_simple
  detail_levels:
    - silhouette
    - trunk_and_canopy
    - bark_and_branch_detail
  audio_label: "old oak near the path"
  relations:
    - beside: path_003
    - north_of: field_001
  epistemic_status: observed
  confidence: 0.92

The object is not “drawn” continuously. It is activated.

The archive stores enough to let the brain rebuild the scene without forcing the machine to render all details all the time.


8. The Deep Rendering Principle

Normal rendering:

machine renders detail → user observes detail

Deep rendering:

machine triggers identity → brain renders detail → machine tracks attention

This is a design principle to be tested, not an established empirical result.

Here, “brain renders detail” means the interface deliberately relies on visible cues, memory, expectation, and imagination instead of forcing the display to carry every detail.

This reverses the usual burden.

The display becomes an index.

The brain becomes the renderer, in the metaphorical sense defined above.

The eyes become the API.


9. Hardware Assumption

The ideal hardware is not exotic.

It is a VR or mixed-reality headset with:

  • high refresh rate;
  • low latency;
  • inward-facing eye cameras;
  • gaze tracking;
  • fixation tracking;
  • blink detection;
  • pupil tracking where safe and meaningful;
  • head movement tracking;
  • open-ear or high-quality audio;
  • optional pass-through cameras;
  • physical emergency stop;
  • optional breathing or pulse sensor;
  • clear session logging.

The headset is not used to create a photorealistic world. It is used to control visual input and track the user’s gaze.

The most important hardware feature is not resolution. It is closed-loop attention tracking.


10. Deep-Lite: Cardboard Mode

The first real deployment wedge may not be an expensive headset.

It may be:

smartphone + cardboard/plastic frame + front camera + headphones

This is important because a modern smartphone already contains most of the required prototype hardware:

screen
front camera
speaker/headphones
IMU / gyro / accelerometer
battery
compute
storage
network
touch fallback

Deep does not require photorealistic VR for the first prototype.

Deep-0C only needs to prove the core loop:

visual marker → perceptual keyframe → icon binding → gaze command → safe return

That can be tested with commodity hardware.

Deep-0C

Deep-0C means:

Deep-0 = first safe prototype
C      = cardboard / commodity camera

The purpose of Deep-0C is not to build perfect eye tracking or perfect immersion.

The purpose is to test whether a person can:

  • bind a detailed perceptual keyframe to a simple icon;
  • navigate the resulting symbolic world by gaze;
  • use head tilt for movement;
  • return safely and consistently;
  • recall the scene after exit.

The first prototype can be cheap, hackable, and globally reproducible.

This matters.

If Deep begins on luxury headsets only, it becomes another temple object. If it begins on phones and cardboard, it becomes a movement.


11. One Camera Is Enough to Open the Door

For Deep-0C, one camera is enough.

The phone’s front camera can track one eye well enough for coarse command intent:

look left
look right
look up
look down
hold gaze
blink
long blink
double blink
centre fixation

This supports the first control grammar:

select
confirm
cancel
inspect
pause
reset
return

Human eyes are usually coordinated for direction, so monocular tracking can provide a workable attention joystick for a symbolic interface.

The first goal is not precise medical-grade gaze science.

The first goal is:

Can one eye provide enough command bandwidth to navigate a symbolic archive?

If yes, Deep has a cheap first door.

Hardware Requirement

A normal cardboard frame may block the selfie camera or place it at the wrong angle.

So Deep-0C needs a modified cardboard-style frame:

phone screen faces the eyes
front camera has clear view of one eye
small internal eye window or light tunnel
stable phone position
simple calibration grid
headphones or open-ear audio
easy physical removal

The frame does not need to be beautiful.

It needs to be reliable.


12. Two Cameras Are for Measuring the Room

One camera can open the door.

Two inward-facing cameras are for proper Deep.

A dedicated headset with two eye cameras can unlock:

binocular gaze
vergence
rough depth focus
better drift correction
better blink confidence
better pupil telemetry
occlusion fallback
higher reliability across users
better support for glasses / face geometry variation

This gives a clean hardware ladder:

Stage 1: Smartphone + cardboard frame + exposed front camera
Stage 2: Smartphone + custom printed frame + better camera angle / light tunnel
Stage 3: Phone + external small eye camera module
Stage 4: Dedicated headset with two inward IR eye cameras
Stage 5: Deep-native headset

The rule is simple:

One eye is enough to open the door.
Two eyes are for measuring the room.

Do not wait for perfect hardware.

Build the symbolic/gaze software now.


13. Head Tilt as the Movement Bus

Eyes are too valuable to waste on locomotion.

Eyes should be used for meaning:

select object
inspect object
confirm
reject
request detail
pause
exit

Head tilt should be used for motion:

move forward
move back
move left
move right
ascend
descend
slow
stop

A smartphone already has the required inertial sensors.

Deep-0C can use the phone IMU as a crude movement controller:

Action Control
Move forward Tilt head slightly forward
Move backward Tilt head slightly back
Move left Roll head left
Move right Roll head right
Slow / stop Return head to neutral
Ascend / more abstract Tilt up or use a dedicated gaze command
Descend / more detail Tilt down or use a dedicated gaze command
Inspect Fixate object
Confirm Hold fixation
Return Look at EXIT or use physical fallback

Important distinction:

A cardboard phone rig mostly provides orientation, not true body translation.

So “move forward/back/up/down” in Deep-0C means:

tilt-controlled symbolic velocity

Not actual 6DOF physical movement.

That is fine.

Deep does not need realistic walking. It needs controlled relocation of attention.

Movement Safety

Neutral head position must always mean stop.

No command should continue forever without confirmation.

Movement should be damped:

small tilt  → slow symbolic movement
larger tilt → faster symbolic movement, capped
neutral     → stop
erratic     → pause or simplify

The user should never feel dragged.

The interface should feel like attention leaning into a space.


14. Eyes for Meaning, Head for Motion

This becomes one of Deep’s core control laws:

Eyes are for meaning.
Head is for motion.
Voice is for guidance.
Hands are optional.
Exit is mandatory.

This separation matters.

If eyes control everything, the user cannot inspect without accidentally moving.

If head controls everything, the user loses precision.

If voice controls everything, the interface becomes slow and physically fragile.

The clean split is:

AUDIO  → system speaks to the brain
EYES   → brain points to meaning
HEAD   → user moves through symbolic space
BUTTON → emergency fallback

This creates a practical low-cost control stack for Deep-0C.


15. Control Modes

Deep should use modes so movement does not conflict with inspection.

Navigate Mode

head tilt moves through the space
eyes select landmarks
audio guides transitions

Inspect Mode

head movement is damped or disabled
eyes inspect object details
perceptual keyframes can re-expand
audio explains or asks questions

Return Mode

all movement stops
objects freeze
EXIT dominates the field
room/passthrough or neutral view returns
audio grounds the user

Possible mode transitions:

fixate object → Inspect Mode
look away to horizon → Navigate Mode
look at EXIT → Return Mode
long blink → Pause
physical stop → Abort / Ground

This should be learned while fully awake before any deeper session begins.


16. Assistive Technology Is the Treasure Cave

Deep should not invent its control grammar from scratch.

There is already a large field of assistive technology for people with severe motor impairments, including ALS, locked-in syndrome, cerebral palsy, spinal cord injury, and other full-body or partial-body disabilities.

That field has spent decades asking the question Deep also asks:

How can a person control a computer when hands and voice are limited or unavailable?

Relevant areas include:

eye-gaze typing
dwell selection
blink-based commands
gaze-controlled wheelchairs
gaze-controlled robotics
eye-tracking AAC systems
head-motion interfaces
electrooculography
brain-computer interfaces
predictive language interfaces

Deep should learn from this work.

The user in Deep may not be disabled, but the interface constraints are similar:

hands unavailable
speech undesirable or too slow
body quiet
eyes available
head orientation available
attention is the main signal

This means disability-access research is not a side note.

It is the control-system foundation.

Dwell Selection

One of the most important ideas is dwell selection:

look at target → hold gaze → target activates

That maps directly to Deep.

A symbolic archive can use dwell to select, inspect, confirm, or return.

Predictive Commands

Assistive systems also show that the interface should not require the user to spell everything out.

It should predict likely intent and offer a small number of gaze-selectable choices.

In Deep, after a user focuses on a tree icon, the system might offer:

INSPECT
MOVE TO
COMPARE
REMEMBER
BACK

This is much faster than making the user navigate a giant menu.

Deep should be sparse, predictive, and forgiving.


17. Deep-Lite Control Grammar

A first Deep-0C command grammar could look like this:

Gaze:
  centre fixation        = hold / stabilize
  fixate object          = select
  hold object            = inspect / confirm
  look left target       = previous
  look right target      = next
  look up target         = more abstract
  look down target       = more detail
  double blink           = reset
  long blink             = pause
  exit fixation 2 sec    = return

Head:
  neutral                = stop
  tilt forward           = move forward
  tilt back              = move back
  roll left              = move left
  roll right             = move right
  small nod              = optional confirm
  small shake            = optional reject

Physical:
  remove headset         = hard exit
  press button/touch     = hard exit

Early sessions should use only a small subset.

Do not overload the user with commands.

The interface should grow like a language.


18. Input and Output

Deep has asymmetric channels.

System to Brain

Best channels:

audio guidance
symbolic visual field
perceptual keyframes
spatial markers
timing and rhythm

Audio remains the clean command layer because speech comprehension can stay natural, familiar, and low-friction.

Brain to System

Best channels:

gaze direction
fixation duration
saccade gestures
blink patterns
pupil changes
head movement
optional breath pattern

Voice is too slow and too physically expensive as a return channel, especially if the body is relaxed or near-immobile.

Eyes are faster.

Eyes can select, confirm, reject, navigate, inspect, and signal overload.


19. The Eye Interface

A Deep interface should not ask the user to speak commands while absorbed.

It should provide a fixed spatial control grammar.

Example:

              CONTINUE

   LEFT        INSPECT        RIGHT

   BACK        HOLD           DETAIL

              EXIT

Possible gaze grammar:

Eye behaviour Meaning
Fixate object Select
Hold fixation Confirm
Quick left-right-left No / reject
Quick up-down-up Yes / accept
Double blink Reset / error
Long blink Pause
Look to edge Move / pan
Fixate EXIT for 2 seconds Begin return
Erratic gaze Possible overload, simplify scene

The exact grammar must be calibrated per user. The point is that the interface should be learned while fully awake before any deeper session begins.

No calibrated eye language, no Deep session.


20. Deep Is Not Photorealism

Photorealism may be a trap.

If the machine renders too much, the user becomes a spectator.

Deep wants the user to become an internal participant.

The right visual style may be:

  • sparse;
  • symbolic;
  • low-detail;
  • stable;
  • high contrast but not flashing;
  • semantically rich;
  • low motion;
  • predictable in layout;
  • responsive to gaze.

Deep should feel less like a game world and more like a living map, memory palace, cockpit, archive, and dream notebook fused into one instrument.


21. Movement in Deep

Movement should not imitate walking unless the hardware and user state support it safely.

VR sickness is often linked to conflict between visual motion and vestibular/body signals. If the screen says “we are moving” but the body says “we are still,” nausea, dizziness, disorientation, and discomfort can appear.

Deep should prefer:

gaze target → soft fade → scene re-centre → audio confirms transition

Instead of:

camera rushes forward through a corridor

No rollercoaster.

No sudden full-field motion.

No aggressive artificial locomotion.

Movement in Deep should feel like attention relocating, not a body being dragged.


22. The Body State

The body should not be the main interface.

The body should become quiet.

But quiet is not the same as paralyzed.

The target body state for early Deep systems should be:

comfortable
stable
still
relaxed
non-essential
easy to exit from

Not:

trapped
unable to move
unable to speak
panicked
dissociated

The system should never punish movement. If the user moves, the system treats it as telemetry.

movement detected → reduce depth → stabilize → ask continue / pause / return

The purpose is not to defeat the body. The purpose is to stop requiring the body for ordinary interface control.


23. The State Machine

Deep should be implemented as a state machine, not a vibe.

CALIBRATE
  ↓
SETTLE
  ↓
INDUCE
  ↓
ANCHOR
  ↓
DEEP
  ↓
RETURN
  ↓
GROUND
  ↓
NORMAL

From every state:

ANY STATE → ABORT → GROUND → NORMAL

The system must always know where the user is in the session.

The system must always know how to leave.


24. State Definitions

CALIBRATE

  • eye tracking works;
  • blink detection works;
  • gaze targets are readable;
  • exit gesture is learned;
  • physical stop is reachable;
  • audio is confirmed;
  • user is fully awake and oriented.

SETTLE

  • reduce visual clutter;
  • introduce stable anchor;
  • slow interaction;
  • establish breathing or attention rhythm if appropriate;
  • verify comfort.

INDUCE

  • narrow attention;
  • present simple visual field;
  • introduce first archive markers;
  • avoid rapid motion;
  • avoid flashing;
  • keep audio calm and concrete.

ANCHOR

  • introduce object markers;
  • use perceptual keyframes;
  • bind object tokens;
  • verify gaze selection;
  • verify user can pause and exit.

DEEP

  • user navigates symbolic archive;
  • objects expand and collapse;
  • gaze controls selection;
  • audio provides guidance;
  • system monitors overload and distress.

RETURN

  • freeze object movement;
  • collapse icons to markers;
  • restore room/passthrough;
  • widen field;
  • shift audio to normal grounding voice;
  • ask for simple confirmation.

GROUND

  • user moves fingers;
  • user turns head;
  • user confirms orientation;
  • session summary appears;
  • headset can be removed.

NORMAL

  • session ends;
  • logs are saved;
  • user can review objects, actions, discomfort, and recall.

25. Return Is Part of the Product

The first commandment of Deep:

The way out is part of the system.

Exit must be easier than entry.

A Deep system should have multiple independent return paths:

Exit channel Mechanism
Gaze exit Fixed EXIT glyph, same location every session
Blink exit Long blink or calibrated blink pattern
Physical exit Real button, not hidden in software
Voice exit User can say stop if able
Timeout exit Session returns automatically
Watchdog exit No response or distress triggers return
Supervisor exit External observer can terminate session in early testing

No single channel should be trusted alone.

The exit system should be boring, repetitive, and reliable.


26. The Return Script

Return should be concrete, not mystical.

Bad:

Awaken, traveler, from the mirror-realm.

Better:

The scene is now still.
The markers are closing.
You are in the room.
Feel the chair.
Move your fingers.
Look at the exit marker.
The headset is returning to normal view.
Session complete.

Exit is not theatre.

Exit is engineering.


27. Safety Constraints

Deep must be designed around conservative constraints from the beginning.

Visual Safety

  • no rapid full-field flashing;
  • no saturated red strobe patterns;
  • no hidden flicker tricks;
  • no surprise brightness explosions;
  • no forced subliminal content;
  • no high-contrast repetitive patterns without testing;
  • no dense motion fields during induction;
  • no unlogged perceptual keyframes.

WCAG guidance around flashing content is relevant here: content flashing more than three times per second can be a seizure risk for susceptible users. Deep should be stricter than ordinary media because it is immersive.

VR Comfort

  • avoid artificial locomotion early;
  • prefer teleport/fade/re-centre;
  • maintain stable horizon when possible;
  • detect head/eye instability;
  • allow immediate pause;
  • keep early sessions short;
  • log discomfort.

Psychological Safety

  • avoid fear induction;
  • avoid forced loss-of-control framing;
  • avoid “you cannot move” language;
  • avoid identity manipulation;
  • avoid false certainty;
  • avoid content that blurs fiction and fact without clear labelling.

Operational Safety

  • always-visible exit;
  • physical stop;
  • maximum duration;
  • no “infinite mode”;
  • full session logging;
  • explicit consent;
  • post-session check-in;
  • content provenance.

28. The Epistemology Layer

Deep content can feel more real than text.

That is dangerous.

Therefore every object should visually carry its epistemic status.

Example visual grammar:

Status Visual treatment
Observed fact Solid
Hypothesis Translucent
Memory Warm glow
Fiction Distinct border
Simulation Grid overlay
Contradiction Broken line
Uncertain Hollow shape
Warning Red pulse, used sparingly
User-created Hand-drawn edge
AI-generated Machine glyph edge

The user must know whether an object is fact, memory, theory, fiction, or generated suggestion.

Deep without epistemology becomes persuasion machinery.

Deep with epistemology becomes cognitive instrumentation.


29. The Deep Archive Format

A future Deep Archive might contain:

deep_archive:
  metadata:
    title: "Field Training Scenario 01"
    version: "0.1"
    author: "example"
    safety_profile: "low_motion_symbolic"
    max_session_minutes: 5

  spaces:
    - id: field_001
      type: symbolic_environment
      background: low_detail_grass_field
      horizon: stable
      objects:
        - tree_001
        - path_001
        - gate_001

  objects:
    - id: tree_001
      category: tree
      marker_far: dot_green
      marker_mid: vertical_shape_green
      perceptual_keyframe: oak_tree_keyframe_01
      token: tree_icon
      detail_levels:
        - silhouette
        - trunk_canopy
        - bark_branch
      audio_label: "old oak"
      epistemic_status: observed
      confidence: 0.92
      relations:
        - beside: path_001

    - id: gate_001
      category: gate
      marker_far: dot_black
      marker_mid: rectangle_dark
      perceptual_keyframe: old_gate_keyframe_01
      token: gate_icon
      epistemic_status: uncertain
      confidence: 0.61
      relations:
        - blocks: path_001

  controls:
    gaze:
      fixation_select_ms: 650
      fixation_confirm_ms: 1200
      exit_fixation_ms: 2000
    blink:
      double_blink: reset
      long_blink_ms: 1500
      long_blink_action: pause

  return_protocol:
    always_available: true
    auto_return_on_no_input_seconds: 20
    auto_return_on_distress: true

This is not a media file.

It is a cognitive scene package.


30. What AI Does in Deep

AI is not the interface.

AI is the scene composer, archivist, narrator, simplifier, and safety assistant.

Given a document, project, memory set, training task, or problem graph, AI can:

extract objects
classify entities
generate symbols
generate keyframes
map relationships
assign confidence
build spaces
write audio guidance
adapt complexity
detect overload
summarise session output

This turns many forms of information into navigable visual-symbolic space.

AI becomes less “answer bot” and more cartographer of thought.


31. First Killer Application: Training

Deep can be used to build mental models faster than slides or flat video.

Example: aviation, medicine, maritime navigation, emergency response, engineering diagnostics.

Instead of reading:

Watch for crosswind drift.

The user sees:

runway marker
wind glyph
aircraft token
drift vector
correction path

The system briefly expands the scene at recognition points, then collapses back to symbols.

The user learns the relationship spatially.

Not memorisation.

Model-building.


32. Second Killer Application: Memory Palaces

Deep can turn memory palaces into programmable archives.

A user can load:

  • a legal case;
  • a codebase;
  • a speech;
  • a song structure;
  • a language vocabulary set;
  • an aircraft checklist;
  • a research map;
  • a personal project.

Deep converts it into spaces, objects, paths, and glyphs.

The user enters the archive, binds objects, follows relations, and exits with structured recall.

This is not note-taking.

It is spatialised memory construction.


33. Third Killer Application: Complex Problem Navigation

Many problems are not linear.

They are constellations:

actors
events
documents
contradictions
hypotheses
missing evidence
dependencies
risks
timelines

Flat screens struggle with this.

Deep can let the user inhabit the structure.

A fraud investigation becomes a room of objects and threads.

A codebase becomes a machine landscape.

A medical differential becomes competing paths and warning glyphs.

A strategy becomes weather, terrain, constraints, and moving pieces.

The user does not merely read the analysis.

The user enters the analysis.


34. Fourth Killer Application: Creative Work

Deep could be monstrous for creative work.

A song becomes a place.

A music video becomes a moving symbolic stage.

A story becomes a navigable myth-map.

A character becomes an object with memory, voice, costume, conflict, and transformations.

A chorus becomes a pulse in space.

The artist can move by gaze, bind scenes with keyframes, collapse them to icons, and re-expand details.

Not prompt engineering.

Scene gardening.


35. Why “Soft” Is the Real Product

The hardware is becoming available.

The difficult part is the software.

The soft must:

invite absorption
calibrate gaze
present sparse markers
time perceptual keyframes
avoid overload
keep audio coherent
interpret eye signals
control motion
log state
preserve exit
return reliably

The software does not “force” Deep.

It creates conditions where the user’s brain can enter a useful state voluntarily.

That is the difference between an instrument and a trap.


36. The Minimal Prototype: Deep-0

Deep-0 should not claim altered states.

Deep-0 should test the core interaction.

Deep-0 Goals

  • Can users bind a detailed object to a simple icon?
  • Can they navigate objects by gaze?
  • Can they recall more after a symbolic VR session than after a flat screen?
  • Can the system safely return the user every time?
  • Can the system detect confusion or overload?

Deep-0 Session

Duration: 3–5 minutes
Environment: simple field or room
Objects: 3–7
Motion: minimal
Controls: gaze + blink + physical stop
Audio: simple guidance
Exit: always visible

Deep-0 Tasks

find object
select object
remember object
compare object
track relationship
return and describe

Deep-0 success is not “the user went into a trance.”

Deep-0 success is:

The user can bind, navigate, remember, and return.


37. What Can Be Built Now

The near-term product should be an awake, ordinary prototype.

It should test the smallest useful loop:

marker → visible keyframe → icon binding → gaze/head command → return → recall check

This can be built now with:

  • a phone or commodity headset;
  • a simple symbolic scene;
  • three to seven objects;
  • visible perceptual keyframes;
  • coarse gaze or camera-based attention detection;
  • head-tilt symbolic movement;
  • audio guidance;
  • a fixed exit glyph;
  • a physical stop path;
  • session logs;
  • post-session recall questions.

It should not attempt:

  • sleep paralysis;
  • clinical claims;
  • subliminal frames;
  • high-risk psychological content;
  • long sessions;
  • photorealistic locomotion;
  • unlogged stimuli;
  • medical or therapeutic use.

The long-term Deep vision can remain ambitious.

The first build should be modest, measurable, and easy to exit.


38. The Research Questions

Deep becomes real only if it survives measurement.

Key questions:

  1. Do perceptual keyframes improve object recall?
  2. Does icon binding reduce cognitive load?
  3. Does gaze-only navigation preserve absorption better than hand controllers?
  4. Does symbolic rendering reduce VR sickness compared to photorealistic movement?
  5. Can users maintain dynamic object constellations more effectively in Deep than on flat screens?
  6. What are the safest and most reliable return cues?
  7. How does individual variation affect depth, comfort, recall, and control?
  8. Can overload be detected from gaze instability, blink rate, head movement, or task errors?
  9. What content categories are inappropriate for Deep?
  10. How should fact, fiction, memory, and hypothesis be visually distinguished?

39. The Design Principles

  1. The user must always be able to return.
  2. Do not force paralysis.
  3. Do not hide stimuli.
  4. Do not use subliminal manipulation.
  5. Use the brain as renderer, not victim.
  6. Use audio for guidance.
  7. Use eyes for control.
  8. Use symbols before spectacle.
  9. Use keyframes for binding, not persuasion.
  10. Represent uncertainty visually.
  11. Prefer stillness over motion.
  12. Prefer attention jumps over artificial locomotion.
  13. Make exit easier than entry.
  14. Log what was shown.
  15. Treat safety as architecture, not disclaimer.

40. The Product Stack

DeepOS
  ├─ Safety Shell
  ├─ User Calibration
  ├─ Gaze Runtime
  ├─ Blink Runtime
  ├─ Visual Archive Renderer
  ├─ Perceptual Keyframe Scheduler
  ├─ Symbolic Object Engine
  ├─ Relationship Graph Engine
  ├─ Audio Guidance Layer
  ├─ Epistemology Renderer
  ├─ Overload Detector
  ├─ Return Protocol
  ├─ Session Logger
  └─ Archive Exporter

The most important module is not the renderer.

It is the Return Protocol.


41. The Social Risk

If Deep works, it will be powerful.

That means it can be abused.

Possible risks:

  • persuasion through felt reality;
  • false memory reinforcement;
  • emotional manipulation;
  • over-trust in AI-generated scenes;
  • content that feels proven because it was spatially experienced;
  • loss of distinction between evidence and simulation;
  • surveillance through gaze and attention data.

Therefore Deep must be built with:

consent
audit logs
content provenance
epistemic visual grammar
local-first privacy where possible
clear session boundaries
user ownership of archives
export and delete rights

Attention data is intimate.

Gaze is not just cursor movement. It is behavioural telemetry.

Deep must treat it as sensitive.


42. The Deep Manifesto

We are tired of flat screens pretending to be enough.

We are tired of chatbots answering with confident fog.

We are tired of interfaces that bury thought in tabs, panels, documents, notifications, and tiny rectangles.

The human brain can do more.

It can hold landscapes.

It can track constellations.

It can bind symbols to memory.

It can hear a voice and build a world.

It can move through an archive without touching a keyboard.

Deep is an attempt to build for that brain.

Not to overwhelm it.

Not to trick it.

Not to trap it.

To give it a structured place where imagination becomes an interface.

The machine does not need to draw every leaf.

It only needs to show the right dot, at the right distance, at the right moment.

Then the tree appears where it always appeared:

inside the human mind.


43. Working Definition

Deep is a reversible gaze-controlled symbolic cognition interface that uses sparse visual markers, visible perceptual keyframes, and audio guidance to let the human brain reconstruct and navigate complex information as an internal spatial model.

Shorter:

Deep is an operating system for directed imagination.

Shortest:

VR supplies the glyphs.
Audio supplies the grammar.
The brain supplies the world.
The eyes supply the API.


44. Final Line

There is nothing impossible.

We are building Deep.

But we build it with one law carved into the foundation:

The way out is part of the system.


References and Grounding Notes

This document is speculative. The concept should be treated as research/design exploration, not validated neuroscience or medical guidance. The following references ground some safety and feasibility constraints:

  1. Cleveland Clinic, “Sleep Paralysis: What It Is, Causes, Symptoms & Treatment.” Notes that sleep paralysis can include inability to move or speak, chest pressure, and hallucinations.
    https://my.clevelandclinic.org/health/diseases/21974-sleep-paralysis

  2. W3C Web Accessibility Initiative, WCAG guidance on flashing content. Notes that flashing content can trigger seizures in susceptible people and recommends strict limits around flashing.
    https://www.w3.org/WAI/WCAG21/Understanding/three-flashes.html

  3. Frontiers in Virtual Reality, “Effects of Linear Visual-Vestibular Conflict on Presence, Perceived Self-Motion, and Cybersickness.” Discusses visual-vestibular mismatch as a contributor to cybersickness in VR.
    https://www.frontiersin.org/journals/virtual-reality/articles/10.3389/frvir.2021.582156/full

  4. Weech, Kenny, Barnett-Cowan, “Presence and Cybersickness in Virtual Reality Are Negatively Related: A Review.” Reviews presence and cybersickness in VR.
    https://pmc.ncbi.nlm.nih.gov/articles/PMC6369189/

  5. Konkoly et al., “Real-time dialogue between experimenters and dreamers.” Current Biology, 2021. Demonstrates limited two-way communication with lucid dreamers during verified REM sleep using signals such as eye movements and facial muscle contractions.
    https://pubmed.ncbi.nlm.nih.gov/33607035/

  6. Baird et al., “The cognitive neuroscience of lucid dreaming.” Reviews lucid dreaming and eye-movement signalling research.
    https://pmc.ncbi.nlm.nih.gov/articles/PMC6451677/

  7. Microsoft Support, “Eye control basics in Windows.” Useful mainstream example of dwell selection, eye-control launchpads, mouse control, scrolling, eye keyboard, and text-to-speech.
    https://support.microsoft.com/en-us/windows/eye-control-basics-in-windows

  8. Nature Communications, 2024, “Enhancing communication in ALS using an LLM-assisted eye-gaze typing system.” Shows how prediction can reduce the burden of eye-gaze typing and improve communication rates.
    https://www.nature.com/articles/s41467-024-53873-3

  9. Frontiers in Virtual Reality, 2023, “Evaluation of head joystick locomotion control in virtual reality.” Discusses head-orientation/head-lean style locomotion, calibration, and comfort constraints.
    https://www.frontiersin.org/journals/virtual-reality/articles/10.3389/frvir.2023.1169654/full

  10. Review literature on eye-gaze controlled assistive robotics and interfaces for severe motor impairment is directly relevant to Deep’s “eyes as API” control layer.
    https://pmc.ncbi.nlm.nih.gov/articles/PMC10909843/


Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment