Skip to content

Instantly share code, notes, and snippets.

@Co0olCat
Last active June 16, 2026 10:28
Show Gist options
  • Select an option

  • Save Co0olCat/8c5c78ee97117f4608bf57e1cbcb5b7a to your computer and use it in GitHub Desktop.

Select an option

Save Co0olCat/8c5c78ee97117f4608bf57e1cbcb5b7a to your computer and use it in GitHub Desktop.
Labyrinth of Reflections Model

Labyrinth of Reflections Model

Training intelligence on obraz instead of words

Author: Timur Yusupov, PhD

Status: Draft v0.5 for public discussion

Date: 2026-06-16

Short name: LRM

Workspace name: Obraz Lattice

Core claim: language is an output channel, not the native substrate of thought

Subtitle: A testable architecture hypothesis for training intelligence on structured multimodal idea-forms instead of language-first traces

Words are an export format. A mind is not built from its own press releases.

Before a person explains, they often see - not necessarily with the eyes, and not as pixels, but as an internal form: a relation, a pressure, a path, a shape, a tension. The idea moves before the sentence arrives. Then the fog settles: what was impossible to say becomes obvious enough to draw, point at, test, or name.

LRM starts from that moment. Not from the word. From the form before the word - from the obraz.


Abstract

Frontier AI systems are mostly trained on text: a serialized, compressed, lossy export of human thought. Text is extraordinarily valuable, but it is not the native form in which humans experience understanding. People often reason through spatial, embodied, relational idea-forms before they can put those ideas into language.

This note proposes the Labyrinth of Reflections Model (LRM): an AGI-oriented architecture hypothesis whose core training substrate is obraz, not words. An obraz is not a picture, bitmap, or frame. It is a structured, grounded, spatial-relational idea-form - proposed concretely as an object-centric, slot-and-relation structure - compiled from continuous multimodal experience. In LRM, perception writes obraz, prediction reshapes obraz, action grounds obraz, and language renders crystallised obraz into words when words are useful.

The practical thesis is deliberately narrower than the metaphor: at matched data and compute, a system trained to predict and transform structured grounded obraz should generalise better on compositional spatial, cross-modal, and action-conditioned tasks than a system trained only on serialized symbolic traces of the same experience. If that prediction fails, the idea fails.

This bet sits beside world models, joint-embedding predictive architectures, embodied AI, and global-workspace theory, and is sharpened against recursive language-model scaffolds. The photonic substrate that motivated the original idea is demoted here to an optional future implementation path; the load-bearing content is the representation, the training signal, the learning rule, and the explicit operability of the workspace.

Provenance: v0.5 follows adversarial review across earlier drafts and incorporates one further editorial pass. The train-on-obraz thesis is kept; the photonic substrate is demoted; the obraz representation is committed to an object-centric structure; the world-model differentiator is closed with a kill condition; the operability claim is made sharper; a short "why now" section is added; duplicated summaries and proverb-as-evidence have been removed.


Why now?

Language-first AI is becoming powerful enough to reveal its own missing organ.

Large context windows, multimodal adapters, tool use, agents, recursive scaffolds, and world models all point in the same direction: intelligence improves when the system is no longer trapped inside a single serial text buffer. Recursive Language Models show that externalising the object of work and operating over it structurally can beat stuffing everything into context. Vision-language systems show that perception helps, but often still route understanding through a language-centred core. World models show that predictive latent structure matters, but often keep that structure hidden inside dynamics rather than exposing it as an operable workspace.

LRM is proposed at that junction. It asks whether the next step is not a bigger window, a better prompt, or another adapter, but a native workspace where grounded idea-forms can be composed, fractured, simulated, inspected, and rendered.

The timing matters because the ingredients are finally testable on ordinary hardware: simulated environments, object-centric representation learning, graph-structured relational models, multimodal streams, action-conditioned prediction, and language renderers. The mirror-lattice metaphor can wait. The representation thesis can be tested now.


1. The claim in one sentence

LRM proposes that general intelligence should be trained on grounded, structured idea-forms - obraz - drawn from continuous multimodal experience, learned by self-supervised prediction in obraz-space, with language attached afterward as one renderer of crystallised structure.

Everything below either sharpens that sentence or tests it.


2. The intuition

Before a person can explain a hard idea, they often sketch it: circles, arrows, boxes, crossings, paths. The drawing is not decoration - it is thought becoming external, faster than sentences allow. And when a solution finally arrives, it tends to feel less like a sentence completing than like a structure settling: many possibilities shimmer, then one arrangement holds and the rest fall away.

LRM takes that pre-verbal stage seriously, against the language-first default:

  • Thought is not primarily a sequence of words.
  • Thought often has spatial, relational, multimodal form.
  • Language is one way to render thought, not the only medium of thought.
  • A general intelligence may need a native workspace where ideas form, move, collide, fracture, and settle before being spoken.

These are intuitions, not evidence - motivation for the hypothesis, not proof of it. The proof, if there is one, lives in §17. The workspace is the Obraz Lattice.


3. What is an obraz?

An obraz is not a picture. It is not a bitmap, frame, JPEG, screenshot, or ordinary visual embedding. It is a structured, composable, partly grounded representation of an idea: a form with internal geometry that the system can hold, transform, compare, merge, split, and break. It has at least six properties.

3.1 Spatial-relational structure. Parts and relations between them: containment, distance, direction, contact, opposition, sequence, causality, pressure, symmetry, transformation. A cup on a table is not the tokens [cup, on, table]; it is a relation field - support, gravity, contact, affordance, possible motion, possible breakage, possible grasp.

3.2 Continuous geometry. Obraz live in a similarity space; nearby forms mean nearby things. A chair, stool, bench, and a rock usable as a seat can be close in one region of the space even when they differ visually or linguistically.

3.3 Partial grounding. An obraz links back to experience - what was seen, heard, touched, predicted, attempted, corrected. The system knows the difference between a word heard once and a relation paid for by failed action.

3.4 Composition. Obraz nest and compose. A room contains a chair; a chair contains legs, seat, back, affordances, actions, memories, social meanings; a plan contains subplans; a proof contains transformations.

3.5 Transformability. The system can operate on obraz - rotate, occlude, unfold, compare, reverse, fracture, complete, analogise, simulate. This is the central difference from a passive vector.

3.6 Crystallisation. An obraz can move from vague to stable: a field of possible relations collapsing into a coherent form with low prediction error and high action-readiness.


4. Why obraz rather than image, latent, concept, or world model?

Each available English word pulls the idea in the wrong direction. Image suggests pixels or mental pictures. Concept suggests something too symbolic and language-like. Representation is too broad and bloodless. Latent suggests hidden activations inside a model, not a workspace the system can inspect and manipulate. World model is the closest, but usually names a learned predictive module rather than the primary medium of thought (the difference is made concrete in §8).

Obraz points to a form that is visual without being pictorial, symbolic without being textual, grounded without being raw sensory data, and dynamic without being merely a vector.

This matters for an obvious objection. People born blind reason, invent, navigate, plan, and abstract; people with aphantasia may report little or no voluntary mental imagery while performing complex thought. So LRM cannot depend on literal pictures in the head. The claim is not that intelligence requires internal cinema - it is that intelligence may require operable spatial-relational forms. Vision is the richest teacher of such structure for many agents, but hearing, touch, balance, proprioception, action, and language can all write obraz.


5. The thesis: train on obraz, not words

Today's most capable systems are largely trained to predict the next token over text. Text carries law, science, culture, jokes, source code, recipes, equations, grief. But it is an export format - what thought looks like after it has been serialized, narrowed, and written down for someone else. Training primarily on text means learning from the fossil record of cognition rather than from cognition's source.

LRM proposes the inverse pipeline: the primary training signal is grounded multimodal experience; experience is compiled into obraz; the system learns by predicting, correcting, and transforming obraz; language is attached as one renderer of crystallised obraz; and action closes the loop by testing whether the obraz actually predicts the world.

This is not a vision-language model. A VLM bolts visual input onto a language-centred system whose centre of gravity stays linguistic. In LRM the Obraz Lattice is primary and language is peripheral. The order is reversed on purpose.


6. Why language-first hits a wall

6.1 Text is serial. It unfolds one token at a time, while human thought often manipulates many relations at once - layout, salience, memory resonance, causal pressure, analogy, constraint. Serializing all of that loses structure.

6.2 Text is lossy. A paragraph about a workshop is not the workshop; a clause is not the dispute; a recipe is not cooking; a map is not terrain - and a sentence about terrain is less than the map.

6.3 Text is secondhand. Most text describes experience after the fact. LRM wants the system to learn from experience-like streams directly: seeing, hearing, touching, acting, predicting, failing, correcting.

6.4 Text hides the pre-verbal stage. The final explanation usually hides the actual process. The proof looks clean only after it settles; the words are a polished fossil of the event.

6.5 Bigger context windows do not solve this. A larger window gives the model more text. It does not change the fact that the system is trained to operate over serialized language rather than over a native structured world workspace.


7. Why Recursive Language Models are relevant but not sufficient

Recursive Language Models (Zhang, Kraska & Khattab, 2026) treat a long prompt as part of an external environment, hand the model a symbolic handle to it, and let it write code that inspects, decomposes, and recursively queries language models over slices of the prompt.

That is valuable, and it cuts both ways here. It shows that a key bottleneck is not simply context length but the monolithic way context is presented: externalising the object and operating over it structurally can beat stuffing everything into one prompt. So part of what looks like a "language" limit is really a buffer limit, and the fix can be structural while staying symbolic.

But RLM remains an inference-time scaffold around a model trained mostly on text. Its recursion lives in the runtime, not in the learned representation. LRM takes the deeper bet:

If RLM shows that externalised, operable structure helps at inference time, LRM asks whether the system should be trained on externalisable, grounded, operable structure from the beginning.

RLM asks how a language model can process a huge object. LRM asks why language is the core object at all.


8. Obraz versus pixels, tokens, latents - and world models

The idea is easy to misread as four familiar things.

  • Not pixels. Pixels are rendered surface. An obraz is the structure behind possible renderings - the operable form a cube keeps whether it is seen, touched, rotated, occluded, drawn, or imagined.
  • Not tokens. Tokens are discrete, serial symbols. Obraz are parallel, spatial, continuous, and transformable. Tokens can describe a maze; an obraz can hold it as something navigable.
  • Not ordinary latents. Every network has latent activations. An obraz must be more: a first-class object an internal controller can address, decompose, recombine, and transform - not hidden state that only the dynamics use.
  • Not (quite) a world-model latent - and here is the differentiator, made concrete. This is the objection that has shadowed every draft, so it gets an answer rather than another disclaimer.

A standard world model learns a latent that rolls forward under its dynamics. That latent is used by the model; it is not operated on by any process the system can direct. The obraz claim is that the workspace exposes its contents as addressable, recombinable parts - objects, relations, affordances - that an internal controller can bind into configurations on demand.

The testable consequence is compositional and transformational generalisation: constructing and reasoning about combinations never co-observed in training - these two objects, never seen together, interacting thus; this relation, learned in one modality, applied in another; this novel action sequence and its predicted result. A monolithic latent can interpolate within its learned dynamics. An operable workspace of recombinable parts can construct configurations outside them.

This is the trained, grounded analogue of the RLM lesson (§7): making the object operable by an explicit controller beat cramming it into one buffer. Obraz applies the same move to the world-representation itself.

It comes with a kill condition. If making the workspace operable yields no compositional advantage over an equivalent fixed world-model latent at matched data and compute, then obraz adds nothing to structured predictive latent, and the word should be dropped. The experiment in §17 is exactly this test.

The key word is operable. LRM is not merely "object slots exist somewhere inside the model." It is the claim that the active workspace exposes objects, relations, affordances, uncertainty, and counterfactual structure to an internal controller that can deliberately:

  • compose parts that were never co-observed;
  • fracture a false relation without destroying the whole structure;
  • simulate transformations before acting;
  • bind relations across modalities;
  • render the result as language, diagram, action, sound, or image.

Without those verbs, obraz collapses back into an ordinary latent with better mythology. With those verbs, it becomes a candidate medium of thought.


9. Minimal architecture sketch

LRM is a direction, not a finished architecture. What follows is a functional decomposition with one load-bearing commitment - the representation - made concrete, and the rest honestly marked open.

The load-bearing commitment: what an obraz is, as a structure. An obraz is proposed as an object-centric, slot-structured representation: a set of entity slots with continuous embeddings, plus explicit typed relations between them (containment, contact, support, direction, causal influence, and so on), carrying affordances and uncertainty. This is the operable, recombinable structure that §3 and §8 require, and it has a real lineage - object-centric and slot-based representation learning, and graph-structured relational inductive biases. The specific instantiation (how many slots, how relations are typed, how binding works) is open. That obraz are slot-and-relation structures rather than flat vectors or token sequences is the bet.

Around that commitment, the functional components:

  1. Multimodal transducers - vision, sound, touch, proprioception, balance, action feedback, and text map into a shared representational space; every modality contributes to obraz geometry rather than being translated into language. (Mechanism open.)
  2. Obraz encoder - binds raw streams into slots and relations, preserving structure, affordance, temporal change, and uncertainty rather than emitting flat labels. (The inductive bias - object-centric binding - is the committed part; the network is open.)
  3. Obraz Lattice workspace - holds active obraz and lets them interact: reinforcement, contradiction, analogy, occlusion, completion, fracture, crystallisation. The slot structure is what makes these operations addressable.
  4. Prediction engine - predicts missing, masked, future, and action-conditioned obraz (§10).
  5. Action loop - acts, predicts the consequence in obraz-space, observes, corrects, so the lattice pays rent to reality.
  6. Renderer layer - renders crystallised obraz as language, diagram, sound, image, or plan. Language is one renderer.
  7. Verifier and stabiliser - reality-checking, provenance, consolidation, contradiction handling, reversible exploration (§13, §19).
world stream
  -> multimodal transducers
  -> obraz encoder (object-centric: slots + typed relations)
  -> Obraz Lattice workspace
  -> prediction + action loop
  -> crystallisation
  -> renderer: language / action / diagram / sound / image

Honest status: components 1-7 are the functional skeleton of almost any embodied predictive agent. The content specific to LRM lives in three places - the representational commitment (slot-and-relation obraz), the training signal (grounded experience, not text), and the learning rule (§10). If those three are wrong, the boxes do not save it.


10. The learning rule

A workspace without an update rule is poetry, not architecture. The proposed rule is self-supervised prediction in obraz-space. The system continually predicts: missing obraz from partial obraz; future obraz from present obraz; hidden relations from visible relations; absent modalities from present modalities; consequences of its own actions; stable structure behind changing appearances. It adjusts to reduce prediction error over the obraz field.

This is related to predictive coding, world models, and joint-embedding predictive architectures. The difference is emphasis: the predicted object is not a token sequence or a raw-pixel reconstruction, but a structured, operable representation.

10.1 Crystallisation as a measurable event. In LRM, crystallisation is operational, not only metaphorical:

Crystallisation occurs when prediction error over an active obraz field collapses to a stable low value while action-readiness and cross-modal consistency rise.

A thought becomes clear when the lattice can predict, transform, and render it without internal contradiction.

10.2 Fracture as learning. False obraz must be broken, overfit structures weakened, misleading associations split apart. The system needs not just memory but the ability to revise memory without destroying everything attached to it. This suggests four movements: formation (weak obraz appears from experience), reinforcement (useful relations strengthened by prediction and action), fracture (misleading relations broken under error pressure), consolidation (stable structures compressed, linked, reused).

The specific loss is open. The commitment is only this: learning is prediction and correction over obraz, grounded by action.


11. Native multimodality and grounding

LRM should be multimodal from the first breath, not later through adapters. Vision is direct. Sound is vibration that can be shaped into spatial structure - rhythm, source, distance, motion, voice, intent. Touch is contact geometry and force. Balance is orientation and acceleration. Proprioception is body-state geometry. Pain is a high-priority signal about damage and boundary. Language is a modality too - but not the king modality. All of them enter the lattice as different ways of shaping the same world model.

Grounding comes from action. A system that only watches can learn much, but not enough; a representation that has never paid the cost of a wrong prediction is only half an idea. The system must learn that pushing moves, rotating hides one face and reveals another, dropping falls, speaking makes another agent react, and assuming a hidden relation is later confirmed or broken. That loop is where obraz stop being decoration and become cognition.


12. The substrate question: photonic, neuromorphic, electronic, or hybrid?

The original intuition behind LRM was photonic - light has high bandwidth, natural quantization, low propagation resistance, and an intuitive fit with reflection and refraction. A 3D lattice of adaptive reflective elements is a powerful image for thought. That image remains useful, but it should not be load-bearing yet.

12.1 Why photonics is tempting. High parallelism, high bandwidth, low-latency propagation, natural interference, potential energy advantages for some matrix-like operations, and a physical analogy to routing and transformation.

12.2 Why photonics is not enough. Passive optics is mostly linear: mirrors, lenses, path delays, occlusion, and interference do not by themselves create the nonlinear learning dynamics intelligence needs. Nonlinearity must come from nonlinear media, detectors, feedback, or optoelectronic control - and once a nonlinear element sits in a cavity with feedback, the natural result is an optical parametric oscillator network relaxing to a low-energy state: a coherent Ising machine, which is a known optimiser, not a general cognitive engine. The energy story is also not automatic; costs reappear in conversion, modulation, memory updates, detectors, stabilisation, and cooling.

Photonics may become an excellent substrate for parts of LRM, but LRM does not depend on photonics being solved first.

12.3 Prototype substrate. The first useful prototype runs on ordinary hardware: simulated environments, multimodal episode streams, structured (object-centric) latent workspaces, action-conditioned prediction, and comparison against serialized symbolic baselines. Only after the representation proves useful should the hardware question return.


13. Dreams, instability, and an unsolved tension

Humans do not only reason in controlled logic. We dream, misread, obsess, form impossible scenes - and sometimes return from those states carrying invention or insight. A system that never enters unstable recombination may be sterile; one that cannot return from it is dangerous. So LRM wants mode separation - perception, simulation, dream, planning, verification, action - and a return path from unstable states.

But here the proposal must be honest rather than confident, because the thing it wants is an open problem, not a design choice:

  • The more generative the dream-recombination - the property that makes it worth having - the less cleanly inspectable and reversible it tends to be. These pull against each other.
  • "The system must know when it is dreaming" presumes a reliable metacognitive monitor over a workspace that is continuously rewriting itself. That monitor is itself unsolved, and is plausibly harder here than in a static model with checkpointable state.
  • Listing requirements (provenance, consolidation, contradiction handling, reality checks) names what is needed; it does not show the set is jointly achievable.

The principle still holds as an aim:

The system may imagine, but it must know when it is imagining.

But it is a target to be earned, not a property the architecture confers for free. Reversible, inspectable, generative instability is one of the hardest things LRM asks for, and the draft should not pretend otherwise.


14. Relation to existing work

LRM is a synthesis and a bet, not an invention from nothing.

14.1 World models and JEPA are the closest technical relatives - they already learn predictive, non-pixel representations of environments. LRM's delta is operability and primacy: make the representation a first-class, structured, inspectable workspace rather than a module beneath a language head (the concrete differentiator and its kill condition are in §8). If that delta adds nothing, the framing should yield to ordinary world-model terminology.

14.2 Object-centric and graph-structured representation learning - slot-based and relational architectures - are the lineage for the specific commitment in §9: obraz as slots plus typed relations.

14.3 Vision-language models connect visual encoders to language models; LRM reverses the hierarchy, making language a renderer of a multimodal workspace.

14.4 Embodied AI aligns strongly, since grounding requires action; LRM adds the claim that action should train structured obraz, not merely produce more labels or reward.

14.5 Global workspace and blackboard architectures propose shared spaces where distributed processes read, write, and compete. LRM can be read as a global workspace whose contents are learned, grounded obraz.

14.6 Recursive Language Models are a structural contrast (§7): externalise and operate over structure, but at runtime, around a text-trained core.

14.7 Optical and neuromorphic computing are possible substrate ancestors - implementation options, not proof of the thesis.

14.8 Mental imagery and dual coding support the claim that cognition is not purely verbal; they motivate, they do not prove.


15. What LRM is not

Not a claim that current photonic chips are AGI; not a claim that a mirror maze becomes conscious; not a claim that language is useless; not a vision-only theory; not an image generator; not a dressed-up next-token model; not a guarantee of low energy; not a denial that blind or aphantasic humans reason; not a claim that dreaming should directly drive action; not a finished architecture. LRM is a testable bet about the native training substrate for general intelligence.


16. Strongest objections

Objection 1 - Is this just a world model with new vocabulary? The strongest objection, and the one §8 answers concretely: the differentiator is operability (addressable, recombinable parts) and its testable consequence is compositional generalisation, with an explicit kill condition. If the operable workspace shows no compositional advantage over a fixed latent, obraz should be dropped. The disagreement is not rhetorical; §17 measures it.

Objection 2 - Does "visual" hide the real claim? A real risk. If visual just means structured and spatial-relational, the visual framing can mislead. The honest version: vision is the richest teacher for many agents, but the required substrate is amodal obraz, not literal imagery (§4).

Objection 3 - Where is the exact loss function? Open. The commitment is to self-supervised, action-conditioned prediction in obraz-space (§10), not to one final objective.

Objection 4 - How do you stop continuous learning from becoming delusion? Not by forbidding instability but by making it inspectable, reversible, and separated from action - and §13 is honest that achieving this is itself unsolved, not a checklist.

Objection 5 - Why not just scale multimodal transformers? Maybe that works. LRM predicts scale alone is insufficient if the core remains serial and language-centred; §17 makes the disagreement measurable.

Objection 6 - Does the human analogy overreach? It can. Dreams, sketches, and imagery are suggestive, not proof. LRM is an engineering hypothesis inspired by human cognition, not neuroscience fact.

Objection 7 - Can obraz be inspected? It must be. If the workspace cannot be rendered into diagrams, causal traces, uncertainty maps, and counterfactuals, LRM becomes another opaque latent system with better mythology. The slot-and-relation commitment (§9) exists partly to make inspection tractable.


17. A discriminating experiment

The experiment separates the representation thesis from the hardware fantasy and runs on ordinary hardware.

17.1 Setup. Two systems are trained on the same stream of grounded multimodal episodes - object interactions with synchronized vision, audio, contact, position, and action metadata - at matched architecture scale, data, and compute.

  • System A (obraz workspace): compiles each episode into structured, object-centric obraz and trains by masked, future, and action-conditioned obraz prediction.
  • System B (serialized trace): compiles the same episode into a serialized symbolic or textual trace and trains by next-token prediction over that trace.

The only intended variable is representation format.

17.2 Test tasks (novel, not in training):

  1. recover an occluded contact event from partial audio-visual evidence;
  2. predict the result of a spatial transformation;
  3. infer a hidden relation between objects never co-observed;
  4. predict the consequence of a novel action sequence;
  5. solve a cross-modal completion (e.g. sound plus partial motion implying an unseen collision);
  6. transfer a relation learned visually into a tactile or action-only setting.

These are deliberately compositional - they probe exactly the operability claim of §8.

17.3 Prediction. If the thesis is right, System A generalises better than System B on spatial transformation, hidden relations, cross-modal completion, and action-conditioned prediction, and the gap widens as compositional complexity increases. If A and B perform equally at matched data and compute, the representation format was not load-bearing, and the LRM claim - including the obraz-over-world-model differentiator - is falsified.

17.4 Why this matters. No photonic brain, robot body, or metaphysics required. The question is bare: does training on structured grounded obraz produce capabilities that serialized symbolic training does not?


18. Prototype roadmap

Phase 1 - Software Obraz Lattice. A simulated environment with simple objects, surfaces, sounds, contacts, and actions; an object-centric structured workspace trained to predict hidden and future relations. Deliverable: obraz-format training beats serialized-trace training on held-out spatial and cross-modal tasks.

Phase 2 - Inspectable workspace. Tools to render internal obraz as diagrams, graphs, trajectories, uncertainty maps, and counterfactual scenes. Deliverable: humans can inspect what the system thinks is happening and why.

Phase 3 - Action grounding. Move from passive episodes to active behaviour; the system chooses interventions, predicts consequences, learns from failure. Deliverable: concepts robust through action, not only observation.

Phase 4 - Language renderer. Attach a language renderer after the workspace; the system explains crystallised structures rather than using language as the native medium. Deliverable: answers grounded in visible internal structure, not only verbal fluency.

Phase 5 - Hardware exploration. Only after Phases 1-4 show representational advantage should photonic, neuromorphic, or hybrid substrates be explored. Deliverable: whether the lattice benefits from non-traditional hardware without changing the representation thesis.


19. Safety and interpretability requirements

A grounded simulator may become more capable and more dangerous, so groundedness is not safety.

  1. State inspection - internal obraz must render into human-understandable forms.
  2. Provenance - learned structures link back to experience, evidence, and update history.
  3. Mode separation - perception, simulation, dream, planning, verification, action.
  4. Action gating - internal structures must not directly drive real-world action without verification.
  5. Reversible exploration - unstable or speculative states must be marked and returnable (and §13 notes this is unsolved, not assumed).
  6. Consolidation - sleep-like compression and cleanup to reduce noise and prevent runaway overwriting.
  7. Contradiction handling - conflicting obraz coexist under uncertainty until resolved, not silently averaged into nonsense.
  8. External reality checks - internal predictions tested against the world.

The system may imagine, but it must know when it is imagining - stated as a requirement to be earned, not a guarantee.


20. Falsifiable research questions

  1. At matched data and compute, does obraz-format training beat serialized-trace training on spatial-transformation tasks?
  2. Does shared obraz encoding recover absent modalities better than late-fusion or language-mediated alternatives?
  3. Can crystallisation be measured as a collapse in prediction error plus a rise in action-readiness?
  4. Does action-conditioned obraz prediction produce more robust concepts than passive ingestion?
  5. Can continual learning revise old obraz without catastrophic forgetting or uncontrolled drift?
  6. Does making the workspace first-class and operable add compositional capability beyond an equivalent hidden world-model latent? (The load-bearing test of the §8 differentiator.)
  7. Can internal obraz be rendered well enough for human debugging?
  8. Does a language renderer attached after the lattice produce more grounded explanations than a language-first system?
  9. Does the advantage widen with task complexity - transformations, occlusions, hidden relations, cross-modal inference?
  10. If photonic or neuromorphic substrates are later used, do they improve efficiency or capability without changing the representation thesis?

21. The short version

Words are an export format. A mind is not built from its own press releases.

LRM proposes that general intelligence should not be language-first. It should be trained on obraz - grounded, structured, spatial-relational idea-forms drawn from continuous multimodal experience, proposed concretely as object-centric slot-and-relation structures the system can address, recombine, and transform. Language is one renderer of crystallised structure, not the substrate of thought.

The serious version is narrow and falsifiable. At matched data and compute, a system that learns to predict and transform obraz should generalise better than a system trained on serialized symbolic traces of the same episodes - specifically on the compositional cases: spatial transformation, occlusion, hidden relations between objects never co-observed, cross-modal transfer, and novel action sequences. That same prediction is the kill condition for the world-model objection: if operability buys no compositional advantage, the framing collapses and the word should be dropped.

If the test fails, LRM is wrong, and we will know.

If it succeeds, the missing organ was never a larger context window, a recursive scaffold, or a better prompt. It was a place to see ideas before saying them.


References and adjacent work

These are adjacent rather than exhaustive - included to situate LRM and to make clear what it does not claim to have invented.

Representation, prediction, and world models

  • Yann LeCun, "A Path Towards Autonomous Machine Intelligence," 2022 - joint-embedding predictive architectures and non-generative predictive representations.
  • David Ha and Jürgen Schmidhuber, "World Models," 2018, arXiv:1803.10122.
  • Rajesh P. N. Rao and Dana H. Ballard, "Predictive coding in the visual cortex," Nature Neuroscience, 1999.

Object-centric and relational structure

  • Francesco Locatello et al., "Object-Centric Learning with Slot Attention," 2020.
  • Peter W. Battaglia et al., "Relational inductive biases, deep learning, and graph networks," 2018.

Cognitive architecture and consciousness

  • Bernard J. Baars, A Cognitive Theory of Consciousness, 1988 - global workspace theory.
  • Stanislas Dehaene, Consciousness and the Brain, 2014 - global neuronal workspace.

Language-model scaffolding and long context

  • Alex L. Zhang, Tim Kraska, Omar Khattab, "Recursive Language Models," preprint, arXiv:2512.24601, 2026.

Mental imagery and multimodal cognition

  • Stephen M. Kosslyn, "Mental Images and the Brain," Cognitive Neuropsychology, 2005.
  • Allan Paivio's dual-coding tradition; later surveys such as C. Kanellopoulou et al., "The Dual-Coding and Multimedia Learning Theories," Education Sciences, 2019.

Optical, photonic, and neuromorphic computing

  • Tingzhao Fu et al., "Optical neural networks: progress and challenges," Light: Science & Applications, 2024.
  • Xing Lin et al., "All-optical machine learning using diffractive deep neural networks," Science, 2018.
  • R. V. Kutluyarov et al., "Neuromorphic Photonics Circuits: Contemporary Review," Nanomaterials, 2023.
  • MIT News, "Photonic processor could enable ultrafast AI computations with extreme energy efficiency," 2024.

Biological energy context

  • BrainFacts / Society for Neuroscience, "How Much Energy Does the Brain Use?", 2019 - background for the commonly cited ~20 W figure, with the caveat that biological and engineered systems are not directly comparable.

Closing note

LRM began as a mirror maze: light, crystals, reflections, receptors, and the suspicion that intelligence might be built from forms rather than words. The mirror maze is still a useful image, but v0.5 does not ask the image to carry the science. The load-bearing claim is smaller and stronger:

Train the system on grounded, operable, object-centric obraz. Make language the renderer, not the substrate. Then test whether that changes generalisation - and drop the framing if it does not.

If the test fails, the labyrinth was only beautiful. If it succeeds, the first real clue was always in front of us: we did not need to say it first. We needed to see it.


Version archaeology

This gist is intended to keep the current version readable while preserving the path that led here. Earlier revisions are useful because they show which parts were raw intuition, which parts survived critique, and which parts were demoted.

Main public gist: https://gist.github.com/Co0olCat/8c5c78ee97117f4608bf57e1cbcb5b7a

@Co0olCat

Copy link
Copy Markdown
Author

Personal note: I use LLMs every day, but I won't pray to them. Magnificent tools — but a tool becoming dominant doesn't make it the foundation. Language is powerful, but it may still be an output channel: the smoke that rises after the fire has already burned somewhere deeper.
LRM is my attempt to look for the fire.

@Co0olCat

Copy link
Copy Markdown
Author

Practical note: experiments before temples
I use current LLMs, but I refuse to treat them as the only possible road to intelligence.
A fair objection is practical: even if an alternative architecture works in a small prototype, turning it into a useful model may require millions of dollars, serious compute, data, engineering, and a team. That is true. But it confuses two different stages.
The first stage is not to build a frontier model. The first stage is to ask whether there is anything worth scaling.
Bad ideas should die cheaply. Good ideas should first show a small, strange advantage: a task where the new representation does something the current paradigm does poorly. Only then does it make sense to look for scale, partners, money, labs, and real-world deployment.
For LRM, the first test may not be “train a smaller LLM on GPUs.” It may be to test a different physical and representational principle: continuous sensory fields, optical or analog dynamics, reactive media, and structured obraz before symbols.
The point is not that an optical output from a sound card will produce AGI tomorrow. The point is that the first useful experiment may be about the principle of representation, not the size of the model.
The question is not: can I build OpenAI in a garage?
The question is: can a garage experiment reveal a lever worth taking to OpenAI, Anthropic, or whoever is willing to scale what first looked impossible?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment