Combining everything, I would change one assumption from our earlier discussion:
Do not make every domain concept a generated Rust/Python/TypeScript struct.
If every new patent teaches us a new table concept and that requires:
LinkML → codegen → compile → deploy
then we've just moved our current bottleneck into a nicer schema system.
The architecture needs a stable kernel + rapidly evolving model.
There are three different rates of change:
SLOW FAST
──────────────────────────────────────────────────────────>
ENGINE / KERNEL WORLD MODEL KNOWLEDGE
months days/hours continuously
GraphNode TableFragment patent #123
GraphEdge IC50Measurement table #42
EntityId Assay "10 nM"
Provenance Compound relationships
TransformationRun rules
SchemaVersion ontology mappings
The mistake would be compiling the middle column completely into the left column.
I'd build Jubust around this:
JUBUST
┌─────────────────┐
│ WORLD MODEL │
│ │
│ LinkML schemas │
│ relationships │
│ transformations │
│ SHACL rules │
│ descriptions │
│ versions │
└────────┬────────┘
│
compile / validate
│
▼
┌────────────┐
│ Graph IR │
└─────┬──────┘
│
┌─────────────┼──────────────┐
▼ ▼ ▼
Visualizer Runtime Agents / CI
│ │
│ ▼
│ Patent instances
│ │
│ provenance/evidence
│ │
└─────────────┼──────────────┐
▼ ▼
Graph PDF
The World Model should be data, not application code.
That is the crucial property that allows it to evolve rapidly.
Rust/Python/TypeScript only need to understand a handful of concepts permanently:
Entity
Relationship
TypeDefinition
PropertyDefinition
TransformationDefinition
TransformationRun
Evidence
SchemaVersion
Something conceptually like:
{
"id": "table_fragment:123",
"type": "jubust:TableFragment",
"schema_version": "2026-09-18.4",
"properties": {},
"relationships": []
}Your Rust backend doesn't necessarily need:
struct IC50Measurement {...}
struct KiMeasurement {...}
struct FooBarNewPatentThing {...}to understand every new concept.
It understands:
Entity
type → IC50Measurement
properties → according to schema
The World Model tells the engine what IC50Measurement means.
For important stable concepts, we can still generate native types.
That's a developer-experience/performance optimization, not the foundation.
Have something roughly like:
world-model/
model/
document.yaml
tables.yaml
chemistry.yaml
bioactivity.yaml
transformations/
tables.yaml
normalization.yaml
constraints/
tables.shacl.ttl
bioactivity.shacl.ttl
generated/
graph-ir.json
json-schema/
python/
rust/
typescript/
examples/
patents/
Treat generated/ as build artifacts.
The important directory is:
world-model/model/
This becomes something developers and eventually agents can change cheaply.
This is a major architectural requirement.
We want:
Edit LinkML
↓
validate schema
↓
compile World Model
↓
run SHACL
↓
generate Graph IR
↓
VISUAL DIFF
↓
publish model version
↓
runtime loads new model
Not:
edit schema
↓
regenerate Rust
↓
compile backend
↓
deploy backend
↓
deploy frontend
↓
hope
The runtime should load a versioned compiled World Model artifact.
For example:
world-model-2026-09-18.4.json
Backend and frontend understand the stable Graph IR/kernel.
They don't need a redeployment because we introduced:
CellViabilityMeasurement
This is where the visualization becomes much more than documentation.
Suppose tomorrow a patent teaches us:
TableFragment
│
▼
PercentInhibitionMeasurement
│
├── concentration
└── percentage
We modify LinkML.
Immediately the system generates:
BEFORE AFTER
Measurement Measurement
│
├── IC50
│
└── + PercentInhibition
│
├── concentration
└── percentage
A PR should eventually contain an automatically generated World Model Diff:
git diff
+
schema diff
+
graph diff
+
SHACL violations
+
affected transformations
+
affected patents
That's substantially more useful for Jubust than looking at code diffs alone.
This matters enormously for patents.
Suppose we encounter:
"Residual Activity [%]"
Don't immediately add:
ResidualActivityMeasurement
to the ontology.
First capture an observation:
UnknownField
raw_label:
"Residual Activity [%]"
source:
patent X
table:
42
bbox:
...
Now the system can show:
KNOWN WORLD MODEL
IC50
Ki
Kd
EC50
UNMAPPED OBSERVATIONS
"Residual Activity [%]" 17 occurrences
"Remaining activity" 8 occurrences
"% inhibition @ 10 μM" 31 occurrences
Then a human or agent proposes:
PercentActivityMeasurement
and maps those observations onto it.
This prevents the ontology from becoming garbage simply because every patent uses different terminology.
This is one of the most important design decisions.
Patents continuously teach the World Model:
CURRENT WORLD MODEL
│
▼
ingest PDF
│
▼
extraction
│
▼
map against World Model
/ \
/ \
known unknown
│ │
▼ ▼
normal flow Observation
│
▼
Discovery Queue
│
human / agent review
│
▼
Model proposal
│
▼
Graph Diff
│
▼
validate
│
▼
World Model vNext
│
└──────► loop
This architecture fits Jubust much better than a conventional fixed ontology.
The ontology/world model isn't designed once.
It grows from evidence.
LinkML answers:
What concepts and relationships exist?
SHACL answers:
Which configurations are invalid?
Example:
LinkML
LogicalTable
fragments → TableFragment[*]
Later:
SHACL
LogicalTable must have fragments
all fragments must belong to same patent
every fragment must have evidence
every evidence node must eventually resolve to PDF + bbox
These rules can evolve without recompiling Rust.
A model release could therefore look like:
World Model v17
├── schema
├── relationships
├── transformations
└── constraints
The World Model becomes a deployable data artifact.
Don't encode transformation implementations into LinkML.
Define the contract:
merge_multi_page_tables
consumes:
TableFragment[*]
produces:
LogicalTable
implementation:
pipeline.tables.merge_multi_page_tables
version:
...
Runtime resolves that to Python.
The graph naturally becomes:
TableFragment
│
▼
┌─────────────────────────┐
│ merge_multi_page_tables │
└─────────────────────────┘
│
▼
LogicalTable
Later we can add:
implemented_by
│
▼
Python CodeSymbol
│
▼
git SHA
We don't need to solve that in prototype #1.
Don't build a graph visualization that's merely pretty.
Eventually it should become the World Model IDE.
Initially:
[Document] [Tables] [Chemistry] [Bioactivity]
Patent
│
Page
│
TableFragment
│
┌──────▼──────┐
│ table merge │
└──────┬──────┘
│
LogicalTable
Click LogicalTable:
LogicalTable
Properties
────────────────
id
fragments
headers
...
Relations
────────────────
composed_of → TableFragment
Produced by
────────────────
merge_multi_page_tables
Constraints
────────────────
3 SHACL rules
Instances
────────────────
12,481
Unmapped nearby observations
─────────────────────────────
17
Eventually the visualization is how we understand and modify Jubust's conceptual world.
Do not start with compounds, SHACL, agents, or runtime schema loading.
The first 2–3 day prototype should be:
LinkML
Patent
│
Page
│
BoundingBox
│
TableFragment
│
▼
merge_multi_page_tables
│
▼
LogicalTable
↓ compile
Graph IR
↓
Interactive visual map
Only prove five things:
- Existing Jubust types can reasonably map onto a tiny LinkML model.
- Relationships are defined once in LinkML.
- Transformation contracts can appear in the same conceptual graph.
- Changing LinkML immediately regenerates the graph and visualization.
- The frontend knows only the stable Graph IR — adding a new type requires zero frontend code.
Number 5 is critical.
Test it explicitly:
Add to LinkML:
FooMeasurement
├── value
└── unit
Run:
jubust world-model buildRefresh the visualization.
If FooMeasurement appears automatically with its properties and relationships without touching Rust or TypeScript, we've proven the most important architectural property.
The next prototype should NOT be more schema work.
It should connect reality:
real patent
↓
instances
↓
same Graph IR
↓
same visualization
Now we can switch between:
SCHEMA VIEW
TableFragment
│
▼
merge_multi_page_tables
│
▼
LogicalTable
and:
INSTANCE VIEW
TableFragment #123
│
▼
TransformationRun #8372
│
▼
LogicalTable #55
Then connect:
LogicalTable #55
↓
TableFragment #123
↓
BoundingBox
↓
Page 37
↓
PDF
At that point the map stops being architecture documentation and starts becoming an actual living Jubust World Model.
The pieces eventually become:
JUBUST WORLD MODEL
LinkML
schema / meaning
│
┌─────────────────┼─────────────────┐
│ │ │
Document Model Scientific Model Transformations
│ │ │
└─────────────────┼─────────────────┘
│
Graph IR
│
┌────────────────┼────────────────┐
│ │ │
Visualization Runtime Agents
│ │ │
│ Instance Graph │
│ │ │
│ Provenance │
│ │ │
└────────────────┼────────────────┘
│
Evidence
│
PDF
+
│
SHACL
rules/invariants
│
▼
Validation / CI
The fundamental principles are:
LinkML describes the language of the Jubust world.
The Graph represents that world.
Graph IR gives software a stable interface to that changing world.
Provenance explains how knowledge was created.
Evidence connects knowledge back to patents.
SHACL defines rules the world must obey.
The visualization becomes our external brain for understanding it.
Agents can eventually propose changes to both code and the World Model.
And most importantly:
The World Model must be able to evolve faster than the application code.