Neonmem Workspace
A vibecode-centric IDE built on inbuilt models.
Generation runs on models shipped inside the editor. A memory slot accepts a Neonmem cartridge, which supplies the agent with the accumulated experience of the project at retrieval time.
Memory takes two forms. A cartridge is a single file holding a typed graph of nodes, edges and quantized vectors, consulted per turn. A binary slot holds the same domain material compiled into model KV-state and restored at load, so it applies without occupying the context window.

The Neonmem cartridge
A single-file, multi-layer memory store with built-in vector storage.
A cartridge holds a project's experience as a typed graph. It is little-endian, page-aligned, and begins with a 64-byte header; the sections below carry the memory itself.
Nodes
Fixed-size records. Each holds a type — observation, question, hypothesis, conclusion, resolution, rule, dead-end, plan, procedure — with a confidence and relevance score, timestamps, an edge count, and an index into the vector store.
Edges
Typed, directed relationships between nodes: led-to, blocked-by, supports, contradicts, part-of. A node stores its first edges inline, with an overflow table for high-degree nodes.
Vectors
One embedding per node, int8-quantized with a per-vector scale (v = q·scale). The vectors are L2-normalized, so cosine similarity is a dot product. Recall is exact, over the vectors in the cartridge.
Integrity
Node text is stored in an LZ4-compressed string table; a footer carries a per-section CRC32 and a content hash. The vectors are part of the file — there is no external database.
Layers
- reflexRules and invariants that always apply.
- short-termThe decisions and observations of the current effort.
- long-termKnowledge retained across sessions.
PRISM — reasoning through the graph
Recall answers what is stored; PRISM reasons over how the stored items relate. It is based on Hopfield/Ising energy-based associative memory. For a query, anchor nodes are found by cosine similarity and their neighborhood is expanded into a subgraph. The subgraph is treated as a network of units: its edges define a coupling matrix J — supporting edges attract, contradicting and blocking edges repel — and node types define a field bias h. The unit states relax by
until they settle into a low-energy configuration. That configuration is read out as attractors — the answer — repelled regions — dead-ends to avoid — and active constraints — the rules that apply.
Embeddings
Text is embedded with the IBM Granite-30M model. A string is tokenized, passed through one forward pass, and the last hidden state at the CLS position is L2-normalized to a 384-dimensional unit vector. The same model runs at write and read time.
Capture
What a turn of work adds is scored by an energy-like salience — σ(β·L)·L — combining novelty, node-type weight, and the relief of resolving an open thread. Items above a write threshold are kept; a plan is a forward-looking commitment and is kept even when it relates to existing memory.
What the agent does
Generation, retrieval, analysis and planning, on local models.
All four run against the models shipped inside the editor. Nothing is sent anywhere.
Code generation
A turn produces files in the project, not a suggestion to copy. Whole-file writes carry the current contents; adding a method to an existing class generates the method alone and inserts it. Markup, configuration and data formats are written the same way as source.
Retrieval-augmented grounding
Each turn embeds the request and recalls from the project's cartridge before the model runs. What is retrieved is reduced to a contract — declarations, annotations and query-scored lines under a token budget — rather than whole files, so it informs the answer without displacing it.
Analysis
A question about existing code is answered from the code; a question about the project's own history is answered from its memory, with no file access. Recorded dead-ends are surfaced as warnings when a question touches one.
Planning
Forward-looking commitments are recorded as plans and kept even when they restate known facts, since a plan about a project naturally resembles it. Standing rules are held in the reflex layer and applied every turn.
What the agent checks before it finishes
Reading generated code establishes less than running it. A file can import cleanly, satisfy every pattern a reviewer would look for, and still not work: a function called with arguments its own declaration cannot accept raises at the first call, and a traversal that enqueues an item it never marks as seen never terminates at all. Both were observed and are checked for.
- runs itAfter a turn, the work is executed where it declares an entry point. A failure — including a program that does not terminate — is returned to the model, and a repair is kept only if the program then runs.
- arityA call whose arguments its own declaration cannot accept, in the same file.
- dependenciesAn import the project's declared Maven or Gradle dependencies cannot provide.
- deliveryA task naming several files that produced only some of them.
A session
Designing a small game, then asking the project what it decided.
What follows is one continuous session on a fresh install, from first launch to the end of the day. The prompts are typed as written; the replies are the replies that came back. Nothing here was staged, and the memory it describes is the one this session produced.
First launch
- Install
The editor opens with no models on disk. A notification offers them; declining leaves everything except the offline agent working, and the offer returns next launch.
- Accept
Roughly five and a half gigabytes are fetched into the install directory, resuming where an interrupted transfer left off. On the machine this was written on it took a little under five minutes.
Framing the project
- You type
We're building a small terminal game in Python called The Labyrinth — single file, no dependencies. The player walks through connected rooms looking for the exit.
- You type
Open question: should rooms be plain dictionaries or proper classes?
- Recorded
The project frame and the open question are written into the project's memory as they are said, typed as what they are — an open question stays open until something settles it.
A decision, and the reason for it
- You type
Let's debate the map itself: a hand-crafted fixed maze versus a procedurally generated one.
- You type
Decision: go with a hand-crafted maze for v1 — it's testable and we can tune the feel. Procedural can come later.
- Recorded
The decision is linked to the question it answers, so the reasoning travels with it rather than being stored as a bare verdict.
A failure, and the rule that came out of it
- You type
Heads up: I earlier tried keeping the player position and room state in module-level globals and it caused nasty state bugs across functions. That approach is a dead end.
- You type
So let's make it a rule: always keep the whole game state in one dict passed around explicitly — never module globals.
- What changed
The code written after this point threads the state through explicitly. Where a move had taken a direction and reached for whatever was in scope, it now takes the game state as an argument and reads the rooms out of it. The rule was not repeated in any later prompt.
Later, asking the project about itself
- You ask
Why did we go with a hand-crafted maze again?
- It answers
The decision to go with a hand-crafted maze was based on the observation that the hand-crafted fixed maze was testable and we could tune the feel.
- You ask
Remind me — is it okay to keep game state in globals?
- It answers
No, it is not okay to keep game state in globals — and it quotes the reason that was recorded, that module-level globals for the player position and room state caused state bugs across functions.
- You ask
What did we decide today?
- It answers
We decided to go with a hand-crafted maze for v1 — it's testable and we can tune the feel.
- You ask
What's our rule about dependencies?
- It answers
The rule about dependencies is that this project should always have zero dependencies.

None of those four answers is in any file. They are in the project's memory because the session put them there, and the questions are routed to it rather than to the code: a question about the project's own record never opens a file. The recorded dead-end is the one that matters most — it comes back as a refusal with its reason attached, which is the difference between a memory and a log.
Download
Neonmem Workspace for Windows.
Two ways to run it, both carrying the editor, the memory engine, and the embedding model. The installer sets it up on your PC; the portable build unzips and runs from a folder — nothing added to your system, keep it on a drive and take it with you.