How memspine works

memspine keeps a log of events for each of your customers, with a link from every derived fact to the events it came from. This page covers what each operation guarantees, how recall ranks, how customers are kept apart, what a crash or a delete leaves on disk, and where another tool fits better. The transcripts were recorded from version 0.2.2.

Events, and links to their sources

The application behind an agent appends an event for each thing worth keeping: a message from the customer, a tool call with its result, a fact the agent derived. This is the only record type. An event is stored as sent and never changed afterwards. A derived event names the events it came from, and the server stores a link to each.

timekindidtext
09:14:02turn75173310Ana: we adopted a greyhound called Pixel last weekend
09:14:31turn9509bd23Ana: she needs a vet, somewhere open on Saturdays
09:14:33tool_call697ea0fafind_vets(city=Lisbon, open=saturday) -> Clinica Arroios, VetLuz
09:15:10turnb8533f46Ana: book Clinica Arroios for this Saturday at ten
09:15:12tool_call740c7d03book_vet(Clinica Arroios, 2026-10-10 10:00) -> confirmed A-2291
09:15:13factd53b6b8bAna owns a greyhound named Pixelfrom 75173310
09:15:14factb41bdb35Ana prefers Saturday appointmentsfrom 9509bd23, b8533f46
09:15:15fact3dce87e5Pixel has a vet appointment at Clinica Arroios on 2026-10-10 at 10:00, booking A-2291from b8533f46, 740c7d03
09:41:50turnaf2d5cc8Ana: my number is 912 555 019, text me the day before
09:41:52factc3d7a7afText Ana the day before the vet appointmentfrom af2d5cc8, 3dce87e5

Ten events from one recorded session, as GET /v1/memory/events returned them, shown here oldest first. The route returns newest first. The lines at the left are the SOURCE_FROM links. Embeddings have four dimensions and were chosen by hand so that the rankings further down can be followed. The two tool calls were stored without one. Every command and its output is in session.txt.

A store that kept only the derived facts, or only their vectors, would be smaller, but a conclusion could not be traced and a source could not be deleted on its own. Here the appointment in row 8 leads to the instruction in row 4 and the booking call in row 5, and removing row 9 takes the phone number off the disk and leaves the reminder in row 10.

kind
A label the application chooses: any string of 1 to 256 bytes. The server filters on it and gives it no meaning. This session uses turn, tool_call and fact.
embedding
Optional. A vector from the application's own model, of exactly the server's dimension (384 unless configured), finite and with non-zero length.
properties
A JSON object, returned as sent. Its source_event_ids lists the events this one was derived from. For each id found in the same customer's data the server stores a SOURCE_FROM link from the derived event to that source. An id it cannot find is skipped and the append still succeeds.

Three of an event's five fields. The other two are text, all of which is indexed for keyword recall, and ts, the time it happened in microseconds, which defaults to the arrival time and is not checked against the server's clock. Text and properties are bounded by the request body: 2 MiB for one event, 32 MiB for a batch of up to 1,000.

Extracting facts and computing embeddings are the application's work. The server never calls a model. Events are immutable, so a correction is a new event followed by a delete of the old one.

Text and properties of any size within the body limit read back exactly as sent. From Cypher, only 16 top-level properties per event can be filtered on by name: booleans, numbers and strings of at most 256 bytes, under keys that are plain identifiers, taken in key order. The cap comes from FFS, the storage engine, which keeps an in-memory index entry for every pair of scalar properties on a node, so an event's cost grows with the square of that count.

The operations, and what is true when each returns

Every request to the memory routes carries a bearer key. The key determines the customer, and no field in a request can name another one. The interface is HTTP with JSON bodies, and the transcripts on this page call it with curl. Rust, Python and TypeScript client libraries exist, and the installer does not include them. An agent uses four operations: append, recall, recall with hops, and forget.

Append

POST /v1/memory/eventPOST /v1/memory/eventswrite scope
Sent
One event, or up to 1,000 in one request. An Idempotency-Key header makes a retry safe.
Returned
The id of each event.
True when it returns
The event, its embedding, its text and its links are in the customer's log, and the log has been fsynced. A batch is one commit and one fsync: all of it is stored or none of it.
Not promised
A second copy on another disk.
$ curl -s localhost:7777/v1/memory/events -H "$K" -H "$J" -d @turns.json; echo
{"event_ids":["75173310-8a48-40bc-8b49-8c2ab8c7dc0c","9509bd23-a670-42e0-bbe0-b87e994b0617","697ea0fa-fa50-4d51-9aa8-417263290f08","b8533f46-b491-4c20-a078-b0ca2e039926","740c7d03-9b97-4854-8799-6d66524cf3fe"]}

$ curl -s localhost:7777/v1/memory/event -H "$K" -H "$J" -d @fact.json; echo
{"event_id":"c3d7a7af-d5dc-4a60-bc3c-2fe5d52e992d"}

Recall

POST /v1/memory/recallread scope
Sent
text, an embedding, or both. k up to 1,000, and 10 if omitted. Optional kind, since and until filters.
Returned
Up to k events in rank order, each with its text, its properties and a score.
True when it returns
The same request over the same data gives the same list in the same order. An embedding is compared with every embedding the customer has stored, so the nearest ones are exact.
Not promised
Sub-linear time, or keyword matches in text outside ASCII letters and digits.
$ curl -s localhost:7777/v1/memory/recall -H "$K" -H "$J" \
  -d '{"text":"vet appointment booking","k":2}' \
  | jq -r '.hits[] | [(.score*1e4|round/1e4), .kind, .event_id[0:8], .text[0:44]] | @tsv'
3.2165	fact	3dce87e5	Pixel has a vet appointment at Clinica Arroi
2.5492	fact	c3d7a7af	Text Ana the day before the vet appointment

Recall with hops

POST /v1/memory/recallread scope
Sent
The same request with hops set to 1, 2 or 3.
Returned
The ranked hits, then up to k more events reached over links from them. Each added event carries via, the id of the hit it was reached from.
True when it returns
Every added event is linked to a ranked hit within hops steps. The walk loads at most 8 × k events, each once, so a source shared by many hits is not walked again for each of them.
Not promised
Every linked event, since additions stop at k. A cap on the links read from one event: the walk reads the whole list of links of each event it passes through.
$ curl -s localhost:7777/v1/memory/recall -H "$K" -H "$J" \
  -d '{"text":"vet appointment booking","k":2,"hops":1}' \
  | jq -r '.hits[] | [.kind, .event_id[0:8], (.via // "-")[0:8], .text[0:52]] | @tsv'
fact	3dce87e5	-	Pixel has a vet appointment at Clinica Arroios on 20
fact	c3d7a7af	-	Text Ana the day before the vet appointment
turn	b8533f46	3dce87e5	Ana: book Clinica Arroios for this Saturday at ten
tool_call	740c7d03	3dce87e5	book_vet(Clinica Arroios, 2026-10-10 10:00) -> confi
timekindidtextrecall
09:14:02turn75173310Ana: we adopted a greyhound called Pixel last weekend
09:14:31turn9509bd23Ana: she needs a vet, somewhere open on Saturdays
09:14:33tool_call697ea0fafind_vets(city=Lisbon, open=saturday) -> Clinica Arroios, VetLuz
09:15:10turnb8533f46Ana: book Clinica Arroios for this Saturday at tenits source
09:15:12tool_call740c7d03book_vet(Clinica Arroios, 2026-10-10 10:00) -> confirmed A-2291its source
09:15:13factd53b6b8bAna owns a greyhound named Pixelfrom 75173310
09:15:14factb41bdb35Ana prefers Saturday appointmentsfrom 9509bd23, b8533f46
09:15:15fact3dce87e5Pixel has a vet appointment at Clinica Arroios on 2026-10-10 at 10:00, booking A-2291from b8533f46, 740c7d03answer
09:41:50turnaf2d5cc8Ana: my number is 912 555 019, text me the day before
09:41:52factc3d7a7afText Ana the day before the vet appointmentfrom af2d5cc8, 3dce87e5answer

The response above, marked on the log. The third column of the transcript is via. The booking call in row 5 has no embedding and was outside the two ranked hits. It is returned because the fact in row 8 names it as a source. Row 9, a source of the second hit, is one step away too and was left out: k is 2 and both additions were taken.

Forget

DELETE /v1/memory/event/<id>write scope
Sent
An event id. Add ?purge=true to rewrite the customer's files before the response.
Returned
deleted, and purged: whether the event's bytes have left those files.
True when it returns
Get, list, recall and Cypher no longer return the event, its links are gone, and that is durable. When purged is true, no file in the customer's directory holds its text or embedding.
Not promised
Without ?purge=true the bytes stay in the log until the next purge pass, which by default runs every 60 seconds. What a delete removes, and when.
$ curl -s -X DELETE "localhost:7777/v1/memory/event/$E9?purge=true" -H "$K"; echo
{"deleted":true,"purged":true}

Reading without ranking

GET /v1/memory/events
A customer's events newest first, ordered by time and then by id, in pages of up to 1,000 with a cursor. Optional kind, since and until.
GET /v1/memory/event/<id>
One event as stored. An id that belongs to another customer, or to a forgotten event, gets a 404.
POST /v1/memory/cypher
A read-only query over (:Event) nodes and [:SOURCE_FROM] links, with parameters. CREATE, MERGE, SET and DELETE get a 400, and a result of more than 10,000 rows is refused. Query shapes known to exhaust the server are refused before they run. The cost of a query that is accepted has no bound.
Recorded: every fact with its sources in one Cypher query, and a query that is refused
cypher
$ curl -s localhost:7777/v1/memory/cypher -H "$K" -H "$J" -d '{"query":
  "MATCH (f:Event)-[:SOURCE_FROM]->(s:Event) WHERE f.kind = $kind RETURN f.text, s.kind, s.text",
  "params":{"kind":"fact"}}' | jq -c '.rows[]'
["Ana owns a greyhound named Pixel","turn","Ana: we adopted a greyhound called Pixel last weekend"]
["Ana prefers Saturday appointments","turn","Ana: she needs a vet, somewhere open on Saturdays"]
["Ana prefers Saturday appointments","turn","Ana: book Clinica Arroios for this Saturday at ten"]
["Pixel has a vet appointment at Clinica Arroios on 2026-10-10 at 10:00, booking A-2291","tool_call","book_vet(Clinica Arroios, 2026-10-10 10:00) -> confirmed A-2291"]
["Pixel has a vet appointment at Clinica Arroios on 2026-10-10 at 10:00, booking A-2291","turn","Ana: book Clinica Arroios for this Saturday at ten"]
["Text Ana the day before the vet appointment","fact","Pixel has a vet appointment at Clinica Arroios on 2026-10-10 at 10:00, booking A-2291"]
["Text Ana the day before the vet appointment","turn","Ana: my number is 912 555 019, text me the day before"]
$ curl -s localhost:7777/v1/memory/cypher -H "$K" -H "$J" \
  -d '{"query":"MATCH (a:Event), (b:Event) RETURN count(*)"}'; echo
{"error":"invalid_request","detail":"the query matches two patterns that share no variable, which pairs every row of one with every row of the other. Connect them through a shared variable"}

The reference lists every route, including the embedding-only search and the admin routes for customers and keys.

How recall ranks

the request carrieshow the customer's events are ranked
textBM25 over each event's text, with k1 of 1.2 and b of 0.75. The tokenizer lower-cases ASCII letters and splits on every character that is not an ASCII letter or digit. It has no stemming and no stop-word list. appointments does not match appointment, the counts as a term, accented words are split at the accent, and text in other scripts produces no terms.
embeddingCosine similarity between the query vector and every embedding the customer has stored. There is no approximate index.
bothThe two rankings are fused by reciprocal rank, with ranks counted from 1.
score = 1 / (60 + rankembedding) + 0.5 / (60 + rankkeyword)

One question, ranked three ways

The session asks when is the vet appointment for Pixel by embedding, by text, and with both. The three scores are on different scales: cosine similarity, BM25, and the fused score.

both, fused
$ curl -s localhost:7777/v1/memory/recall -H "$K" -H "$J" \
  -d '{"text":"when is the vet appointment for Pixel",
       "embedding":[0.3,0.9,0.3,0.0],"k":10}' \
  | jq -r '.hits[] | [(.score*1e4|round/1e4), .kind, .event_id[0:8], .text[0:44]] | @tsv'
0.0243	fact	3dce87e5	Pixel has a vet appointment at Clinica Arroi
0.0236	fact	c3d7a7af	Text Ana the day before the vet appointment
0.0228	turn	af2d5cc8	Ana: my number is 912 555 019, text me the d
0.0161	turn	9509bd23	Ana: she needs a vet, somewhere open on Satu
0.0159	turn	b8533f46	Ana: book Clinica Arroios for this Saturday 
0.0156	fact	b41bdb35	Ana prefers Saturday appointments
0.0152	turn	75173310	Ana: we adopted a greyhound called Pixel las
0.0149	fact	d53b6b8b	Ana owns a greyhound named Pixel

The first row is first by embedding and third by keyword: 1 / (60 + 1) + 0.5 / (60 + 3). The fourth is second by embedding and under the keyword floor, so it scores 1 / (60 + 2) alone.

The two rankings that were fused, as the server returned them
by embedding
$ curl -s localhost:7777/v1/memory/recall -H "$K" -H "$J" \
  -d '{"embedding":[0.3,0.9,0.3,0.0],"k":10}' \
  | jq -r '.hits[] | [(.score*1e4|round/1e4), .kind, .event_id[0:8], .text[0:44]] | @tsv'
0.9957	fact	3dce87e5	Pixel has a vet appointment at Clinica Arroi
0.9908	turn	9509bd23	Ana: she needs a vet, somewhere open on Satu
0.9535	turn	b8533f46	Ana: book Clinica Arroios for this Saturday 
0.5721	fact	b41bdb35	Ana prefers Saturday appointments
0.536	fact	c3d7a7af	Text Ana the day before the vet appointment
0.3996	turn	75173310	Ana: we adopted a greyhound called Pixel las
0.3015	fact	d53b6b8b	Ana owns a greyhound named Pixel
0.1626	turn	af2d5cc8	Ana: my number is 912 555 019, text me the d

The fact that holds the appointment is first. The two tool calls have no embedding and cannot appear.

by keyword
$ curl -s localhost:7777/v1/memory/recall -H "$K" -H "$J" \
  -d '{"text":"when is the vet appointment for Pixel","k":10}' \
  | jq -r '.hits[] | [(.score*1e4|round/1e4), .kind, .event_id[0:8], .text[0:44]] | @tsv'
4.6866	fact	c3d7a7af	Text Ana the day before the vet appointment
3.1517	turn	af2d5cc8	Ana: my number is 912 555 019, text me the d
2.5925	fact	3dce87e5	Pixel has a vet appointment at Clinica Arroi
2.0447	turn	b8533f46	Ana: book Clinica Arroios for this Saturday 
1.3526	fact	d53b6b8b	Ana owns a greyhound named Pixel
1.1752	turn	75173310	Ana: we adopted a greyhound called Pixel las
0.9173	turn	9509bd23	Ana: she needs a vet, somewhere open on Satu
0.8109	tool_call	740c7d03	book_vet(Clinica Arroios, 2026-10-10 10:00) 

The reminder is first. The turn with the phone number is second because it shares is and the with the question. The floor is half of the best score, so only the first three rows enter the fusion.

The appointment comes first because both rankings place it high. The phone-number turn comes third, ahead of two turns about the vet, on the strength of two common words: its keyword score cleared the floor. Of these ten events one contains is and two contain the, so both score as rare words. In a larger log common words weigh less, but nothing removes them.

The rules behind the example

  • A keyword candidate scoring under half of the best keyword score is dropped before fusion. BM25 returns every event that shares any term with the query, and rank fusion ignores how weak a match is. With plain rank fusion and no floor, multi-hop recall on the benchmark was 0.704, below the 0.734 of embeddings alone. Both figures are over all ten conversations with facts and query expansion done by the client.
  • An event stored without an embedding cannot appear in the embedding ranking. It scores 1 / (60 + rankkeyword), what an embedding-only candidate at that rank would, so the two kinds interleave.
  • Ties go to the candidate the embedding ranked, then to the better single rank, then to the event stored first.
  • kind, since and until filter after ranking. The candidate pool widens fourfold until k events pass or the customer's events run out.
  • The weight and the floor are fixed at 0.5, and a request cannot change them. They were chosen on five conversations of the benchmark and checked on the other five. The results page has the sweep, and the reference has this example with sliders for both.

Hops

With hops, the server walks SOURCE_FROM links out from all the ranked hits at once, one step at a time, in both directions: from a fact to its sources, and from a source to the facts derived from it. Events nearest to a hit are added first. The time filters apply to what the walk reaches. The kind filter does not, because the purpose of the walk is to cross from a fact to the turns and tool calls behind it. How far it goes and how much it adds are listed under recall with hops.

How much recall finds depends on what the application stores and sends. On a public benchmark of long conversations the server retrieved 0.889 of the evidence when the client extracted facts and sent each question as up to four queries, with 20 results per query. From raw turns, with the question as the only query and 20 results, it retrieved 0.794. The results page has the method and the categories where it is weakest.

One database per customer

Each customer (a tenant, in the settings and error messages) is a separate FFS database in its own directory, opened on that customer's first request. A request reads and writes the database its key resolves to. An id that belongs to another customer is not found there, and a search cannot reach another customer's text.

two customers, one server
$ echo "$A"; echo "$B"
Authorization: Bearer 00000000-0000-0000-0000-0000000000a1
Authorization: Bearer 00000000-0000-0000-0000-0000000000b2

$ curl -s localhost:7796/v1/memory/event -H "$A" -H "$J" -d @note.json; echo
{"event_id":"ff268751-73de-4b7e-abcb-d2dce1785afe"}

$ curl -s localhost:7796/v1/memory/event/$ID -H "$A" | jq -c '[.kind, .text]'
["turn","Ana: my number is 912 555 019, text me the day before"]

$ curl -s localhost:7796/v1/memory/event/$ID -H "$B"; echo
{"error":"not_found","detail":"no event with that id"}

$ curl -s localhost:7796/v1/memory/recall -H "$B" -H "$J" \
  -d '{"text":"number 912 555 019","k":5}'; echo
{"hits":[]}
$ ls data
00000000-0000-0000-0000-0000000000a1
00000000-0000-0000-0000-0000000000b2
memspine.lock

$ grep -rl "912 555 019" data
data/00000000-0000-0000-0000-0000000000a1/events.ffs.qlog

Recorded on 0.2.2. $ID is the id the append returned. Customer B has a directory because its first request opened a database for it. The phone number is in customer A's files only.

From a key to a customer

keyswhat is accepted
no control databaseAny UUID sent as the bearer token names a customer and can read and write. Nothing in it is secret. This mode is for local use, and the recordings on this page run in it.
with a control databaseCreated once with the memspine-admin tool. From then on, issued keys only. A key belongs to one customer, carries scopes (read, write), may expire and can be revoked. The server stores a SHA-256 hash of it and shows the key once. If the control database later goes missing, requests get a 503 and the UUID mode does not come back.

Where a request can be turned away

  1. KeyResolved to one customer and its scopes.otherwise 401
  2. Token bucket1,000 requests a second for that customer, with a burst of 100.otherwise 429
  3. Calls in flight32 engine calls running or queued for that customer.otherwise 429
  4. DatabaseThe customer's own directory, opened on first use.503 if the cap is reached and every open database is busy

Defaults. A key that lacks the scope a route needs gets a 403.

The in-flight bound exists because the token bucket counts requests and ignores how long each one runs. A customer whose calls are slow could stay inside its rate and still fill the worker pool that all customers share.

one customer over its limit
$ MEMSPINE_RATE_LIMIT_RPS=1 MEMSPINE_RATE_LIMIT_BURST=3 \
    MEMSPINE_BIND=127.0.0.1:7796 MEMSPINE_EMBEDDING_DIM=4 MEMSPINE_DATA_DIR=./data \
    memspine-server > memspine.log 2>&1 &

$ for i in 1 2 3 4 5; do curl -s -o /dev/null -w "%{http_code} " \
    localhost:7796/v1/whoami -H "$A"; done; echo
200 200 200 429 429 

$ curl -si localhost:7796/v1/whoami -H "$A" | tr -d "\r" | grep -i -e "^HTTP" -e "^retry-after" -e "^{"
HTTP/1.1 429 Too Many Requests
retry-after-ms: 1000
retry-after: 1
{"error":"rate_limited","detail":"tenant rate limit exceeded","retry_after_ms":1000}

$ curl -s -o /dev/null -w "%{http_code}\n" localhost:7796/v1/whoami -H "$B"
200

The same server restarted with a limit of one request a second and a burst of three, so that five requests in a row exceed it. Customer A is refused and customer B is served.

An open database holds about eight file descriptors and keeps its indexes in memory, so the server caps how many are open: (descriptor limit − 64) / 12, which is 80 at a limit of 1,024. At the cap it closes the least recently used one that has no request in flight. A customer with no request for an hour is closed too, and reopens on its next request.

Customers share one process, so a large customer's scan uses the same cores and memory as its neighbours' requests. The limits above bound how many calls it can run at once and leave the weight of each call unbounded.

What is on disk before the reply

  1. CheckedThe key, the limits and the event's shape. Nothing is written until all of it passes.dies here: nothing was written
  2. One batchUnder the customer's write lock: the event, its embedding, its text and its links.dies here: not confirmed, may or may not be there after a restart
  3. fsyncThe customer's log is flushed to disk.dies here: the same
  4. ReplyThe id is returned.dies after: the event is there after a restart

The path of an append, and what a crash at each step leaves.

Appends for one customer run one at a time. Reads run concurrently and do not wait for them, so a read can return an event in the moment between its batch and its fsync, before the append has been confirmed. A restart replays the log. A clean shutdown also writes a checkpoint, which only shortens the next start.

a hard kill after a confirmed append
$ curl -s localhost:7777/v1/memory/event -H "$K" -H "$J" -d @last.json; echo
{"event_id":"1eb38657-2649-47e6-b082-bf90e6da5c72"}

$ kill -9 $(lsof -tiTCP:7777 -sTCP:LISTEN); sleep 0.3; curl -s -m 2 localhost:7777/health; echo "curl exit $?"
curl exit 7

$ MEMSPINE_BIND=127.0.0.1:7777 MEMSPINE_EMBEDDING_DIM=4 MEMSPINE_DATA_DIR=./data \
    memspine-server > memspine.log 2>&1 &

$ curl -s localhost:7777/v1/memory/event/$E11 -H "$K"; echo
{"event_id":"1eb38657-2649-47e6-b082-bf90e6da5c72","kind":"turn","ts":1791366150000000,"text":"Ana: thanks, that is all for today","properties":{}}

A confirmed append, a hard kill, a restart, and the event read back. Durability was tested this way, by killing the process. It has not been tested by cutting power, and the data is one copy on one disk.

If a write to the storage engine fails, memspine stops trusting what that customer's open database holds in memory, because FFS does not undo everything an aborted batch did there. It drops the database without a checkpoint and opens it again from the log, answering 503 for the moment that takes. A customer whose log cannot be written does not open at all.

A retried append

With an Idempotency-Key header the event id is a version-5 UUID computed from the customer id and the key. The engine does not write an id twice, so a retry returns the first id and stores nothing. No table of keys is kept and the protection does not expire. The body of the retry is not compared with the first.

the same request twice
$ curl -s localhost:7777/v1/memory/event -H "$K" -H "$J" \
  -H "Idempotency-Key: turn-0941" -d @turn.json; echo
{"event_id":"af2d5cc8-2a80-583b-b7ed-941419d7a27f"}

$ curl -s localhost:7777/v1/memory/event -H "$K" -H "$J" \
  -H "Idempotency-Key: turn-0941" -d @turn.json; echo
{"event_id":"af2d5cc8-2a80-583b-b7ed-941419d7a27f"}

Both calls return one id, and one event is stored. The third group of the id starts with 5, the UUID version. A batch under a key derives an id for each position: replayed with the same number of events or fewer it returns the first ids, and with more it gets a 400.

What a delete removes, and when

The log is append-only, so a delete can only add a record to it. Getting the bytes off the disk is a second step that rewrites the customer's whole log.

momentget, list, recall, Cyphertext and embedding on disk
before DELETEreturnedin the customer's log
DELETE returnsa marker file, then one batch and one fsync of the lognot returned, links gone, durablestill in the log
the purge runson the next pass, every 60 s by default, or before the reply with ?purge=truenot returnedthe log is being rewritten and the customer's requests wait
the purge endsnot returnedin no file of the customer's directory

A pass also runs at startup, and the server purges when it closes an idle customer and at a clean shutdown. The marker file is written before the delete, so a crash between the two steps still ends in a purge after the restart, whether or not that customer sends another request.

forget, then look for what is left
$ grep -rl "912 555 019" data
data/00000000-0000-0000-0000-000000000042/events.ffs.qlog

$ curl -s -X DELETE "localhost:7777/v1/memory/event/$E9?purge=true" -H "$K"; echo
{"deleted":true,"purged":true}

$ grep -rl "912 555 019" data; echo "grep exit $?"
grep exit 1

$ curl -s -o /dev/null -w "%{http_code}\n" localhost:7777/v1/memory/event/$E9 -H "$K"
404

$ curl -s localhost:7777/v1/memory/event/$F10 -H "$K" | jq -c '.properties'
{"source_event_ids":["af2d5cc8-2a80-583b-b7ed-941419d7a27f","3dce87e5-4b2c-4cc0-96bc-338ce95f158f"]}

$ curl -s localhost:7777/v1/memory/cypher -H "$K" -H "$J" -d '{"query":
  "MATCH (f:Event)-[:SOURCE_FROM]->(s:Event) WHERE f.event_id = $id RETURN s.kind, s.text",
  "params":{"id":"'$F10'"}}' | jq -c '.rows[]'
["fact","Pixel has a vet appointment at Clinica Arroios on 2026-10-10 at 10:00, booking A-2291"]

The phone number is found in the customer's log. The event is deleted with ?purge=true, and the same search of the data directory finds nothing. The id then returns 404.

timekindidtextafter
09:14:02turn75173310Ana: we adopted a greyhound called Pixel last weekend
09:14:31turn9509bd23Ana: she needs a vet, somewhere open on Saturdays
09:14:33tool_call697ea0fafind_vets(city=Lisbon, open=saturday) -> Clinica Arroios, VetLuz
09:15:10turnb8533f46Ana: book Clinica Arroios for this Saturday at ten
09:15:12tool_call740c7d03book_vet(Clinica Arroios, 2026-10-10 10:00) -> confirmed A-2291
09:15:13factd53b6b8bAna owns a greyhound named Pixelfrom 75173310
09:15:14factb41bdb35Ana prefers Saturday appointmentsfrom 9509bd23, b8533f46
09:15:15fact3dce87e5Pixel has a vet appointment at Clinica Arroios on 2026-10-10 at 10:00, booking A-2291from b8533f46, 740c7d03
09:41:50turnaf2d5cc8Ana: my number is 912 555 019, text me the day beforedeleted
09:41:52factc3d7a7afText Ana the day before the vet appointmentfrom 3dce87e5kept

The same ten rows with the deleted one marked. Row 10 is kept, and its link to row 9 is gone. The turn appended in the kill test is left out.

What remains

  • The reminder in row 10 still lists the deleted id in its source_event_ids, as the fifth command shows. Other events' properties are not rewritten. The link is gone: the last query finds one source where there were two.
  • Row 9 was appended under an Idempotency-Key. Its id is kept, with no content, so that a late retry of that append stores nothing. Cypher patterns without a label see that record, so match (n:Event).
  • Backups taken before the purge.

A purge blocks the customer while it runs, and its time grows with the log. At 100,000 events the three purges in the run reported below took a median of 25 s and at most 39 s. Earlier runs on the same machine measured 4 to 12 s. The test suite repeats the check in the transcript after a purge: it searches the data directory for the text and for the embedding's bytes.

What it takes to run

memspine is one server process and a data directory. Once issued keys are turned on, the control database is one more file in that directory. The server makes no outbound calls, so computing embeddings is the application's cost. It speaks plain HTTP, and TLS goes in a reverse proxy in front of it. The work per request grows with that customer's event count. The sizes of other customers do not enter into it.

Backing up
Back up the data directory. Each customer's directory is that customer's complete state and is copied whole.
Removing a customer
Delete that customer's directory while the server is stopped. No route and no admin command removes one.
Instances and deploys
One server per data directory. The server holds an exclusive lock on a file in it and a second server refuses to start, as recorded below, so a deploy stops the old server before it starts the new one.
Changing the embedding model
The dimension is fixed per server and recorded per customer. A customer that has stored embeddings refuses to open under another dimension. The model itself is not recorded, so vectors from two models of the same dimension mix without warning.
Watching it
GET /health answers without a key. GET /metrics is in Prometheus text format, also without a key, so allow it at the proxy for the scraper only or move it to its own address. Every response has an x-memspine-trace-id header, and the server's log lines for that request carry the same id.

The reference has the files in the data directory and the settings, under storage and backup.

a second server, same directory
$ MEMSPINE_BIND=127.0.0.1:7791 MEMSPINE_EMBEDDING_DIM=4 MEMSPINE_DATA_DIR=./data \
    memspine-server 2>&1 | tail -1
Error: another memspine-server is using ./data; run one server per data directory

Time, memory and disk at two sizes

one customer, 384 dimensions10,000 events100,000 events
search, k = 10, p50 / p992.8 ms / 6.3 ms35 ms / 52 ms
fused recall, p50 / p993.7 ms / 7.5 ms45 ms / 50 ms
single append, p50 / p9916 ms / 20 ms16 ms / 20 ms
batch ingest, 1,000 per request17,304 /s9,788 /s
delete alone, p5020 ms24 ms
purge, median / longest3.4 s / 4.5 s25 s / 39 s
cold open437 ms3.3 s
on disk27.2 MiB269.7 MiB
peak resident memory214 MiB1.9 GiB

The scale mode of the benchmark tool on the 0.2.1 engine, 7 October 2026, on an arm64 laptop running macOS. The engine is called in process, with no HTTP or JSON. Random unit vectors, eight-word texts and two short properties per event, 100 queries, 50 single appends and three deletes, each followed by a purge. Memory is the peak of the whole run. Sizes in between were not run.

The rows bound by fsync (append, ingest, delete, purge) varied between runs on this machine while other work was using its disk. An earlier run on the 0.1.0 engine measured a single append at 8.0 ms and batch ingest at 32,580 events a second at 10,000 events, about twice as fast as here. Earlier runs measured purges of 0.6 to 2.5 s at 10,000 events and 4 to 12 s at 100,000. Search and fused recall took 12 to 24 percent longer here than in that earlier run.

  • A single append costs one fsync, so unbatched writes to one customer are bounded by the disk's fsync rate. The batch route stores 1,000 events per fsync.
  • An open customer holds its embeddings and indexes in memory. The run at 100,000 events, which includes three purges, peaked at 1.9 GiB resident, about 20 KiB per event. An earlier run on the 0.1.0 engine, before the benchmark included purges, peaked at 1.2 GiB. Budget for the customers that are open at once. One that has been idle for an hour is closed and pays the cold-open time on its next request.

Where another tool fits better

A customer will hold far more than 100,000 events. Search is an exact scan and nothing larger has been measured.

You need a second copy on another host. There is no replication.

You want the service to embed text, extract facts or summarise. It stores and returns what the application sends.

Your text is not in ASCII letters and you have no embedding model. Keyword recall will not match it.

You would give Cypher to customers you do not trust. The screen refuses the query shapes found to be explosive and leaves the cost of the rest unbounded.

A customer cannot wait for seconds while a purge of its log runs.

Running agents for your customers?

memspine is built by ERP.AI. If you are putting memory under agents that serve your customers, talk to us.