Summary
of each question's evidence retrieved on LoCoMo at 20 events per query, when the client extracts facts and rewrites the question (about 34 turns scored)
of the evidence when the client stores raw turns and sends the question as asked (20 turns scored)
of end-to-end answers accepted by a judge model, in a June run on an earlier engine with embedding-only retrieval
median search over one customer's 100,000 events, in process on a laptop; 2.8 ms at 10,000
Two more results are pass or fail. A server killed after confirming a write restarted with that write, and after a purge the forgotten text was in no file of the data directory. Both are recorded below.
The weakest results are open-domain questions, at 0.608 recall at best and 0.375 of answers accepted in the June run, where their recall was 0.550, and fused ranking on raw turns, where the gain over embeddings alone is under 0.01 and does not hold on the held-out conversations.
The retrieval and timing runs were made on the night of 7 October 2026, between the 0.2.1 and 0.2.2 releases. The current release is 0.2.2.
Retrieval on LoCoMo
LoCoMo is a public benchmark of 10 long conversations between two people, 5,882 turns in all, each spread over 19 to 32 dated sessions. Every question names the turns that hold its answer. The dataset has 1,986 questions in five categories. The published convention leaves out the 446 adversarial ones and scores the other 1,540: single-hop (841), temporal (321), multi-hop (282) and open-domain (96).
Each conversation is stored as one customer's memory, one event per turn. Evidence recall@20 is the share of a question's evidence turns that come back when every query returns 20 events, averaged over the 1,536 questions that name evidence (four open-domain questions name none). The score is computed from turn ids, and no model judges it.
Two kinds of client
memspine does not extract facts or rewrite queries. A client either does that work or leaves it out, and the benchmark harness plays both.
- with client-side work
- After each session the harness asks gpt-4o-mini for dated facts and stores each as a
factevent linked to its source turns. With each question it also sends up to three shorter queries written by the same model. Every query returns 20 events and the union is scored. That is 3.9 queries per question, and the union holds about 43 distinct turns under keyword-only ranking, 32 under embedding-only and 34 under fused, so the three rankings are not scored on the same number of turns. A retrieved fact counts as the turns it cites. - raw turns
- Only the turns are stored, and the question is the only query. Exactly 20 turns are scored.
Each client was run three ways: keyword only (BM25 over event text, no embedding sent), embedding only (exact cosine similarity, text-embedding-3-small at 384 dimensions), and both fused by reciprocal rank.
One question, scored six ways
Melanie artworksMelanie paintingsMelanie art portfolio| ranking | D1:12 | D8:6 | D13:8 | turns scored | recall |
|---|---|---|---|---|---|
| with client-side work: four queries, 20 events each | |||||
| keyword only | not retrieved | not retrieved | turn | 56 | 0.333 |
| embedding only | not retrieved | through a fact | not retrieved | 28 | 0.333 |
| both, fused | not retrieved | through a fact | turn | 38 | 0.667 |
| raw turns: the question as the only query | |||||
| keyword only | not retrieved | not retrieved | turn | 20 | 0.333 |
| embedding only | not retrieved | turn | turn | 20 | 0.667 |
| both, fused | not retrieved | not retrieved | turn | 20 | 0.333 |
One line of each of the six result files. Turns scored is the number of distinct turns in the scored union. Recall is the retrieved share of the three evidence turns.
With client-side work, the fused ranking keeps the turn that keyword ranking found and the one that embedding ranking reached through a fact, and recall rises from 0.333 to 0.667. On raw turns the keyword term displaces a turn that embeddings alone had returned, and recall falls from 0.667 to 0.333. One evidence turn comes back in no run. The question was picked because it shows both movements, and the next section counts them.
All 1,536 questions
By ranking, all 1,536 scored questions.
By question category, fused ranking.
| ranking | multi- hop 282 | temporal 321 | open- domain 92 | single- hop 841 | all 1,536 |
|---|---|---|---|---|---|
| with client-side facts and query rewriting, 32 to 43 turns scored per question | |||||
| keyword only | 0.585 | 0.902 | 0.532 | 0.885 | 0.812 |
| embedding only | 0.734 | 0.940 | 0.550 | 0.921 | 0.869 |
| both, fused | 0.746 | 0.946 | 0.608 | 0.946 | 0.889 |
| raw turns, the question as the only query, 20 turns scored | |||||
| keyword only | 0.279 | 0.670 | 0.310 | 0.689 | 0.587 |
| embedding only | 0.609 | 0.835 | 0.484 | 0.861 | 0.787 |
| both, fused | 0.590 | 0.844 | 0.531 | 0.873 | 0.794 |
Evidence recall@20 on LoCoMo, 10 conversations. The result files were written on the night of 7 October 2026, after the 0.2.1 release and before 0.2.2. The embedding-only and fused rows with client-side work score the same on every question as runs made on the 0.1.0 engine earlier that day. Embeddings and model responses were replayed from a cache recorded in June. Marked: the largest gain from fusing, keyword-only recall on multi-hop questions from raw turns, and the loss from fusing on multi-hop for raw turns.
Open-domain questions are the weakest category in every row but one, and no cause has been established. The exception is keyword-only recall on raw turns, where multi-hop is lower (0.279 against 0.310) because multi-hop questions share too few words with their evidence. A caller with no embedding model needs the client-side work: keyword-only recall is 0.812 with it and 0.587 without.
What each part of the client's work buys
| the client does | multi- hop | temporal | open- domain | single- hop | all |
|---|---|---|---|---|---|
| facts and rewritingabout 32 turns scored | 0.734 | 0.940 | 0.550 | 0.921 | 0.869 |
| rewriting onlyabout 38 turns scored | 0.727 | 0.865 | 0.542 | 0.884 | 0.830 |
| facts onlyabout 18 turns scored | 0.648 | 0.926 | 0.514 | 0.907 | 0.840 |
| neither20 turns scored | 0.609 | 0.835 | 0.484 | 0.861 | 0.787 |
Evidence recall@20, embedding-only ranking. The two middle rows were run in June on the engine before 0.1.0, whose vector index was approximate. The first row scored the same on both engines. Under each label is the mean number of distinct turns scored per question. No keyword-only or fused run has one part of the work without the other.
Taking the facts away from the full client costs temporal questions 7.5 points and single-hop 3.7. Taking the rewriting away costs multi-hop 8.6 and open-domain 3.6.
Facts alone reach 0.840 while scoring fewer turns than the raw client, about 18 against 20: a fact takes one of the 20 places, and some facts cite a turn already counted. Their gain therefore does not come from a larger budget. Rewriting alone reaches 0.830 on about 38 turns, and no run scored raw turns at that budget.
The price is one language-model call per session at ingest and one per question at recall, plus an embedding for every fact and every extra query, all made by the client.
Fusing helps when the client does the fact and query work
The fused ranking adds a half-weight keyword term to the embedding's reciprocal rank, after dropping keyword candidates that score under half of the best keyword match. A mean can rise because a few questions moved a lot, so fused and embedding-only ranking were compared one question at a time.
Questions for which fused ranking retrieved more, or less, of the evidence than embedding-only ranking. An exact sign test on the questions that changed gives p below 0.00001 with client-side work and 0.67 on raw turns. The test treats questions as independent, which questions about one conversation are not.
With client-side work, fusing retrieved more evidence for 70 questions and less for 24. Questions with all their evidence retrieved rose from 1,238 to 1,270, and questions with none fell from 117 to 88. On the five held-out conversations the count is 35 to 14, and fused recall is higher in each of the ten.
On raw turns the mean rises from 0.787 to 0.794, the count is 71 to 65, and multi-hop falls from 0.609 to 0.590. Fused recall is higher in five conversations of ten, four of them among the five the settings were tuned on, and on the held-out half it is 0.785 against 0.790 for embeddings alone. For a client that sends the question as asked, these runs do not show that fusing helps.
How the weight and the floor were chosen
The weight of the keyword term and the floor under it were chosen on the first five conversations and checked on the other five, in runs with client-side work.
| weight / floor | tuned on | held out |
|---|---|---|
| embedding only | 0.859 | 0.878 |
| 0.25 / 0.5 | 0.874 | 0.898 |
| 0.5 / 0.5, the default | 0.882 | 0.896 |
| 0.75 / 0.5 | 0.881 | 0.896 |
| 1.0 / 0.5 | 0.879 | 0.887 |
| 0.5 / 0.3 | 0.881 | 0.892 |
| 0.5 / 0.7 | 0.869 | 0.897 |
Recall@20 on each half of the conversations, with client-side work.
On the held-out half the default lifts recall from 0.878 to 0.896, and all four categories rise. Neighbouring settings land between 0.887 and 0.898, so the gain does not depend on the exact values.
With no floor, fusing lowered multi-hop recall over all ten conversations to 0.704, against 0.734 for embeddings alone and 0.746 with the floor. BM25 ranks every event that shares any word with the query, and rank fusion ignores how weak those matches are.
By conversation
| conversation | questions | with client-side work | raw turns | ||
|---|---|---|---|---|---|
| embedding | fused | embedding | fused | ||
| settings tuned on these | |||||
| conv-26 | 150 | 0.831 | 0.864 | 0.754 | 0.779 |
| conv-30 | 81 | 0.866 | 0.872 | 0.820 | 0.811 |
| conv-41 | 152 | 0.883 | 0.907 | 0.839 | 0.873 |
| conv-42 | 199 | 0.836 | 0.858 | 0.742 | 0.765 |
| conv-43 | 178 | 0.885 | 0.906 | 0.789 | 0.808 |
| held out | |||||
| conv-44 | 123 | 0.874 | 0.885 | 0.789 | 0.771 |
| conv-47 | 150 | 0.882 | 0.905 | 0.790 | 0.789 |
| conv-48 | 191 | 0.881 | 0.903 | 0.792 | 0.792 |
| conv-49 | 156 | 0.852 | 0.874 | 0.756 | 0.761 |
| conv-50 | 156 | 0.900 | 0.910 | 0.824 | 0.806 |
Evidence recall@20 by conversation, under the dataset's own ids. Marked: fused below embedding only. For conv-48 the difference is in the fourth decimal place.
End-to-end answers
J scores the answer built from the retrieved events: an answer model reads them, a judge model marks the answer against the gold answer, and J is the share marked correct.
J has been scored once, on 11 June 2026, on the engine before 0.1.0, with embedding-only retrieval and client-side facts and rewriting. The answer model and the judge were both gpt-4o-mini, at about 1,150 tokens per question on the answer path. The judge prompt was a published rubric, used verbatim.
| category | questions | recall@20 | F1 | J |
|---|---|---|---|---|
| single-hop | 841 | 0.921 | 0.606 | 0.851 |
| multi-hop | 282 | 0.734 | 0.433 | 0.762 |
| temporal | 321 | 0.940 | 0.586 | 0.760 |
| open-domain | 96 | 0.550 | 0.211 | 0.375 |
| all | 1,540 | 0.869 | 0.546 | 0.786 |
The June run. Recall is averaged over the 1,536 questions that name evidence, and F1 is word overlap between the answer and the gold answer.
J over the same 1,540 answers.
The first judge prompt tried asked whether the answer "matches in meaning", and the same answers scored 0.631. The published rubric tells the judge to be generous: an answer on the same topic as the gold answer counts, and dates match loosely.
The prompt alone is worth 15.5 points, so the 0.786 can be set beside another score only if that score was judged with the same prompt.
The 0.786 is one pass with no error bar. It depends on client code that is not part of memspine: the harness's answer prompt, and its layout of the retrieved events, facts first and then turns in date order under session headers. It also predates the current engine, which replaced an approximate vector index with exact search.
In embedding-only mode the engine since 0.1.0 reproduces June's evidence recall on every question and returns a byte-identical list for 1,416 of the 1,529 distinct questions (11 are asked twice). Of the other 113, 34 differ only in order and 79 in turns that are not evidence, so J should carry over. It was not re-run, and fused retrieval has no J. A full pass costs about 1.8 million tokens on the answer path, with the judge's calls on top.
Where the rejected answers came from
| category | rejected | their evidence, retrieved | all | part | none | refusals |
|---|---|---|---|---|---|---|
| single-hop | 125 | 84 | 6 | 35 | 65 | |
| temporal | 77 | 63 | 8 | 6 | 39 | |
| multi-hop | 67 | 21 | 42 | 4 | 19 | |
| open-domain | 60 | 27 | 10 | 21 | 49 | |
| all | 329 | 195 | 66 | 66 | 172 |
The answers the judge rejected in the June run. All, part and none say how much of the question's evidence had been retrieved. A refusal is the answer "No information available". Two rejected open-domain answers belong to questions that name no evidence.
Of the 329 rejected answers, 195 scored full evidence recall: every evidence turn was in the context the answer model read, as the turn itself or as a fact derived from it. Raising evidence recall would not change them. Refusals account for 172, and 102 of those had full evidence recall.
In two categories most rejected answers lacked some of their evidence: multi-hop, 46 of 67 (42 had part, 4 none), and open-domain, 31 of the 58 that name evidence (10 part, 21 none).
Timings at 10,000 and 100,000 events
The scale run fills one customer's memory with synthetic events: a random 384-dimension unit vector, eight words of text and two short properties each. It then times 50 single appends, 100 searches, 100 fused recalls, three forgets with a purge after each, and a cold open. The engine is called in process, so no figure includes HTTP or JSON.
| one customer, 384 dimensions | 10,000 events | 100,000 events |
|---|---|---|
| batch ingest, 1,000 per request | 17,304 /s | 9,788 /s |
| single append, p50 / p99 | 16 ms / 20 ms | 16 ms / 20 ms |
| search, k = 10, p50 / p99 | 2.8 ms / 6.3 ms | 35 ms / 52 ms |
| fused recall, k = 10, p50 / p99 | 3.7 ms / 7.5 ms | 45 ms / 50 ms |
| forget, p50 | 20 ms | 24 ms |
| purge, p50 / max | 3.4 s / 4.5 s | 25 s / 39 s |
| cold open | 437 ms | 3.3 s |
| on disk | 27.2 MiB | 269.7 MiB |
| peak memory of the run | 214 MiB | 1.9 GiB |
memspine-bench --mode scale on the 0.2.1 engine, 7 October 2026, on an Apple M-series laptop that was busy with other work. Sizes between the two were not run. Peak memory covers the whole run, purges included. The 100,000-event run on 0.1.0, before forget and purge were part of it, peaked at 1.2 GiB.
The write rows varied several-fold between runs, and this run is at the slow end. Batch ingest, append, forget and purge are bound by fsync, and other builds were using the laptop's disk. A second run at 10,000 events gave a single append of 8.0 ms and a purge of 2.2 s. Across earlier runs a purge took 0.6 to 2.5 s at 10,000 events and 4 to 12 s at 100,000, and on the 0.1.0 engine batch ingest ran at 32,580 and 18,657 events per second. Search moved little: 2.4 ms and 31 ms on 0.1.0.
Each bar is the 100,000-event figure divided by the 10,000-event figure, taken from the table as printed and rounded. Where a row gives two figures the first is used. The ratios for append, forget and purge carry the variation described above.
- Search compares the query with every embedding the customer has stored, so its time grows in proportion to that customer's event count.
- A single append costs one fsync and a forget two, so unbatched writes are bounded by the disk's fsync rate. The batch route stores up to 1,000 events per fsync.
- A purge rewrites the customer's whole log and blocks that customer while it runs: a median of 25 s and at most 39 s at 100,000 events in this run.
- An open customer holds its embeddings and indexes in memory. After an hour idle, by default, it is closed, and its next request pays the cold open.
Crashes and deletion
An append replies only after its batch is in the customer's log and the log has been fsynced. The direct check is to confirm a write, kill the server without warning, restart it and read the event back.
$ curl -s localhost:7777/v1/memory/event -H "$K" -H "$J" -d @last.json; echo
{"event_id":"1eb38657-2649-47e6-b082-bf90e6da5c72"}
$ kill -9 $(lsof -tiTCP:7777 -sTCP:LISTEN); sleep 0.3; curl -s -m 2 localhost:7777/health; echo "curl exit $?"
curl exit 7
$ MEMSPINE_BIND=127.0.0.1:7777 MEMSPINE_EMBEDDING_DIM=4 MEMSPINE_DATA_DIR=./data \
memspine-server > memspine.log 2>&1 &
$ curl -s localhost:7777/v1/memory/event/$E11 -H "$K"; echo
{"event_id":"1eb38657-2649-47e6-b082-bf90e6da5c72","kind":"turn","ts":1791366150000000,"text":"Ana: thanks, that is all for today","properties":{}}
Recorded against the released 0.2.2 binaries on macOS arm64. $K and $J hold the authorization and content-type headers, and $E11 the id the first command returned. The whole session is in session.txt. This is the only check that kills the process, and it reads one event back.
Forgetting has two steps. DELETE makes the event unreadable and removes its links in one batch, with two fsyncs: first a marker file that records the purge owed, then the log. A purge then rewrites the log so the bytes leave the disk. A purge pass starts within 60 seconds by default and takes the customers that owe one in turn. With ?purge=true the purge runs before the reply, and the reply says whether it succeeded ("purged": true).
$ grep -rl "912 555 019" data
data/00000000-0000-0000-0000-000000000042/events.ffs.qlog
$ curl -s -X DELETE "localhost:7777/v1/memory/event/$E9?purge=true" -H "$K"; echo
{"deleted":true,"purged":true}
$ grep -rl "912 555 019" data; echo "grep exit $?"
grep exit 1
From the same session. A phone number is found in the customer's log file, its event is deleted with a purge, and the same search of the data directory finds nothing.
What each claim rests on
| claim | how it is checked | what the check leaves out |
|---|---|---|
| A confirmed write survives the server dying | A test aborts the server with no shutdown step and starts a new one on the same directory. Properties, embeddings, the keyword index, links and an event with no embedding are read back. The recorded session above kills the process with kill -9 and reads one event back. | Power loss. Loss of the disk: there is one copy. |
| One customer cannot read another customer's events | A server test appends an event for one customer and requests it as another, which gets a 404. A revoked key gets a 401, and a key without the scope a route needs gets a 403. The recorded session on the product page finds a stored phone number in one customer's directory only. | Customers share one process, one worker pool and one heap. The effect of one customer's load on another's latency has not been timed. |
| A purged event's bytes are off the disk | Tests search every file in the data directory for the forgotten text, after ?purge=true and after the timer has run. An engine test also searches for the bytes of the embedding. | Backups taken earlier. Other events that cited the forgotten one keep its id in their properties. |
| A purge owed before a crash still runs | A marker file is written before the delete. A test drops the engine after a forget and starts a new one that receives no request; it finds the marker and the bytes are gone. | A customer that cannot be opened, for instance after a change of embedding dimension, keeps the bytes. Its purge is retried every 15 minutes. |
| Data written by 0.2.0 opens | A data directory written by the released 0.2.0 binary, and left unable to open by it, was opened by 0.2.1, which rebuilt the customer from its log. Tests build the older storage format and open it. | Data from before 0.1.0 is not read. A customer holding many strings over 4,096 bytes is slow to open the first time. |
| The published installer works | After each publish of this site the install command is run on Ubuntu, AlmaLinux 8 (glibc 2.28) and macOS arm64. The installed server stores, recalls, lists and purges an event, and its data directory is searched for the text. | Other platforms: there are two builds, Linux x86_64 and macOS arm64. The Linux build is stated to need glibc 2.17, and the oldest tested is 2.28. |
What has broken
All four tagged releases came out on two consecutive days. After 0.2.0, two independent reviews went looking for defects, and a finding counted once a second reviewer had reproduced it. 0.2.1 and 0.2.2 are the fixes. The last part of the review finished after 0.2.1 had gone out, and it found a defect that 0.2.1 itself had introduced.
- 0.1.0
The engine is rebuilt on FFS. Keyword and fused recall, batch append, reading an event back.
- 0.2.0
Forget and purge, listing, and closing customers that sit idle.
- 0.2.1
26 fixes for what the review found in 0.1.0 and 0.2.0.
- 0.2.2
9 fixes from the part of the review that finished late, one of them for a defect of 0.2.1.
The defects a caller could have met, newest fix first:
| fixed in | defect | trigger | now |
|---|---|---|---|
| 0.2.2 | Customers' data damaged and confirmed appends lost. Present in 0.2.1 only. | A second server started on a data directory already in use. With no request sent to it, it still damaged any customer that owed a purge. | The server locks its data directory and each open customer. A second one refuses to start. |
| 0.2.1, 0.2.2 | Customer ids accepted as keys | The control database going missing while the server ran (0.2.1), or a restart while it was missing (0.2.2). | The data directory records that it has had a control database. Requests get a 503 until it is back. |
| 0.2.1 | A customer's data could not be opened again | Any stored string over about 8 KB, then two purges. | Long values are stored where the engine does not index them. A customer already in that state is rebuilt from its log on first open. |
| 0.2.1 | Appends confirmed, then lost at restart | The customer's log could not be opened for writing, for instance with no file descriptors left. | Opening a customer writes a record and checks that the log grew. |
| 0.2.1 | One query aborted the server | About a thousand nested brackets, or a long chain of operators. | Queries are screened on the engine's own tokens, and each runs on its own 16 MB stack. |
| 0.2.1 | One query grew the server by gigabytes | UNWIND, or two patterns with no variable in common. | Refused with a 400. |
| 0.2.1 | One customer stalled the others | Many slow calls inside its rate limit. | 32 engine calls per customer running or queued, then a 429. |
| 0.2.1 | A retry restored a forgotten event | Replaying the Idempotency-Key it was appended under. | The id is kept and the replay stores nothing. |
| 0.2.1, 0.2.2 | Fused recall could not return a keyword match that had no embedding | k of 61 or less, once the customer held at least k embedded events. After 0.2.1, still true of events stored by an earlier version. | Such an event scores as an embedding-only match at the same rank. |
| 0.2.1 | Search returned fewer than k events | After a forget, until the customer was reopened. | A deleted event leaves the vector index at once. |
| 0.2.1 | Forgotten bytes stayed on disk | A crash before the purge, and no later request for that customer. | Purges owed are found at startup and on every pass. |
| 0.1.0 | The admin key was written to the server log | Every start with an admin key set: the configuration was logged whole. | The logged configuration redacts the key. |
| 0.1.0 | Issued keys were never rate limited | Any request made with an issued key. | The limiter counts against the customer the key resolves to. |
| 0.1.0 | A crash could drop up to 99 recent events from the vector index | The index was saved every 100th embedded event. Builds before 0.1.0 only. | Embeddings are written to the log with the event and fsynced before the reply. |
Every entry is in changelog.txt. One gap remains from this list: events forgotten under 0.2.0 left no record of their id, so a retry of one of their keys after upgrading still stores the event.
The code is covered by 162 Rust tests, 52 of which drive the HTTP server end to end, plus 29 tests in the Python client's suite and 21 in the TypeScript one. Every change runs formatting, lints with warnings denied, the tests, and a container image that is started and written through.
How the runs were made
The harness is memspine-bench, a crate in the source repository. That repository is private and the installer ships only the server and the admin tool, so these commands show how the runs are invoked. Running them takes the repository; talk to ERP.AI.
cargo run -p memspine-bench --release -- \
--mode recall --retrieval hybrid
cargo run -p memspine-bench --release -- \
--mode recall --retrieval hybrid \
--no-facts --no-expand
cargo run -p memspine-bench --release -- \
--mode scale --events 100000
The recorded runs also passed --workdir bench-runs/v0.2.0, and the scale runs were wrapped in /usr/bin/time -l, which is where the peak-memory row comes from.
- --retrieval
vector(the default),hybridortext: the embedding-only, fused and keyword-only rows.--text-weightand--text-floorset the fusion numbers.- --no-facts --no-expand
- Together, the raw-turn rows. Either flag alone, with
--retrieval vector, gives a middle row of the client-work table. - the dataset
locomo10.jsonfrom the benchmark's public repository, saved underbench-data/.- the model cache
- Embeddings and model responses are cached in
llm-cache.jsonlin the working directory. The retrieval runs replayed the cache recorded in June with the model API key set to a dummy value, so a missing entry would have failed the run instead of calling a model. - --mode full
- Adds the answer model and the judge, and needs a real key.
- --mode scale
- Needs neither the dataset nor a key.
--eventssets the size.
Every run writes one line per question: the queries sent, the turn ids retrieved, the recall, and in a full run the answer and the verdict. The worked question, the counts of questions that changed, the turns scored, the per-conversation figures and the breakdown of rejected answers on this page were computed from those files. They sit in the private repository with the harness and are not published here.
Not measured, or known to be weak
J beyond the June run. It has not been scored for fused retrieval, for the current engine, or with lineage hops in the answer context. With embedding-only ranking and client-side work on the 0.1.0 engine, one hop moved evidence recall by at most 0.003 in any category.
Open-domain questions. 0.608 recall at best and 0.375 J, the lowest category in every row here except keyword-only recall on raw turns. No cause has been established.
The client's work, taken apart. No raw-turn run returned more than 20 turns per question, against about 34 with client-side work, so how much of the gap between 0.794 and 0.889 is the larger budget is unknown. Keyword-only and fused ranking were run with both parts of the work or with neither.
A second benchmark or embedding model. Every retrieval figure comes from one English dataset and one embedding model at 384 dimensions.
Larger customers, load and server hardware. Search is an exact scan and nothing above 100,000 events was run. Every timing is one caller against one customer, in process, on a laptop that was doing other work. Requests over HTTP, concurrent callers and many customers at once have not been timed.
Power loss and disk failure. Durability was tested by killing or aborting the server. There is one copy of the data and no replication.
Keyword recall outside ASCII. Only ASCII letters and digits are indexed. Text in other scripts is stored and returned, and matches no keyword query.
The cost of a Cypher query. A screen refuses the shapes found to be explosive. The engine has no row or time budget, and the parser has not been fuzzed.