Presenter notes and evidence
Interview pitch: I use turbopuffer for graph-informed semantic retrieval over specialist reference corpora today. I chose it with a larger workload in mind: many independent customer memory indexes, growing histories, and uneven activity. Unlimited namespace count, object-backed durable storage, cached serving, and native embeddings fit that direction. My next phase is recursive customer memory: each follow-up recalls relevant context, new observations become searchable memory, and qualified insights can feed a versioned shared knowledge graph.
- Verified account footprint: On October 2, 2026, the live account contained nine namespaces and approximately 208,000 records, with about 0.87 GB of logical data. Most records are domain-reference entries. Each namespace reports a current index and a configured native-embedding source field and vector index. These counts describe a current reference corpus, not production-scale customer memory.
- What I use today: Meaning-based retrieval over separate reference pools, metadata boundaries, and a native-embedding integration. The application combines retrieved references with structured context and controls generation. Jev is used for rubric scoring and classification. Proprietary domain details, source identities, record examples, model versions, and individual dataset sizes are intentionally omitted.
- Future customer memory and multi-tenancy: Keep shared reference corpora separate. Give each customer an independent turbopuffer namespace containing meaning-searchable memory records: preferences, prior context, and observed outcomes. Unlimited namespace count supports growth in separate customer indexes; storage and throughput still have limits. The application authenticates the customer and authorizes namespace routing. A query searches eligible record vectors within the selected namespace and returns top matches; it does not put the customer's entire history into the language model or guarantee an exhaustive scan of every vector. Recent conversation and a current customer profile complement retrieved memories.
- Recursive questioning: A follow-up becomes a new focused query, informed by the current conversation. Stored document vectors can be reused instead of re-embedding all prior history. New or changed memory records and new query text still require compatible embeddings. Native embeddings remove a provider integration; lower embedding cost or faster embedding computation has not been established.
- turbopuffer in the knowledge feedback loop: In the proposed loop, turbopuffer persists customer memory and retrieves semantically related feedback for application processing. Jev helps qualify candidate signals through scoped classification and scoring; uncertain decisions can be reviewed. An ETL layer extracts permitted, de-identified candidate insights and applies evidence, provenance, conflict, and versioning checks before shared graph updates. The next interaction combines updated graph context with another scoped turbopuffer memory query. Jev's graph-feedback role and this memory workflow are proposed extensions; its existing role is scoring and classification. Confidence is a decision signal, not proof that a claim is true. The loop accumulates application knowledge; it does not train the language model. Measured improvement requires evaluation.
- Why turbopuffer: My target workload has many separate customer histories, some of which are infrequently queried. Object-backed durable storage plus SSD and memory caching separates corpus growth from the active serving working set. Independent namespace indexes and shared managed serving fit that direction. Unlimited namespace count does not mean unlimited per-namespace capacity or limitless throughput. Cold versus warm latency matters for an interactive agent.
- Why not Redis: The replacement diagram illustrates Redis Cloud Flex with Search, currently a Pro preview supporting HASH records, TAG filters, and vector indexes on tiered storage. This deployment requires RAM/flash sizing and configured persistence. Its preview does not yet support JSON documents or NUMERIC fields; Redis Software has a different feature set. Redis can implement customer memory and may suit an existing Redis stack. My preference is object-backed durable retrieval, while any savings require a sized comparison of data, index overhead, replicas, query load, persistence, and embeddings.
- Why not PostgreSQL + pgvector: The illustrated baseline integrates an embedding provider, PostgreSQL tenant filters, vector indexes, and SQL database capacity. pgvector itself searches supplied vectors; managed PostgreSQL integrations can automate embedding generation. PostgreSQL is credible at my current scale and may be preferable for joins and SQL consolidation. My preference is a dedicated managed retrieval service as customer histories grow.
- Why not Milvus / Zilliz: The illustrated Zilliz Cloud design uses a shared collection and partition key to route by tenant ID. This strategy supports many logical tenants under a common schema; multiple tenants can occupy a physical partition. Independent customer collections are another option within deployment limits. Zilliz also has embedding functions and separated storage/compute. I favor independent namespace indexes without a namespace-count cap; I would evaluate Zilliz's collection and partition strategies before claiming a performance or cost advantage.
- Why not MongoDB: The illustrated Atlas design combines document storage, automated Voyage embeddings, and synchronized indexes on dedicated Search Nodes. Search capacity can scale independently. This offers integration benefits for an existing MongoDB application; my graph-based application instead adds a dedicated retrieval service. Automated embeddings are not an exclusive turbopuffer feature.
- Why not S3 Vectors: The illustrated direct-API design generates query vectors with an embedding provider and retrieves full source records from S3 objects using vector IDs. Small source text can instead fit within vector metadata, subject to payload limits; AWS integrations can also manage embedding and retrieval orchestration. S3 Vectors is itself managed, durable, and object-backed. My preference is the current text-and-retrieval API workflow, rather than a claim that object-backed search is unique.
- How to compare: In chapter seven, the technology selector replaces the retrieval architecture. Only the selected store appears in the diagram. The customer task and application graph / Jev / ETL responsibilities are held constant; the embedding, tenant organization, storage, and serving arrangement changes. The diagrams show representative implementations, not mandatory configurations or measured speed differences.
- What remains to prove: Compare retrieval relevance under filters, response quality, p95 latency across repeated questions, embedding costs, engineering effort, and total cost as the corpus and active customer count grow. The animation is schematic and encodes no measured speed or cost advantage.
- Architecture precedent: turbopuffer's published TELUS case study describes a unique namespace per personalized copilot and reports more than 25,000 namespaces. This is a relevant precedent for many separate customer memory indexes, not a performance or cost measurement of my application. Graph-informed retrieval and customer-scoped semantic memory also have established precedents in GraphRAG and LangGraph documentation.
The customer-memory and shared-knowledge feedback designs are future extensions, clearly distinguished from today's reference retrieval. This presentation contains illustrative labels, not actual customer records.
Present or share: Open this file in Chrome, Edge, or Safari. Use Play story, Back / Next, or the chapter selector. Arrow keys work when a presentation control has focus. Use the browser's full-screen command for screen sharing. The presentation and presenter notes work offline; documentation links require internet. Download this file and open it in a browser to explore the interactive presentation.
Technical references: namespaces and storage/compute separation, service limits, turbopuffer architecture, native embeddings, query semantics, TELUS multi-tenancy case study, Jev typed decisions, Jev confidence, LangGraph memory patterns, graph-informed retrieval, pgvector, Redis Cloud Flex Search preview, Zilliz tenant strategies, Zilliz embedding functions, MongoDB deployment architecture, MongoDB automated embeddings, and S3 Vectors. Reviewed October 2, 2026.