Repository navigation
[geaflow/ai-memory] Define VectorStore SPI with model metadata #851
Description
Activity
Hi ,
I have gone ahead and implemented this.
I created the VectorStore SPI interface and the VectorStoreMetadata class and added the required test cases as requested.I will open PR soon.
I have opened PR. Please let me know if any changes, enhancement or code structure improvements is required.
Thanks.I have opened PR. Please let me know if any changes, enhancement or code structure improvements is required. Thanks.
@Sumit2342 I'm very happy to see your PR. I've been busy with work today and have a meeting tonight. I plan to review your PR tomorrow.
Hello @Sumit2342, I'm glad to see your contribution. However, this VectorStore SPI layer is an intermediate layer in the graph memory, bridging the gap between higher and lower layers. Therefore, we hope it can support different types of search abstractions. Below are some suggestions for your reference.
1. Background and Motivation
1.1 Problem Statement
The current
EmbeddingIndexStorecouples storage, index building, and similarity computation into a single monolithic class. Every layer in the pipeline is non-replaceable:// Current architecture — tightly coupled public class EmbeddingIndexStore implements IndexStore { // Storage: in-memory HashMap + append-only JSONL private final HashMap<GraphEntity, List<EmbeddingResult>> embeddingMap = new HashMap<>(); // Index building: scans all vertices/edges, calls EmbeddingService public void buildIndex(MemoryGraph graph) { ... } // Retrieval: returns EmbeddingVector wrapping stored double[] public List<IVector> getEntityIndex(GraphEntity entity) { ... } // Similarity: EmbeddingVector.match() does brute-force linear scan externally }
This means:
- No pluggable storage backend: the HashMap and JSONL file are hardcoded; switching to a different persistence mechanism requires rewriting the entire class.
- No chunk vector support: the store only accepts
GraphEntityas keys. Phase 1 requires indexing document chunks (GM-AI-P1-018), which have a fundamentally different key type (chunk_idvsGraphEntity). - No metadata contract: there is no formal definition of what metadata accompanies a stored vector (model name, dimension, distance metric, format version). Without this, cross-store queries cannot validate compatibility.
- No test surface:
EmbeddingIndexStorehas zero unit tests. The coupling makes it impossible to test storage behavior in isolation.
// Example of the coupling problem: // To test "does this store reject wrong-dimension vectors?", // you must construct a full MemoryGraph, build all indexes, // and run a complete search pipeline. There is no isolated test point.
Core contradiction: Phase 1 requires two independent vector consumers —
ChunkVectorIndex(GM-AI-P1-018) andEntityVectorIndex(GM-AI-P1-019) — that share the same storage protocol but serve different source types. Without a unifiedVectorStoreSPI, these two issues will each invent their own storage, producing incompatible implementations.1.2 Quantified Impact
Scenario Issue without VectorStore SPI Consequence ChunkVectorIndex (018) and EntityVectorIndex (019) built independently Each defines its own storage format, metadata, and query semantics Two incompatible vector stores; hybrid retrieval (GM-AI-P1-036) cannot fuse results ANN adapter (020) needs to be plugged in later No standardized query interface; each consumer implements its own search ANN integration requires rewriting every consumer's retrieval path Local persistence (GM-AI-P1-014) needs vector storage EmbeddingIndexStore's JSONL is append-only with no delete support Cannot implement idempotent upsert or soft-delete without modifying the store Testing without remote embedding service No way to inject fake vectors into the store All tests require a live EmbeddingService, blocking offline CI1.3 Design Goals
This proposal designs a
VectorStoreSPI for geaflow-ai, addressing the following problems:- Decouple storage from source type: A single store handles both chunk vectors and entity vectors, distinguished by metadata, not by class hierarchy.
- Define a formal metadata contract: Six required fields (model_name, dimension, distance, index_version, created_at, format_version) ensure cross-store compatibility and enable fail-closed validation.
- Provide reference implementations: An
InMemoryVectorStorefor testing and aLocalVectorStorefor persistence, both validated by the same contract test suite. - Preserve backward compatibility: Existing
EmbeddingIndexStorebehavior is not modified; it is wrapped byEntityVectorIndex(GM-AI-P1-019) in a later issue.
1.4 Non-Goals (Scope Out)
- ANN indexing: No HNSW, FAISS, or ScaNN integration. Phase 1 uses brute-force linear scan. ANN adapter is GM-AI-P1-020.
- Chunk semantics: The VectorStore does not know what a "chunk" is. Chunk-specific logic (source_ref interpretation, text_hash validation) belongs in
ChunkVectorIndex(GM-AI-P1-018). - Entity semantics: Similarly, entity-specific logic (vertex/edge type distinction, canonical ID resolution) belongs in
EntityVectorIndex(GM-AI-P1-019). - Fusion and rerank: Score normalization, RRF, and reranking are GM-AI-P1-036 and GM-AI-P1-037.
- Schema validation: Vertex/edge schema validation is GM-AI-P1-008.
- Distributed storage: No multi-node replication, sharding, or consensus. Local-only for Phase 1.
- Remote embedding calls: The VectorStore does not call any embedding service. It receives pre-computed vectors.
2. Constraints
2.1 Architectural Constraints
C-1: VectorStore must not extend IndexStore
// Current interface — source-specific public interface IndexStore { List<IVector> getEntityIndex(GraphEntity entity); } // New interface — source-agnostic public interface VectorStore { void upsert(VectorRecord record); List<VectorHit> search(VectorQuery query); // ... }
These are different abstraction layers.
IndexStoreis source-specific (keyed byGraphEntity);VectorStoreis source-agnostic (keyed byString vectorId). The adapter between them isEntityVectorIndex(GM-AI-P1-019), which implementsIndexStoreand delegates toVectorStoreinternally.C-2: Existing EmbeddingIndexStore must not be modified
The
EmbeddingIndexStoreclass, its tests (if any), and its public contract remain untouched.EntityVectorIndexwraps it as a compatibility layer. This ensures no regression in existing GraphMemoryServer behavior.C-3: All metadata fields must be validated at upsert time
Dimension mismatch, missing model_name, and missing distance metric must be detected at write time (fail-closed), not at query time. This prevents corrupted stores that silently return garbage results.
C-4: Vector IDs must be stable and deterministic
Vector IDs must not depend on wall-clock time, insertion order, or JVM object identity. They must be reproducible across runs for the same input data. This is required by the idempotency contract (GM-AI-P1-011).
2.2 Compatibility Constraints
C-5: Backward compatible — no silent behavior changes
The
GraphMemoryServer.search()method currently usesinstanceofchecks to dispatch betweenEmbeddingOperatorandSessionOperator. This dispatch logic must continue to work. The newVectorStoreis an additional path, not a replacement.C-6: Must work offline without remote embedding credentials
All contract tests and reference implementations must be runnable without any
EmbeddingServiceconnection. Fake embeddings (deterministicdouble[]) must be sufficient.C-7: Persistence format must be forward-compatible
The
format_versionfield is not decorative. On load, theLocalVectorStoremust check it. Unknown versions must fail closed (throw exception), not skip silently.2.3 Performance Constraints
C-8: Upsert and search must be in-memory operations
No network IO or disk IO during the hot path. Disk writes are batched and async. This aligns with the existing
EmbeddingIndexStorepattern (in-memory HashMap + periodic JSONL flush).C-9: Linear scan is acceptable for Phase 1
The
search()method uses brute-force cosine similarity. The performance contract is: O(N) per query where N is the number of stored vectors. ANN acceleration is deferred to GM-AI-P1-020.
3. Current State Analysis
3.1 Existing Index Architecture
Class Path Role Issue IndexStoreindex/IndexStore.javaSingle-method interface: getEntityIndex(GraphEntity)No vector search, no metadata EmbeddingIndexStoreindex/EmbeddingIndexStore.javaEmbedding storage + index building + JSONL persistence Zero tests, tightly coupled EntityAttributeIndexStoreindex/EntityAttributeIndexStore.javaKeyword-based index (verbalized text) Out of scope IndexStoreCacheindex/IndexStoreCache.javaStatic singleton aggregating stores Only registers EntityAttributeIndexStore IVectorindex/vector/IVector.javaVector interface with match()methodAd-hoc similarity, no metadata EmbeddingVectorindex/vector/EmbeddingVector.javaDouble-array vector, cosine similarity No dimension validation KeywordVectorindex/vector/KeywordVector.javaString-array vector, overlap counting Asymmetric normalization 3.2 How Indexes Are Used in GraphMemoryServer
// Source: GraphMemoryServer.java, lines 87-96 public List<GraphEntity> search(String query) { for (IndexStore store : indexStores) { if (store instanceof EntityAttributeIndexStore) { // keyword-based retrieval SessionOperator operator = new SessionOperator(query, (EntityAttributeIndexStore) store); // ... } else if (store instanceof EmbeddingIndexStore) { // embedding-based retrieval EmbeddingOperator operator = new EmbeddingOperator(query, (EmbeddingIndexStore) store); // ... } } }
Key finding: The dispatch uses
instanceofchecks, not polymorphism. Adding a new index type requires modifying this method. TheVectorStoreSPI does not fix this dispatch problem (that is GM-AI-P1-019's job), but it provides the storage foundation that the new index types will use.3.3 Three Missing Capabilities
- Source-agnostic vector storage — current storage is keyed by
GraphEntity, not byString vectorId - Metadata validation — no dimension check, no model compatibility check, no format version check
- Pluggable persistence — JSONL is hardcoded; no way to swap in a different backend
4. Design
4.1 Overall Architecture
┌─────────────────────────────────────────────────────────────────┐ │ Consumers (GM-AI-P1-018, 019, 020) │ │ ┌──────────────────┐ ┌──────────────────┐ ┌───────────────┐ │ │ │ ChunkVectorIndex │ │ EntityVectorIndex │ │ ANN Adapter │ │ │ │ (chunk_id, src) │ │ (entity_id, type) │ │ (future) │ │ │ └────────┬─────────┘ └────────┬─────────┘ └──────┬────────┘ │ │ │ │ │ │ ├───────────┼──────────────────────┼────────────────────┼───────────┤ │ ▼ ▼ ▼ │ │ ┌─────────────────────────────────────────────────────────────┐ │ │ │ VectorStore SPI │ │ │ │ upsert(VectorRecord) search(VectorQuery) markDeleted() │ │ │ └──────────────────────────┬──────────────────────────────────┘ │ │ │ │ ├─────────────────────────────┼────────────────────────────────────┤ │ ┌─────────────────┼─────────────────┐ │ │ ▼ ▼ ▼ │ │ ┌────────────────┐ ┌────────────────┐ ┌──────────────────┐ │ │ │ InMemoryVector │ │ LocalVector │ │ Future: Remote │ │ │ │ Store │ │ Store (JSONL) │ │ VectorStore │ │ │ │ (testing) │ │ (persistence) │ │ (Phase 2+) │ │ │ └────────────────┘ └────────────────┘ └──────────────────┘ │ └─────────────────────────────────────────────────────────────────┘4.2 Data Structures
4.2.1 VectorRecord
public class VectorRecord { /** Stable unique ID. Must not depend on wall-clock time. */ private final String vectorId; /** The embedding vector. Length must match store dimension. */ private final double[] embedding; /** Source discriminator: "chunk", "entity", or custom type. */ private final String sourceType; /** The source identifier (chunk_id or entity_id). */ private final String sourceId; /** Additional metadata (source_ref, text_hash, extractor_version, etc.). */ private final Map<String, String> metadata; // Constructor validates: vectorId non-null, embedding non-null and non-empty, // sourceType non-null, sourceId non-null. // metadata may be null (defaults to empty map). }
Design decision:
sourceTypeandsourceIdare top-level fields (not buried in metadata) because every consumer needs to know the source kind for dispatch. Themetadatamap holds optional, source-specific attributes.4.2.2 VectorQuery
public class VectorQuery { /** The query embedding vector. */ private final double[] queryVector; /** Maximum number of results to return. */ private final int topK; /** Optional metadata filters (e.g., model_name, sourceType). */ private final Map<String, String> filterMetadata; // Constructor validates: queryVector non-null, topK > 0. }
4.2.3 VectorHit
public class VectorHit { /** The matched vector's ID. */ private final String vectorId; /** Similarity score (higher = more similar for cosine). */ private final double score; /** The full record of the matched vector. */ private final VectorRecord record; }
4.2.4 VectorStoreMetadata
Six required fields, each with a specific engineering rationale:
Field Type Purpose Failure mode if missing/wrong model_nameStringIdentifies the embedding model Vectors from different models live in incompatible semantic spaces; cross-model queries return meaningless results dimensionintVector dimensionality Dimension mismatch causes dot()out-of-bounds or silent corruptiondistanceDistanceMetricenumCOSINE, L2, DOT_PRODUCT Different metrics produce incomparable scores; fusion (GM-AI-P1-036) would produce wrong rankings index_versionStringIndex format version Old format may not deserialize after code upgrade created_atlong(epoch millis)Creation timestamp Cannot detect stale indexes or trigger re-indexing format_versionintSerialization format version Unknown format silently corrupts data; must fail closed public enum DistanceMetric { COSINE, L2, DOT_PRODUCT } public class VectorStoreMetadata { private final String modelName; private final int dimension; private final DistanceMetric distance; private final String indexVersion; private final long createdAt; private final int formatVersion; // All fields are required (no defaults). Construction fails if any is null/invalid. }
Design decision TD-1: All metadata fields are mandatory — Optional metadata leads to downstream consumers silently assuming defaults, which produces hard-to-debug incompatibilities. A store that does not know its own dimension cannot safely upsert vectors. The cost of requiring all fields is a slightly verbose constructor; the benefit is that every store is self-describing.
4.2.5 VectorStoreException
public class VectorStoreException extends RuntimeException { public enum ErrorCode { DIMENSION_MISMATCH, // vector.length != store.dimension METADATA_INCOMPLETE, // required metadata field missing MODEL_MISMATCH, // query filter model_name != store model_name FORMAT_VERSION_UNKNOWN, // unknown format_version on load RECORD_NOT_FOUND, // vectorId does not exist PERSISTENCE_ERROR // disk I/O failure } private final ErrorCode errorCode; private final String detail; }
4.3 Core Interface
public interface VectorStore { /** * Insert or update a single vector record. * Upsert is idempotent: same vectorId overwrites the previous record. * * @throws VectorStoreException(DIMENSION_MISMATCH) if record.embedding.length != metadata.dimension * @throws VectorStoreException(METADATA_INCOMPLETE) if required metadata fields are missing */ void upsert(VectorRecord record); /** * Batch insert/update. Atomic per record; partial failure does not * roll back already-committed records. */ void upsertBatch(List<VectorRecord> records); /** * Search for nearest vectors using the configured distance metric. * Results are ordered by score descending (best match first). * * @return empty list if no results match, never null * @throws VectorStoreException(MODEL_MISMATCH) if filterMetadata specifies * a model_name different from the store's model_name */ List<VectorHit> search(VectorQuery query); /** * Soft-delete a vector by id. The record remains on disk but is * excluded from future search results. * * @throws VectorStoreException(RECORD_NOT_FOUND) if vectorId does not exist */ void markDeleted(String vectorId); /** * Returns metadata about this store instance. */ VectorStoreMetadata getMetadata(); /** * Release resources. Implementations must flush pending writes before returning. */ void close(); }
4.4 Reference Implementations
4.4.1 InMemoryVectorStore
Primary purpose: contract validation and unit testing.
public class InMemoryVectorStore implements VectorStore { private final VectorStoreMetadata metadata; private final ConcurrentHashMap<String, VectorRecord> store = new ConcurrentHashMap<>(); private final Set<String> deletedIds = ConcurrentHashMap.newKeySet(); // search(): brute-force linear scan // - Compute cosine similarity against all non-deleted records // - Sort by score descending, return topK // - O(N) per query, acceptable for Phase 1 // upsert(): validate dimension, then store.put(vectorId, record) // markDeleted(): deletedIds.add(vectorId) }
- Not persistent; resets on process restart.
- Suitable for all contract tests that do not require persistence.
- The
search()implementation is the reference for cosine similarity scoring.
4.4.2 LocalVectorStore
Primary purpose: local persistent storage, replacing the current EmbeddingIndexStore JSONL approach.
Storage format: append-only JSONL with per-record checksum Line format: {"vectorId":"...", "embedding":[...], "sourceType":"...", "sourceId":"...", "metadata":{...}, "__checksum":"...", "__deleted":false} On startup: 1. Read all lines from JSONL file 2. Validate checksum for each line 3. Lines with checksum mismatch → quarantine (write to .quarantine file, log warning) 4. Truncated lines (invalid JSON) → quarantine 5. Build in-memory index from valid lines 6. Check format_version: unknown version → throw VectorStoreException(FORMAT_VERSION_UNKNOWN) On upsert: 1. Validate dimension and metadata 2. Append new line to JSONL (with checksum) 3. Update in-memory index 4. If same vectorId existed: the old line is NOT modified (append-only). The in-memory index always points to the latest line for a given vectorId. On markDeleted: 1. Append a delete marker line: {"vectorId":"...", "__deleted":true, "__checksum":"..."} 2. Update in-memory index (remove from searchable set)Design decision TD-2: Append-only with logical delete — In-place modification of JSONL is fragile (partial writes corrupt the file). Append-only with a delete marker ensures that crash recovery is always possible: the last occurrence of a vectorId wins. The cost is that the file grows over time; a compaction step can be added in Phase 2.
4.5 Distance Metric Implementation
public class DistanceUtils { /** * Compute distance between two vectors according to the specified metric. * Returns a similarity score where higher = more similar. * * @throws VectorStoreException(DIMENSION_MISMATCH) if vectors have different lengths */ public static double compute(double[] a, double[] b, DistanceMetric metric) { if (a.length != b.length) { throw new VectorStoreException(ErrorCode.DIMENSION_MISMATCH, "Vector lengths differ: " + a.length + " vs " + b.length); } switch (metric) { case COSINE: return cosineSimilarity(a, b); case L2: return 1.0 / (1.0 + euclideanDistance(a, b)); // invert so higher = better case DOT_PRODUCT: return dotProduct(a, b); default: throw new IllegalArgumentException("Unknown metric: " + metric); } } private static double cosineSimilarity(double[] a, double[] b) { double dot = 0.0, normA = 0.0, normB = 0.0; for (int i = 0; i < a.length; i++) { dot += a[i] * b[i]; normA += a[i] * a[i]; normB += b[i] * b[i]; } if (normA == 0.0 || normB == 0.0) return 0.0; return dot / (Math.sqrt(normA) * Math.sqrt(normB)); } }
Thanks @kitalkuyo-gita for detailed design document.
I have already started refactoring my implementation to match this exact blueprint.
I will push the updated code to my PR for your review.Thanks again for the review.
Hi @kitalkuyo-gita ,
I have updated my PR. Can you please review it again and let me know if any other changes or enhancement is required.
Thanks.
Priority: P0
Difficulty: Intermediate
Context:
EmbeddingIndexStorestores embeddings forGraphEntityin JSONL and memory map. Phase 1 needs a generic vector store contract for both chunks and graph entities.Scope:
VectorStoreinterface.model_name,dimension,distance,index_version,created_at,format_version.Constraints:
EmbeddingIndexStorebehavior.Acceptance Criteria:
Suggested paths:
geaflow-ai/src/main/java/org/apache/geaflow/ai/index/vectorstore