Skip to content

[geaflow/ai-memory] Define VectorStore SPI with model metadata #851

Description

@kitalkuyo-gita

Priority: P0
Difficulty: Intermediate

Context: EmbeddingIndexStore stores embeddings for GraphEntity in JSONL and memory map. Phase 1 needs a generic vector store contract for both chunks and graph entities.

Scope:

  • Add VectorStore interface.
  • Define metadata fields: model_name, dimension, distance, index_version, created_at, format_version.
  • Add tests for dimension mismatch, missing metadata, and model mismatch.

Constraints:

  • Do not implement ANN in this issue.
  • Do not break existing EmbeddingIndexStore behavior.

Acceptance Criteria:

  • Loading vectors with wrong dimension fails closed.
  • Metadata is persisted in local fixture format.

Suggested paths:

  • geaflow-ai/src/main/java/org/apache/geaflow/ai/index/vectorstore

Activity

  1. Sumit2342 commented on Aug 24, 2026

    @Sumit2342

    Hi ,
    I have gone ahead and implemented this.
    I created the VectorStore SPI interface and the VectorStoreMetadata class and added the required test cases as requested.

    I will open PR soon.

  2. Sumit2342 commented on Aug 24, 2026

    @Sumit2342

    I have opened PR. Please let me know if any changes, enhancement or code structure improvements is required.
    Thanks.

  3. kitalkuyo-gita commented on Aug 25, 2026

    @kitalkuyo-gita
    MemberAuthor

    I have opened PR. Please let me know if any changes, enhancement or code structure improvements is required. Thanks.

    @Sumit2342 I'm very happy to see your PR. I've been busy with work today and have a meeting tonight. I plan to review your PR tomorrow.

  4. kitalkuyo-gita commented on Aug 27, 2026

    @kitalkuyo-gita
    MemberAuthor

    Hello @Sumit2342, I'm glad to see your contribution. However, this VectorStore SPI layer is an intermediate layer in the graph memory, bridging the gap between higher and lower layers. Therefore, we hope it can support different types of search abstractions. Below are some suggestions for your reference.

    1. Background and Motivation

    1.1 Problem Statement

    The current EmbeddingIndexStore couples storage, index building, and similarity computation into a single monolithic class. Every layer in the pipeline is non-replaceable:

    // Current architecture — tightly coupled
    public class EmbeddingIndexStore implements IndexStore {
        // Storage: in-memory HashMap + append-only JSONL
        private final HashMap<GraphEntity, List<EmbeddingResult>> embeddingMap = new HashMap<>();
    
        // Index building: scans all vertices/edges, calls EmbeddingService
        public void buildIndex(MemoryGraph graph) { ... }
    
        // Retrieval: returns EmbeddingVector wrapping stored double[]
        public List<IVector> getEntityIndex(GraphEntity entity) { ... }
    
        // Similarity: EmbeddingVector.match() does brute-force linear scan externally
    }

    This means:

    • No pluggable storage backend: the HashMap and JSONL file are hardcoded; switching to a different persistence mechanism requires rewriting the entire class.
    • No chunk vector support: the store only accepts GraphEntity as keys. Phase 1 requires indexing document chunks (GM-AI-P1-018), which have a fundamentally different key type (chunk_id vs GraphEntity).
    • No metadata contract: there is no formal definition of what metadata accompanies a stored vector (model name, dimension, distance metric, format version). Without this, cross-store queries cannot validate compatibility.
    • No test surface: EmbeddingIndexStore has zero unit tests. The coupling makes it impossible to test storage behavior in isolation.
    // Example of the coupling problem:
    // To test "does this store reject wrong-dimension vectors?",
    // you must construct a full MemoryGraph, build all indexes,
    // and run a complete search pipeline. There is no isolated test point.

    Core contradiction: Phase 1 requires two independent vector consumers — ChunkVectorIndex (GM-AI-P1-018) and EntityVectorIndex (GM-AI-P1-019) — that share the same storage protocol but serve different source types. Without a unified VectorStore SPI, these two issues will each invent their own storage, producing incompatible implementations.

    1.2 Quantified Impact

    Scenario Issue without VectorStore SPI Consequence
    ChunkVectorIndex (018) and EntityVectorIndex (019) built independently Each defines its own storage format, metadata, and query semantics Two incompatible vector stores; hybrid retrieval (GM-AI-P1-036) cannot fuse results
    ANN adapter (020) needs to be plugged in later No standardized query interface; each consumer implements its own search ANN integration requires rewriting every consumer's retrieval path
    Local persistence (GM-AI-P1-014) needs vector storage EmbeddingIndexStore's JSONL is append-only with no delete support Cannot implement idempotent upsert or soft-delete without modifying the store
    Testing without remote embedding service No way to inject fake vectors into the store All tests require a live EmbeddingService, blocking offline CI

    1.3 Design Goals

    This proposal designs a VectorStore SPI for geaflow-ai, addressing the following problems:

    1. Decouple storage from source type: A single store handles both chunk vectors and entity vectors, distinguished by metadata, not by class hierarchy.
    2. Define a formal metadata contract: Six required fields (model_name, dimension, distance, index_version, created_at, format_version) ensure cross-store compatibility and enable fail-closed validation.
    3. Provide reference implementations: An InMemoryVectorStore for testing and a LocalVectorStore for persistence, both validated by the same contract test suite.
    4. Preserve backward compatibility: Existing EmbeddingIndexStore behavior is not modified; it is wrapped by EntityVectorIndex (GM-AI-P1-019) in a later issue.

    1.4 Non-Goals (Scope Out)

    • ANN indexing: No HNSW, FAISS, or ScaNN integration. Phase 1 uses brute-force linear scan. ANN adapter is GM-AI-P1-020.
    • Chunk semantics: The VectorStore does not know what a "chunk" is. Chunk-specific logic (source_ref interpretation, text_hash validation) belongs in ChunkVectorIndex (GM-AI-P1-018).
    • Entity semantics: Similarly, entity-specific logic (vertex/edge type distinction, canonical ID resolution) belongs in EntityVectorIndex (GM-AI-P1-019).
    • Fusion and rerank: Score normalization, RRF, and reranking are GM-AI-P1-036 and GM-AI-P1-037.
    • Schema validation: Vertex/edge schema validation is GM-AI-P1-008.
    • Distributed storage: No multi-node replication, sharding, or consensus. Local-only for Phase 1.
    • Remote embedding calls: The VectorStore does not call any embedding service. It receives pre-computed vectors.

    2. Constraints

    2.1 Architectural Constraints

    C-1: VectorStore must not extend IndexStore

    // Current interface — source-specific
    public interface IndexStore {
        List<IVector> getEntityIndex(GraphEntity entity);
    }
    
    // New interface — source-agnostic
    public interface VectorStore {
        void upsert(VectorRecord record);
        List<VectorHit> search(VectorQuery query);
        // ...
    }

    These are different abstraction layers. IndexStore is source-specific (keyed by GraphEntity); VectorStore is source-agnostic (keyed by String vectorId). The adapter between them is EntityVectorIndex (GM-AI-P1-019), which implements IndexStore and delegates to VectorStore internally.

    C-2: Existing EmbeddingIndexStore must not be modified

    The EmbeddingIndexStore class, its tests (if any), and its public contract remain untouched. EntityVectorIndex wraps it as a compatibility layer. This ensures no regression in existing GraphMemoryServer behavior.

    C-3: All metadata fields must be validated at upsert time

    Dimension mismatch, missing model_name, and missing distance metric must be detected at write time (fail-closed), not at query time. This prevents corrupted stores that silently return garbage results.

    C-4: Vector IDs must be stable and deterministic

    Vector IDs must not depend on wall-clock time, insertion order, or JVM object identity. They must be reproducible across runs for the same input data. This is required by the idempotency contract (GM-AI-P1-011).

    2.2 Compatibility Constraints

    C-5: Backward compatible — no silent behavior changes

    The GraphMemoryServer.search() method currently uses instanceof checks to dispatch between EmbeddingOperator and SessionOperator. This dispatch logic must continue to work. The new VectorStore is an additional path, not a replacement.

    C-6: Must work offline without remote embedding credentials

    All contract tests and reference implementations must be runnable without any EmbeddingService connection. Fake embeddings (deterministic double[]) must be sufficient.

    C-7: Persistence format must be forward-compatible

    The format_version field is not decorative. On load, the LocalVectorStore must check it. Unknown versions must fail closed (throw exception), not skip silently.

    2.3 Performance Constraints

    C-8: Upsert and search must be in-memory operations

    No network IO or disk IO during the hot path. Disk writes are batched and async. This aligns with the existing EmbeddingIndexStore pattern (in-memory HashMap + periodic JSONL flush).

    C-9: Linear scan is acceptable for Phase 1

    The search() method uses brute-force cosine similarity. The performance contract is: O(N) per query where N is the number of stored vectors. ANN acceleration is deferred to GM-AI-P1-020.


    3. Current State Analysis

    3.1 Existing Index Architecture

    Class Path Role Issue
    IndexStore index/IndexStore.java Single-method interface: getEntityIndex(GraphEntity) No vector search, no metadata
    EmbeddingIndexStore index/EmbeddingIndexStore.java Embedding storage + index building + JSONL persistence Zero tests, tightly coupled
    EntityAttributeIndexStore index/EntityAttributeIndexStore.java Keyword-based index (verbalized text) Out of scope
    IndexStoreCache index/IndexStoreCache.java Static singleton aggregating stores Only registers EntityAttributeIndexStore
    IVector index/vector/IVector.java Vector interface with match() method Ad-hoc similarity, no metadata
    EmbeddingVector index/vector/EmbeddingVector.java Double-array vector, cosine similarity No dimension validation
    KeywordVector index/vector/KeywordVector.java String-array vector, overlap counting Asymmetric normalization

    3.2 How Indexes Are Used in GraphMemoryServer

    // Source: GraphMemoryServer.java, lines 87-96
    public List<GraphEntity> search(String query) {
        for (IndexStore store : indexStores) {
            if (store instanceof EntityAttributeIndexStore) {
                // keyword-based retrieval
                SessionOperator operator = new SessionOperator(query, (EntityAttributeIndexStore) store);
                // ...
            } else if (store instanceof EmbeddingIndexStore) {
                // embedding-based retrieval
                EmbeddingOperator operator = new EmbeddingOperator(query, (EmbeddingIndexStore) store);
                // ...
            }
        }
    }

    Key finding: The dispatch uses instanceof checks, not polymorphism. Adding a new index type requires modifying this method. The VectorStore SPI does not fix this dispatch problem (that is GM-AI-P1-019's job), but it provides the storage foundation that the new index types will use.

    3.3 Three Missing Capabilities

    1. Source-agnostic vector storage — current storage is keyed by GraphEntity, not by String vectorId
    2. Metadata validation — no dimension check, no model compatibility check, no format version check
    3. Pluggable persistence — JSONL is hardcoded; no way to swap in a different backend

    4. Design

    4.1 Overall Architecture

    ┌─────────────────────────────────────────────────────────────────┐
    │  Consumers (GM-AI-P1-018, 019, 020)                            │
    │  ┌──────────────────┐  ┌──────────────────┐  ┌───────────────┐  │
    │  │ ChunkVectorIndex  │  │ EntityVectorIndex │  │ ANN Adapter   │  │
    │  │ (chunk_id, src)   │  │ (entity_id, type) │  │ (future)      │  │
    │  └────────┬─────────┘  └────────┬─────────┘  └──────┬────────┘  │
    │           │                      │                    │           │
    ├───────────┼──────────────────────┼────────────────────┼───────────┤
    │           ▼                      ▼                    ▼           │
    │  ┌─────────────────────────────────────────────────────────────┐ │
    │  │                    VectorStore SPI                          │ │
    │  │  upsert(VectorRecord)  search(VectorQuery)  markDeleted()  │ │
    │  └──────────────────────────┬──────────────────────────────────┘ │
    │                             │                                    │
    ├─────────────────────────────┼────────────────────────────────────┤
    │           ┌─────────────────┼─────────────────┐                  │
    │           ▼                 ▼                 ▼                  │
    │  ┌────────────────┐ ┌────────────────┐ ┌──────────────────┐      │
    │  │ InMemoryVector │ │ LocalVector    │ │ Future: Remote   │      │
    │  │ Store          │ │ Store (JSONL)  │ │ VectorStore      │      │
    │  │ (testing)      │ │ (persistence)  │ │ (Phase 2+)       │      │
    │  └────────────────┘ └────────────────┘ └──────────────────┘      │
    └─────────────────────────────────────────────────────────────────┘
    

    4.2 Data Structures

    4.2.1 VectorRecord

    public class VectorRecord {
        /** Stable unique ID. Must not depend on wall-clock time. */
        private final String vectorId;
    
        /** The embedding vector. Length must match store dimension. */
        private final double[] embedding;
    
        /** Source discriminator: "chunk", "entity", or custom type. */
        private final String sourceType;
    
        /** The source identifier (chunk_id or entity_id). */
        private final String sourceId;
    
        /** Additional metadata (source_ref, text_hash, extractor_version, etc.). */
        private final Map<String, String> metadata;
    
        // Constructor validates: vectorId non-null, embedding non-null and non-empty,
        // sourceType non-null, sourceId non-null.
        // metadata may be null (defaults to empty map).
    }

    Design decision: sourceType and sourceId are top-level fields (not buried in metadata) because every consumer needs to know the source kind for dispatch. The metadata map holds optional, source-specific attributes.

    4.2.2 VectorQuery

    public class VectorQuery {
        /** The query embedding vector. */
        private final double[] queryVector;
    
        /** Maximum number of results to return. */
        private final int topK;
    
        /** Optional metadata filters (e.g., model_name, sourceType). */
        private final Map<String, String> filterMetadata;
    
        // Constructor validates: queryVector non-null, topK > 0.
    }

    4.2.3 VectorHit

    public class VectorHit {
        /** The matched vector's ID. */
        private final String vectorId;
    
        /** Similarity score (higher = more similar for cosine). */
        private final double score;
    
        /** The full record of the matched vector. */
        private final VectorRecord record;
    }

    4.2.4 VectorStoreMetadata

    Six required fields, each with a specific engineering rationale:

    Field Type Purpose Failure mode if missing/wrong
    model_name String Identifies the embedding model Vectors from different models live in incompatible semantic spaces; cross-model queries return meaningless results
    dimension int Vector dimensionality Dimension mismatch causes dot() out-of-bounds or silent corruption
    distance DistanceMetric enum COSINE, L2, DOT_PRODUCT Different metrics produce incomparable scores; fusion (GM-AI-P1-036) would produce wrong rankings
    index_version String Index format version Old format may not deserialize after code upgrade
    created_at long (epoch millis) Creation timestamp Cannot detect stale indexes or trigger re-indexing
    format_version int Serialization format version Unknown format silently corrupts data; must fail closed
    public enum DistanceMetric {
        COSINE,
        L2,
        DOT_PRODUCT
    }
    
    public class VectorStoreMetadata {
        private final String modelName;
        private final int dimension;
        private final DistanceMetric distance;
        private final String indexVersion;
        private final long createdAt;
        private final int formatVersion;
    
        // All fields are required (no defaults). Construction fails if any is null/invalid.
    }

    Design decision TD-1: All metadata fields are mandatory — Optional metadata leads to downstream consumers silently assuming defaults, which produces hard-to-debug incompatibilities. A store that does not know its own dimension cannot safely upsert vectors. The cost of requiring all fields is a slightly verbose constructor; the benefit is that every store is self-describing.

    4.2.5 VectorStoreException

    public class VectorStoreException extends RuntimeException {
        public enum ErrorCode {
            DIMENSION_MISMATCH,      // vector.length != store.dimension
            METADATA_INCOMPLETE,     // required metadata field missing
            MODEL_MISMATCH,          // query filter model_name != store model_name
            FORMAT_VERSION_UNKNOWN,  // unknown format_version on load
            RECORD_NOT_FOUND,        // vectorId does not exist
            PERSISTENCE_ERROR        // disk I/O failure
        }
    
        private final ErrorCode errorCode;
        private final String detail;
    }

    4.3 Core Interface

    public interface VectorStore {
    
        /**
         * Insert or update a single vector record.
         * Upsert is idempotent: same vectorId overwrites the previous record.
         *
         * @throws VectorStoreException(DIMENSION_MISMATCH) if record.embedding.length != metadata.dimension
         * @throws VectorStoreException(METADATA_INCOMPLETE) if required metadata fields are missing
         */
        void upsert(VectorRecord record);
    
        /**
         * Batch insert/update. Atomic per record; partial failure does not
         * roll back already-committed records.
         */
        void upsertBatch(List<VectorRecord> records);
    
        /**
         * Search for nearest vectors using the configured distance metric.
         * Results are ordered by score descending (best match first).
         *
         * @return empty list if no results match, never null
         * @throws VectorStoreException(MODEL_MISMATCH) if filterMetadata specifies
         *         a model_name different from the store's model_name
         */
        List<VectorHit> search(VectorQuery query);
    
        /**
         * Soft-delete a vector by id. The record remains on disk but is
         * excluded from future search results.
         *
         * @throws VectorStoreException(RECORD_NOT_FOUND) if vectorId does not exist
         */
        void markDeleted(String vectorId);
    
        /**
         * Returns metadata about this store instance.
         */
        VectorStoreMetadata getMetadata();
    
        /**
         * Release resources. Implementations must flush pending writes before returning.
         */
        void close();
    }

    4.4 Reference Implementations

    4.4.1 InMemoryVectorStore

    Primary purpose: contract validation and unit testing.

    public class InMemoryVectorStore implements VectorStore {
        private final VectorStoreMetadata metadata;
        private final ConcurrentHashMap<String, VectorRecord> store = new ConcurrentHashMap<>();
        private final Set<String> deletedIds = ConcurrentHashMap.newKeySet();
    
        // search(): brute-force linear scan
        // - Compute cosine similarity against all non-deleted records
        // - Sort by score descending, return topK
        // - O(N) per query, acceptable for Phase 1
    
        // upsert(): validate dimension, then store.put(vectorId, record)
        // markDeleted(): deletedIds.add(vectorId)
    }
    • Not persistent; resets on process restart.
    • Suitable for all contract tests that do not require persistence.
    • The search() implementation is the reference for cosine similarity scoring.

    4.4.2 LocalVectorStore

    Primary purpose: local persistent storage, replacing the current EmbeddingIndexStore JSONL approach.

    Storage format: append-only JSONL with per-record checksum
    
    Line format:
      {"vectorId":"...", "embedding":[...], "sourceType":"...", "sourceId":"...",
       "metadata":{...}, "__checksum":"...", "__deleted":false}
    
    On startup:
      1. Read all lines from JSONL file
      2. Validate checksum for each line
      3. Lines with checksum mismatch → quarantine (write to .quarantine file, log warning)
      4. Truncated lines (invalid JSON) → quarantine
      5. Build in-memory index from valid lines
      6. Check format_version: unknown version → throw VectorStoreException(FORMAT_VERSION_UNKNOWN)
    
    On upsert:
      1. Validate dimension and metadata
      2. Append new line to JSONL (with checksum)
      3. Update in-memory index
      4. If same vectorId existed: the old line is NOT modified (append-only).
         The in-memory index always points to the latest line for a given vectorId.
    
    On markDeleted:
      1. Append a delete marker line: {"vectorId":"...", "__deleted":true, "__checksum":"..."}
      2. Update in-memory index (remove from searchable set)
    

    Design decision TD-2: Append-only with logical delete — In-place modification of JSONL is fragile (partial writes corrupt the file). Append-only with a delete marker ensures that crash recovery is always possible: the last occurrence of a vectorId wins. The cost is that the file grows over time; a compaction step can be added in Phase 2.

    4.5 Distance Metric Implementation

    public class DistanceUtils {
    
        /**
         * Compute distance between two vectors according to the specified metric.
         * Returns a similarity score where higher = more similar.
         *
         * @throws VectorStoreException(DIMENSION_MISMATCH) if vectors have different lengths
         */
        public static double compute(double[] a, double[] b, DistanceMetric metric) {
            if (a.length != b.length) {
                throw new VectorStoreException(ErrorCode.DIMENSION_MISMATCH,
                    "Vector lengths differ: " + a.length + " vs " + b.length);
            }
            switch (metric) {
                case COSINE: return cosineSimilarity(a, b);
                case L2: return 1.0 / (1.0 + euclideanDistance(a, b));  // invert so higher = better
                case DOT_PRODUCT: return dotProduct(a, b);
                default: throw new IllegalArgumentException("Unknown metric: " + metric);
            }
        }
    
        private static double cosineSimilarity(double[] a, double[] b) {
            double dot = 0.0, normA = 0.0, normB = 0.0;
            for (int i = 0; i < a.length; i++) {
                dot += a[i] * b[i];
                normA += a[i] * a[i];
                normB += b[i] * b[i];
            }
            if (normA == 0.0 || normB == 0.0) return 0.0;
            return dot / (Math.sqrt(normA) * Math.sqrt(normB));
        }
    }
  5. Sumit2342 commented on Aug 27, 2026

    @Sumit2342

    Thanks @kitalkuyo-gita for detailed design document.

    I have already started refactoring my implementation to match this exact blueprint.
    I will push the updated code to my PR for your review.

    Thanks again for the review.

  6. Sumit2342 commented on Aug 28, 2026

    @Sumit2342

    Hi @kitalkuyo-gita ,
    I have updated my PR. Can you please review it again and let me know if any other changes or enhancement is required.
    Thanks.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions