Skip to content

About

A Vector Database built from scratch using Spring Boot featuring HNSW, KD-Tree, Brute Force Search, RAG, and Ollama integration.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

VectorDB — Build a Vector Database from Scratch using Spring Boot

A fully working Vector Database built from scratch using Spring Boot with a modern web UI.

Implements HNSW, KD-Tree, and Brute Force search algorithms side-by-side, plus a complete Retrieval-Augmented Generation (RAG) pipeline powered by Ollama.

Built as an educational project to demonstrate how production vector databases like Pinecone, Weaviate, Chroma, and Milvus work internally.


What This Project Does

Feature Description
3 Search Algorithms HNSW (production-grade), KD-Tree, Brute Force
3 Distance Metrics Cosine Similarity, Euclidean Distance, Manhattan Distance
16D Demo Vectors Interactive demo vectors across multiple categories
768D Real Embeddings Ollama's nomic-embed-text model generates production-grade embeddings
2D PCA Visualization Interactive semantic space visualization
Document Chunking Long documents automatically split into overlapping chunks
RAG Pipeline Ask questions over your own documents using local AI
Spring Boot REST API Complete REST backend with CRUD, search, document management, and AI endpoints

How It Works

Your Text
     │
     ▼
Spring Boot Backend
     │
     ▼
Ollama (nomic-embed-text)
     │
     ▼
768-Dimensional Embedding
     │
     ▼
Vector Database
(HNSW / KDTree / Brute Force)
     │
     ▼
Nearest Chunks Retrieved
     │
     ▼
Ollama (llama3.2)
     │
     ▼
Answer

Technologies Used

  • Java
  • Spring Boot
  • Spring MVC
  • Maven
  • Ollama
  • nomic-embed-text
  • llama3.2
  • HTML
  • CSS
  • JavaScript
  • Canvas API
  • REST APIs

Prerequisites

  • Java 21 (or your project version)
  • Maven
  • Git
  • Ollama

Verify:

java -version
mvn -version
git --version

Install Ollama

ollama pull nomic-embed-text
ollama pull llama3.2
ollama list

Clone Repository

git clone https://github.com/YOUR_USERNAME/VectorDB.git
cd VectorDB

Configure Application

application.yml

server:
  port: 8081

ollama:
  url: http://localhost:11434
  embed-model: nomic-embed-text
  model: llama3.2

Run Application

Start Ollama

ollama serve

Run Spring Boot

mvn spring-boot:run

Open

http://localhost:8081

Using the Application

Tab 1: Search (Demo Vectors)

  • Type any concept in the search box: binary tree, sushi, basketball, calculus
  • Choose your algorithm: HNSW, KD-Tree, or Brute Force
  • Choose distance metric: Cosine, Euclidean, or Manhattan
  • Click ⚡ SEARCH — results appear with distances, the matching point glows on the scatter plot
  • Click ▶ COMPARE ALL ALGOS to run all 3 algorithms and compare their speed

The scatter plot shows all 20 vectors projected to 2D using PCA. Notice how the 4 semantic categories (CS, Math, Food, Sports) form distinct clusters — this is what "semantic similarity" looks like visually.

Tab 2: Documents (Real Embeddings)

This uses Ollama to generate real 768-dimensional embeddings from any text.

  1. Type a title (e.g., Operating Systems Notes)
  2. Paste any text — lecture notes, textbook paragraphs, Wikipedia articles
  3. Click ⚡ EMBED & INSERT
  4. Long documents are automatically split into overlapping 250-word chunks
  5. Each chunk gets its own embedding and is stored in a separate HNSW index

Tab 3: Ask AI (RAG Pipeline)

  1. Make sure you have inserted some documents in Tab 2 first
  2. Type a question about your documents
  3. Click 🤖 ASK AI

What happens behind the scenes:

1. Your question → embedded with nomic-embed-text (768D vector)
2. HNSW search → finds 3 most semantically similar chunks
3. Retrieved chunks → sent as context to llama3.2
4. llama3.2 → generates an answer based only on your documents

The answer streams in with a typewriter effect. Click the context chips to see exactly which chunks the AI used.


Documents

  • Upload documents
  • Automatic chunking
  • 768D embeddings
  • Stored inside VectorDB

Ask AI

Pipeline:

Question
 ↓
Embedding
 ↓
HNSW Retrieval
 ↓
Top-K Context
 ↓
Prompt
 ↓
llama3.2
 ↓
Answer

REST APIs

Vector APIs

Method Endpoint
GET /vector/items
POST /vector/insert
DELETE /vector/delete/{id}
POST /search
GET /hnsw/info

RAG APIs

Method Endpoint
POST /rag/insert
GET /rag/list
DELETE /rag/delete/{id}
POST /rag/search
POST /rag/ask
GET /rag/status

Project Structure

src
└── main
    ├── java
    │   ├── algorithm
    │   ├── config
    │   ├── controller
    │   ├── dto
    │   ├── model
    │   ├── rag
    │   ├── repository
    │   ├── service
    │   └── VectorApplication.java
    └── resources
        ├── static
        │   └── index.html
        └── application.yml

pom.xml
README.md

Architecture

Browser
   ↓
Spring Controllers
   ↓
Service Layer
   ↓
Embedding Service
   ↓
Vector Repository
   ↓
Search Algorithms
 ├── HNSW
 ├── KD Tree
 └── Brute Force
   ↓
RAG Service
   ↓
Ollama

Algorithms Deep Dive

HNSW (Hierarchical Navigable Small World)

Nodes are inserted into a multilayer graph. Each node randomly gets assigned a maximum layer. Layer 0 has all nodes with many connections; higher layers have fewer nodes (exponentially fewer) with longer-range connections.

Insert: Start at the top layer, greedily find the nearest node, drop a layer, repeat. At each layer from your assigned max down to 0, run a beam search (ef_construction=200) and connect to the M nearest neighbors bidirectionally.

Search: Same greedy descent from top layer. At layer 0, expand to ef nearest candidates using a priority queue.

Why it's fast: The upper layers act like a highway — you quickly get to the right neighborhood, then zoom in at layer 0.

KD-Tree (K-Dimensional Tree)

Binary space partitioning. Each node splits space along one dimension (cycling through all dimensions). Search prunes entire subtrees when the closest possible point in that subtree can't beat the current best — the "ball within hyperslab" check.

Weakness: Degrades with high dimensions (curse of dimensionality). Works well for ≤20D, becomes close to brute force at 768D.

Why HNSW Wins at High Dimensions

KD-Tree pruning relies on axis-aligned distance bounds. In high dimensions, almost all the space is near the boundary of the hypersphere — no subtrees get pruned. HNSW's graph-based approach doesn't have this problem.


Common Issues

Problem Fix
Ollama offline Run ollama serve
Missing embedding model ollama pull nomic-embed-text
Missing LLM ollama pull llama3.2
Port conflict Change server.port
Maven build issue Verify Java & Maven installation

Use a Smaller/Faster LLM

If llama3.2 is too slow on your laptop, you can switch to the smaller 1B model:

ollama pull llama3.2:1b

Then update your application.yml (or application.properties) to use the new model:

application.yml

ollama:
  model: llama3.2:1b

Or if you're using application.properties:

ollama.model=llama3.2:1b

Restart the Spring Boot application after making the change.


Future Enhancements

  • Hybrid Search
  • Metadata Filtering
  • Persistent Database
  • Authentication
  • PDF/DOCX Support
  • Streaming Responses

License

MIT — use this however you want.

About

A Vector Database built from scratch using Spring Boot featuring HNSW, KD-Tree, Brute Force Search, RAG, and Ollama integration.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages