A fully working Vector Database built from scratch using Spring Boot with a modern web UI.
Implements HNSW, KD-Tree, and Brute Force search algorithms side-by-side, plus a complete Retrieval-Augmented Generation (RAG) pipeline powered by Ollama.
Built as an educational project to demonstrate how production vector databases like Pinecone, Weaviate, Chroma, and Milvus work internally.
| Feature | Description |
|---|---|
| 3 Search Algorithms | HNSW (production-grade), KD-Tree, Brute Force |
| 3 Distance Metrics | Cosine Similarity, Euclidean Distance, Manhattan Distance |
| 16D Demo Vectors | Interactive demo vectors across multiple categories |
| 768D Real Embeddings | Ollama's nomic-embed-text model generates production-grade embeddings |
| 2D PCA Visualization | Interactive semantic space visualization |
| Document Chunking | Long documents automatically split into overlapping chunks |
| RAG Pipeline | Ask questions over your own documents using local AI |
| Spring Boot REST API | Complete REST backend with CRUD, search, document management, and AI endpoints |
Your Text
│
▼
Spring Boot Backend
│
▼
Ollama (nomic-embed-text)
│
▼
768-Dimensional Embedding
│
▼
Vector Database
(HNSW / KDTree / Brute Force)
│
▼
Nearest Chunks Retrieved
│
▼
Ollama (llama3.2)
│
▼
Answer
- Java
- Spring Boot
- Spring MVC
- Maven
- Ollama
- nomic-embed-text
- llama3.2
- HTML
- CSS
- JavaScript
- Canvas API
- REST APIs
- Java 21 (or your project version)
- Maven
- Git
- Ollama
Verify:
java -version
mvn -version
git --versionollama pull nomic-embed-text
ollama pull llama3.2
ollama listgit clone https://github.com/YOUR_USERNAME/VectorDB.git
cd VectorDBapplication.yml
server:
port: 8081
ollama:
url: http://localhost:11434
embed-model: nomic-embed-text
model: llama3.2Start Ollama
ollama serveRun Spring Boot
mvn spring-boot:runOpen
http://localhost:8081
- Type any concept in the search box:
binary tree,sushi,basketball,calculus - Choose your algorithm: HNSW, KD-Tree, or Brute Force
- Choose distance metric: Cosine, Euclidean, or Manhattan
- Click ⚡ SEARCH — results appear with distances, the matching point glows on the scatter plot
- Click ▶ COMPARE ALL ALGOS to run all 3 algorithms and compare their speed
The scatter plot shows all 20 vectors projected to 2D using PCA. Notice how the 4 semantic categories (CS, Math, Food, Sports) form distinct clusters — this is what "semantic similarity" looks like visually.
This uses Ollama to generate real 768-dimensional embeddings from any text.
- Type a title (e.g.,
Operating Systems Notes) - Paste any text — lecture notes, textbook paragraphs, Wikipedia articles
- Click ⚡ EMBED & INSERT
- Long documents are automatically split into overlapping 250-word chunks
- Each chunk gets its own embedding and is stored in a separate HNSW index
- Make sure you have inserted some documents in Tab 2 first
- Type a question about your documents
- Click 🤖 ASK AI
What happens behind the scenes:
1. Your question → embedded with nomic-embed-text (768D vector)
2. HNSW search → finds 3 most semantically similar chunks
3. Retrieved chunks → sent as context to llama3.2
4. llama3.2 → generates an answer based only on your documents
The answer streams in with a typewriter effect. Click the context chips to see exactly which chunks the AI used.
- Upload documents
- Automatic chunking
- 768D embeddings
- Stored inside VectorDB
Pipeline:
Question
↓
Embedding
↓
HNSW Retrieval
↓
Top-K Context
↓
Prompt
↓
llama3.2
↓
Answer
| Method | Endpoint |
|---|---|
| GET | /vector/items |
| POST | /vector/insert |
| DELETE | /vector/delete/{id} |
| POST | /search |
| GET | /hnsw/info |
| Method | Endpoint |
|---|---|
| POST | /rag/insert |
| GET | /rag/list |
| DELETE | /rag/delete/{id} |
| POST | /rag/search |
| POST | /rag/ask |
| GET | /rag/status |
src
└── main
├── java
│ ├── algorithm
│ ├── config
│ ├── controller
│ ├── dto
│ ├── model
│ ├── rag
│ ├── repository
│ ├── service
│ └── VectorApplication.java
└── resources
├── static
│ └── index.html
└── application.yml
pom.xml
README.md
Browser
↓
Spring Controllers
↓
Service Layer
↓
Embedding Service
↓
Vector Repository
↓
Search Algorithms
├── HNSW
├── KD Tree
└── Brute Force
↓
RAG Service
↓
Ollama
Nodes are inserted into a multilayer graph. Each node randomly gets assigned a maximum layer. Layer 0 has all nodes with many connections; higher layers have fewer nodes (exponentially fewer) with longer-range connections.
Insert: Start at the top layer, greedily find the nearest node, drop a layer, repeat. At each layer from your assigned max down to 0, run a beam search (ef_construction=200) and connect to the M nearest neighbors bidirectionally.
Search: Same greedy descent from top layer. At layer 0, expand to ef nearest candidates using a priority queue.
Why it's fast: The upper layers act like a highway — you quickly get to the right neighborhood, then zoom in at layer 0.
Binary space partitioning. Each node splits space along one dimension (cycling through all dimensions). Search prunes entire subtrees when the closest possible point in that subtree can't beat the current best — the "ball within hyperslab" check.
Weakness: Degrades with high dimensions (curse of dimensionality). Works well for ≤20D, becomes close to brute force at 768D.
KD-Tree pruning relies on axis-aligned distance bounds. In high dimensions, almost all the space is near the boundary of the hypersphere — no subtrees get pruned. HNSW's graph-based approach doesn't have this problem.
| Problem | Fix |
|---|---|
| Ollama offline | Run ollama serve |
| Missing embedding model | ollama pull nomic-embed-text |
| Missing LLM | ollama pull llama3.2 |
| Port conflict | Change server.port |
| Maven build issue | Verify Java & Maven installation |
If llama3.2 is too slow on your laptop, you can switch to the smaller 1B model:
ollama pull llama3.2:1bThen update your application.yml (or application.properties) to use the new model:
application.yml
ollama:
model: llama3.2:1bOr if you're using application.properties:
ollama.model=llama3.2:1bRestart the Spring Boot application after making the change.
- Hybrid Search
- Metadata Filtering
- Persistent Database
- Authentication
- PDF/DOCX Support
- Streaming Responses
MIT — use this however you want.