____ ____ ____ _____ ____ _ _
/ ___| _ \/ ___|| ____| | __ ) ___ _ __ ___| |__ _ __ ___ __ _ _ __| | _____
| | _| | | \___ \| _| _____ | _ \ / _ \ '_ \ / __| '_ \| '_ ` _ \ / _` | '__| |/ / __|
| |_| | |_| |___) | |__|_____|| |_) | __/ | | | (__| | | | | | | | | (_| | | | <\__ \
\____|____/|____/|_____| |____/ \___|_| |_|\___|_| |_|_| |_| |_|\__,_|_| |_|\_\___/
🌐 Live Visualizations: https://wittenyeh.github.io/gdse-benchmarks/
An embedded graph database benchmark framework for testing Neo4j, JanusGraph, OrientDB, Sqlg, ArangoDB, and Aster latency performance using native APIs instead of query engines.
╔════════════════════════════════════════════════════════════════════════════╗
║ HOST (Python) ║
╠════════════════════════════════════════════════════════════════════════════╣
║ ║
║ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ║
║ │ Benchmark │──▶│ Workload │──▶│ Docker │ ║
║ │ Launcher │ │ Compiler │ │ Manager │ ║
║ └──────────────┘ └──────────────┘ └──────┬───────┘ ║
║ │ ║
║ ┌──────────────┐ ┌──────────────┐ │ ║
║ │ Report │ │ Dataset │ │ ║
║ │ Generator │ │ Loader │ │ ║
║ └──────────────┘ └──────────────┘ │ ║
╚═══════════════════════════════════════════════│════════════════════════════╝
│
HTTP Requests (Ports 50080 - 50085)
┌──────────────────────────────────────┴───────────────┐
▼ ▼
╔══════════════════════════════════════╗ ╔═══════════════════════════════════╗
║ DOCKER TARGETS (Java API) ║ ║ DOCKER TARGETS (C++ API) ║
╠══════════════════════════════════════╣ ╠═══════════════════════════════════╣
║ ┌──────────────────────────────────┐ ║ ║ ┌───────────────────────────────┐ ║
║ │ Neo4j Embedded │ ║ ║ │ ArangoDB │ ║
║ │ ├─ BenchmarkServer (:50080) │ ║ ║ │ ├─ BenchmarkServer (:50082) │ ║
║ │ └─ Neo4jBenchmarkExecutor │ ║ ║ │ └─ ArangoDBBenchmarkExecutor │ ║
║ └──────────────────────────────────┘ ║ ║ └───────────────────────────────┘ ║
║ ┌──────────────────────────────────┐ ║ ║ ┌───────────────────────────────┐ ║
║ │ JanusGraph+BDB │ ║ ║ │ Aster │ ║
║ │ ├─ BenchmarkServer (:50081) │ ║ ║ │ ├─ BenchmarkServer (:50085) │ ║
║ │ └─ JanusGraphBenchmarkExecutor │ ║ ║ │ └─ AsterBenchmarkExecutor │ ║
║ └──────────────────────────────────┘ ║ ║ └───────────────────────────────┘ ║
║ ┌──────────────────────────────────┐ ║ ╚═══════════════════════════════════╝
║ │ OrientDB │ ║
║ │ ├─ BenchmarkServer (:50083) │ ║
║ │ └─ OrientDBBenchmarkExecutor │ ║
║ └──────────────────────────────────┘ ║
║ ┌──────────────────────────────────┐ ║
║ │ Sqlg │ ║
║ │ ├─ BenchmarkServer (:50084) │ ║
║ │ └─ SqlgBenchmarkExecutor │ ║
║ └──────────────────────────────────┘ ║
╚══════════════════════════════════════╝
╔════════════════════════════════════════════════════════════════════════════╗
║ COMMON LIBRARIES (Dependencies) ║
╠══════════════════════════════════════╦═════════════════════════════════════╣
║ [ common-java ] ║ [ common-cpp ] ║
║ ├─ BenchmarkExecutor<T> ║ ├─ BenchmarkExecutor<Derived> ║
║ │ (Abstract base class) ║ │ (CRTP base class) ║
║ ├─ NodeIdMapping<T> ║ ├─ NodeIdMapping<T> ║
║ │ (Array-based ID mapping) ║ │ (Vector-based ID mapping) ║
║ ├─ CsvGraphReader ║ ├─ CsvGraphReader ║
║ │ (CSV parsing) ║ │ (CSV parsing) ║
║ ├─ ProgressCallback ║ ├─ ProgressCallback ║
║ │ (Progress reporting) ║ │ (Progress reporting) ║
║ └─ TypeConverter ║ └─ BenchmarkUtils ║
║ (Type conversion utilities) ║ (Utility functions) ║
╚══════════════════════════════════════╩═════════════════════════════════════╝
The benchmark framework consists of two main components:
- BenchmarkLauncher: Orchestrates the entire benchmark workflow
- WorkloadCompiler: Generates native API workload JSON files with structured parameters
- DockerManager: Manages Docker containers and communicates with benchmark servers
- ReportGenerator: Collects benchmark results and saves to JSON files
- BenchmarkServer: HTTP API server that receives execution requests and returns metrics
- WorkloadDispatcher: Reads JSON workload files and dispatches operations to executor
- BenchmarkExecutor: Executes operations serially using native APIs, measures per-operation latency
- Each container is isolated and provides clean benchmarking environment
- Execution time is pure (excludes network transmission time)
- Native API Execution: Direct database API calls (Neo4j Embedded API, TinkerPop Structure API, ArangoDB Fuerte C++ API) instead of query engines;
- Latency Testing: Serial execution with per-operation latency tracking in microseconds;
- 10 Benchmark Operations: ADD_VERTEX, UPSERT_VERTEX_PROPERTY, REMOVE_VERTEX, ADD_EDGE, UPSERT_EDGE_PROPERTY, REMOVE_EDGE, GET_NBRS, GET_VERTEX_BY_PROPERTY, GET_EDGE_BY_PROPERTY, LOAD_GRAPH;
- Support for Neo4j Embedded, JanusGraph, ArangoDB, OrientDB, Sqlg, and Aster with native APIs;
- Multi-Language Support: Java (Neo4j, JanusGraph) and C++ (ArangoDB) implementations;
- Docker-based isolation for clean benchmarking and reproducible results;
- Interactive HTML visualizations with Plotly;
- Python 3.8+
- Docker
- Java 21 (for Neo4j), Java 17 (for JanusGraph)
# Clone the repository
git clone <repository-url>
cd graph-database-benchmark
# Initialize submodules (for datasets)
git submodule update --init --recursive
# Install Python dependencies
pip install -r requirements.txt
# Build Docker images
./build.shThe easiest way to run benchmarks using named arguments:
# Single database, single dataset
./run_benchmark.sh --database neo4j --dataset coAuthorsDBLP
# Single database, multiple datasets
./run_benchmark.sh --database neo4j --dataset coAuthorsDBLP,delaunay_n13
# Multiple databases, single dataset
./run_benchmark.sh --database neo4j,janusgraph --dataset coAuthorsDBLP
# Multiple databases, multiple datasets (runs all combinations)
./run_benchmark.sh --database neo4j,janusgraph --dataset coAuthorsDBLP,delaunay_n13
# Custom workload configuration
./run_benchmark.sh --database neo4j --dataset coAuthorsDBLP --workload workloads/templates/quick_test.json
# Short form options
./run_benchmark.sh -d neo4j,janusgraph -s coAuthorsDBLP,delaunay_n13
# Show help
./run_benchmark.sh --helpLegacy positional arguments are still supported:
./run_benchmark.sh neo4j coAuthorsDBLP
./run_benchmark.sh neo4j coAuthorsDBLP workloads/templates/quick_test.jsonFor more control, use the Python launcher directly:
# Basic usage
python host/benchmark_launcher.py \
--database-name neo4j \
--dataset-name coAuthorsDBLP \
--workload-config workloads/templates/example_workload.json \
--output-dir reports
# Test multiple datasets
python host/benchmark_launcher.py \
--database-name janusgraph \
--dataset-name coAuthorsDBLP delaunay_n13 \
--output-dir reports
# With reproducible seed
python host/benchmark_launcher.py \
--database-name neo4j \
--dataset-name coAuthorsDBLP \
--seed 42 \
--output-dir reports
# Force rebuild Docker image
python host/benchmark_launcher.py \
--database-name neo4j \
--dataset-name coAuthorsDBLP \
--rebuild| Argument | Default | Description |
|---|---|---|
--database-name |
(required) | Database to benchmark: neo4j, janusgraph, or arangodb |
--database-config |
config/database-config.json |
Path to database configuration file |
--dataset-config |
config/datasets.json |
Path to dataset configuration file |
--dataset-name |
(required) | Dataset name(s) to test (can specify multiple) |
--workload-config |
workloads/templates/example_workload.json |
Path to workload configuration file |
--workload-name |
(all workloads) | Specific workload name to run |
--output-dir |
reports/ |
Output directory for benchmark reports |
--seed |
(none) | Random seed for reproducibility |
--rebuild |
off | Force rebuild the Docker image |
The benchmark framework provides visualization tools to generate interactive HTML plots from benchmark reports.
Compare how different databases perform with varying batch sizes for each task:
./visualize.sh batchsize \
--database neo4j janusgraph arangodb aster \
--workload batchsize_comparison \
--dataset delaunay_n13 \
--output-dir plotsThis generates one interactive plot per task (e.g., ADD_VERTEX, ADD_EDGE) showing how latency changes with batch size across databases.
Output files:
plots/batchsize_delaunay_n13_ADD_VERTEX.htmlplots/batchsize_delaunay_n13_ADD_EDGE.html- etc.
Compare overall performance across databases for multiple datasets:
./visualize.sh performance \
--database neo4j janusgraph arangodb aster \
--workload large_structural_workload \
--dataset delaunay_n13 movielens-small \
--output-dir plotsThis generates one interactive grouped bar chart per dataset comparing average latency across all tasks.
The visualize/ directory contains Python scripts for generating plots:
plot_batchsize_comparison.py: Batch size comparison plotsplot_performance_comparison.py: Performance comparison plotsexamples.sh: Example commands and usage demonstrations
All plots are interactive HTML files using Plotly, allowing you to zoom, pan, and hover for detailed information.
The deploy/ directory contains scripts for deploying visualizations to GitHub Pages:
# Generate index page and deploy to GitHub Pages
./deploy/deploy.sh
# Deploy to a custom subdirectory
./deploy/deploy.sh my-benchmarksThe deployment script will:
- Generate a beautiful index page with all plots
- Clone your GitHub Pages repository
- Copy all HTML files to the specified subdirectory
- Commit and push changes
Your visualizations will be available at: https://wittenyeh.github.io/gdse-benchmarks/
See deploy/README.md for detailed documentation.
Maps dataset names to their .mtx file paths:
{
"root_dir": "./graph-datasets/",
"datasets": {
"coAuthorsDBLP": "coAuthorsDBLP/coAuthorsDBLP.mtx",
"delaunay_n13": "delaunay_n13/delaunay_n13.mtx",
"cit-Patents": "cit-Patents/cit-Patents.mtx"
}
}Defines database-specific settings:
{
"neo4j": {
"docker_image": "bench-neo4j",
"dockerfile_path": "./docker/neo4j/Dockerfile",
"container_name": "neo4j-benchmark",
"api_port": 8080,
"query_language": "cypher",
"runtime": "java",
"config": {
"heap_size": "4G",
"page_cache": "2G"
}
},
"janusgraph": {
"docker_image": "bench-janusgraph",
"dockerfile_path": "./docker/janusgraph/Dockerfile",
"container_name": "janusgraph-benchmark",
"api_port": 8081,
"query_language": "gremlin",
"runtime": "java",
"config": {
"heap_size": "4G",
"storage_backend": "berkeleyje"
}
}
}Example workload with multiple tasks:
{
"tasks": [
{ "name": "load_graph" },
{ "name": "add_vertex", "ops": 5000 },
{ "name": "upsert_vertex_property", "ops": 2000 },
{ "name": "remove_vertex", "ops": 1000 },
{ "name": "add_edge", "ops": 5000 },
{ "name": "upsert_edge_property", "ops": 2000 },
{ "name": "remove_edge", "ops": 1000 },
{ "name": "get_nbrs", "ops": 10000, "direction": "OUT" },
{ "name": "get_vertex_by_property", "ops": 1000 },
{ "name": "get_edge_by_property", "ops": 1000 }
]
}| Parameter | Type | Default | Description |
|---|---|---|---|
name |
string | required | Task name (see Supported Operations below) |
ops |
integer | required | Number of operations to execute |
direction |
string | "OUT" | Direction for GET_NBRS: "OUT", "IN", or "BOTH" |
The benchmark supports 10 operations using native database APIs:
- Description: Bulk-loads the entire dataset from MTX file
- Implementation: Two-pass streaming approach
- Pass 1: Collect unique node IDs
- Pass 2: Stream through file again to create edges
- Optimization: Batch commits every 10,000 operations, schema index creation
- Error Handling: Fails if dataset file is invalid or inaccessible
- Description: Adds new vertices using native API
- Neo4j API:
tx.createNode(Label.label("MyNode")).setProperty("id", vertexId) - TinkerPop API:
g.addV("MyNode").property("id", vertexId) - Parameters: List of vertex IDs
- Error Handling: Silent failure if vertex already exists (idempotent)
- Description: Updates or inserts vertex properties using native API
- Neo4j API:
node.setProperty(key, value) - TinkerPop API:
vertex.property(key, value) - Parameters: List of vertex updates (id + properties map)
- Error Handling: Silent failure if vertex doesn't exist
- Description: Removes vertices using native API
- Neo4j API:
node.delete()(after deleting relationships) - TinkerPop API:
g.V(vertexId).drop() - Parameters: List of vertex IDs
- Error Handling: Silent failure if vertex doesn't exist (idempotent)
- Description: Adds edges between vertices using native API
- Neo4j API:
srcNode.createRelationshipTo(dstNode, RelationshipType.withName(label)) - TinkerPop API:
g.V(srcId).addE(label).to(g.V(dstId)) - Parameters: Edge label and list of (src, dst) pairs
- Error Handling: Silent failure if source or destination vertex doesn't exist
- Description: Updates or inserts edge properties using native API
- Neo4j API:
relationship.setProperty(key, value) - TinkerPop API:
edge.property(key, value) - Parameters: Edge label and list of edge updates (src, dst, properties map)
- Error Handling: Silent failure if edge doesn't exist
- Description: Removes edges using native API
- Neo4j API:
relationship.delete() - TinkerPop API:
g.V(srcId).outE(label).where(inV().hasId(dstId)).drop() - Parameters: Edge label and list of (src, dst) pairs
- Error Handling: Silent failure if edge doesn't exist (idempotent)
- Description: Gets neighbors of vertices using native API
- Neo4j API:
node.getRelationships(direction)thenrel.getOtherNode(node) - TinkerPop API:
g.V(vertexId).out()/in()/both() - Parameters: Direction ("OUT", "IN", "BOTH") and list of vertex IDs
- Error Handling: Returns empty result if vertex doesn't exist
- Description: Queries vertices by property using native API
- Neo4j API: Iterate through
tx.getAllNodes()and filter by property - TinkerPop API:
g.V().hasLabel(label).has(key, value) - Parameters: List of property queries (key, value pairs)
- Error Handling: Returns empty result if no matches found
- Description: Queries edges by property using native API
- Neo4j API: Iterate through relationships and filter by property
- TinkerPop API:
g.E().hasLabel(label).has(key, value) - Parameters: Edge label and list of property queries (key, value pairs)
- Error Handling: Returns empty result if no matches found
Workloads are defined as JSON files with structured parameters:
{
"task_type": "ADD_VERTEX",
"ops_count": 3,
"parameters": {
"ids": [10001, 10002, 10003]
}
}{
"task_type": "GET_NBRS",
"ops_count": 3,
"parameters": {
"direction": "OUT",
"ids": [10001, 10002, 10005]
}
}The benchmark framework follows a silent failure approach:
- Add operations: Skip if entity already exists (idempotent)
- Delete operations: No-op if entity doesn't exist (idempotent)
- Read operations: Return empty results if entity doesn't exist
This design ensures:
- Consistent metrics: Failed operations don't skew latency measurements
- Realistic workloads: Real-world systems often handle missing entities gracefully
- Continuous execution: Benchmark doesn't stop on individual operation failures
Benchmark results are saved as JSON files with the naming pattern: bench_{database}_{dataset}.json
Example output:
{
"metadata": {
"database": "neo4j",
"dataset": "delaunay_n13",
"datasetPath": "./graph-datasets/delaunay_n13/delaunay_n13.mtx",
"timestamp": "2026-02-12T09:21:04.532605396Z",
"serverThreads": 8,
"javaVersion": "17.0.18",
"osName": "Linux",
"osVersion": "5.19.0-1010-nvidia-lowlatency"
},
"results": [
{
"task": "load_graph",
"status": "success",
"durationSeconds": 1.356723384,
"totalOps": 32739,
"clientThreads": 0
},
{
"task": "add_nodes_latency",
"status": "success",
"durationSeconds": 15.26413909,
"totalOps": 5000,
"clientThreads": 1,
"latency": {
"minUs": 1309.674,
"maxUs": 1121683.218,
"meanUs": 3050.2426564,
"medianUs": 2310.056,
"p50Us": 2310.056,
"p90Us": 5032.378,
"p95Us": 6839.976,
"p99Us": 10948.379
}
},
{
"task": "read_nbrs_throughput",
"status": "success",
"durationSeconds": 8.234567,
"totalOps": 50000,
"clientThreads": 16,
"throughputQps": 6071.23
}
]
}Latency Metrics (for _latency tasks):
minUs,maxUs,meanUs: Minimum, maximum, and mean latency in microsecondsmedianUs,p50Us: Median (50th percentile) latencyp90Us,p95Us,p99Us: 90th, 95th, and 99th percentile latencies- Calculated as:
batch_execution_time / batch_size
Throughput Metrics (for _throughput tasks):
throughputQps: Queries per second- Calculated as:
totalOps / durationSeconds - Measures the rate of operations completed during parallel execution
- Parameter Configuration: BenchmarkLauncher reads database, dataset, and workload configurations
- Workload Compilation: WorkloadCompiler generates native API workload JSON files with structured parameters (no query strings)
- Docker Container Startup: DockerManager builds/starts Docker container with mounted dataset and compiled workload directories
- Benchmark Execution: Container's WorkloadDispatcher:
- Reads JSON workload files
- Dispatches operations to BenchmarkExecutor
- Executor calls native APIs serially, measures per-operation latency
- Measures pure execution time (no network overhead)
- Returns results via HTTP API
- Result Collection: Host collects results and saves to JSON file
- Visualization: Use
visualize.shto generate interactive plots from benchmark reports
Neo4j (Embedded API):
- Direct calls to
tx.createNode(),node.setProperty(),relationship.delete(), etc. - Each operation wrapped in its own transaction
- No Cypher query parsing overhead
JanusGraph (TinkerPop Structure API):
- Direct calls to
g.addV(),vertex.property(),g.V().drop(), etc. - Transaction management via
g.tx().commit() - No Gremlin query parsing overhead
ArangoDB (Fuerte C++ API):
- Direct calls to ArangoDB HTTP API via Fuerte driver
- Batch operations using AQL
FOR ... IN @array INSERT/UPDATE/REMOVEpatterns - Optimized batch insertion for graph loading
- No AQL query parsing overhead for batch operations
Aster (RocksGraph API):
- Direct calls to RocksGraph API (RocksDB-based graph extension)
- Native C++ embedded graph operations:
AddVertex(),AddEdge(),DeleteEdge(),GetAllEdges() - Property graph support with
AddVertexProperty(),AddEdgeProperty(),GetVerticesWithProperty() - Adaptive edge update policy for optimized performance
- Zero query parsing overhead (pure embedded API)
OrientDB (Native API):
- Direct calls to OrientDB embedded API
- Transaction management and batch operations
- Native graph traversal and property operations
Sqlg (TinkerPop Structure API):
- Direct calls to TinkerPop Structure API with H2 backend
- Transaction management via
g.tx().commit() - No Gremlin query parsing overhead
| Database | Version | Release Date | Backend/Driver | Benchmark Mode | Requirements |
|---|---|---|---|---|---|
| Neo4j | 2026.01.4 | January 2026 | Embedded | Embedded | Java 21 |
| JanusGraph | 1.2.0-20251114-142114.b424a8f | November 14, 2025 | BerkeleyDB | Embedded | Java 17 |
| ArangoDB | 3.12.7-2 | December 2024 | Fuerte C++ driver | Client-Server | C++17 |
| Aster | Latest | 2025 | Embedded Graph | Embedded | C++17 |
| OrientDB | 3.2.49 | January 2026 | Embedded | Embedded | Java 17 |
| Sqlg | 3.1.6 | February 2026 | H2 Database | Embedded | Java 17 |
Benchmark Mode:
- Embedded: Database runs in the same process as the benchmark executor (Neo4j, JanusGraph, Aster, OrientDB, Sqlg)
- Client-Server: Database runs as a separate server process, benchmark executor connects via client API (ArangoDB)
- Native API Performance: Direct database API calls without query parsing overhead
- Separation of Concerns: Host handles orchestration; containers handle pure execution
- Latency Focus: Serial execution to accurately measure per-operation latency
- Clean Benchmarking: Docker isolation ensures reproducible, interference-free measurements
- Pure Execution Time: Timing excludes network transmission and only measures API call execution
- Extensibility: Easy to add new databases by implementing executor interfaces
graph-database-benchmark/
├── host/ # Python host implementation
│ ├── benchmark_launcher.py # Main entry point
│ ├── compiler/ # Workload compiler
│ │ └── workload_compiler.py
│ ├── dataset/ # Dataset loader
│ │ └── dataset_loader.py
│ ├── db/ # Docker manager
│ │ └── docker_manager.py
│ └── report/ # Report generator
│ └── report_generator.py
├── common/ # Shared Java components
│ └── src/main/java/com/graphbench/
│ ├── api/ # Core interfaces
│ │ ├── BenchmarkExecutor.java
│ │ └── WorkloadDispatcher.java
│ └── workload/ # Workload data models
│ ├── WorkloadTask.java
│ ├── AddVertexParams.java
│ ├── AddEdgeParams.java
│ └── ...
├── common-cpp/ # Shared C++ components (headers-only)
│ ├── CMakeLists.txt
│ └── include/graphbench/
│ ├── benchmark_executor.hpp # CRTP base class
│ ├── property_benchmark_executor.hpp
│ ├── progress_callback.hpp
│ └── benchmark_utils.hpp
├── docker/ # Docker container implementations
│ ├── neo4j/ # Neo4j embedded benchmark (Java)
│ │ ├── Dockerfile
│ │ ├── pom.xml
│ │ └── src/main/java/com/graphbench/neo4j/
│ │ ├── BenchmarkServer.java
│ │ └── Neo4jBenchmarkExecutor.java
│ ├── janusgraph/ # JanusGraph embedded benchmark (Java)
│ │ ├── Dockerfile
│ │ ├── pom.xml
│ │ └── src/main/java/com/graphbench/janusgraph/
│ │ ├── BenchmarkServer.java
│ │ └── JanusGraphBenchmarkExecutor.java
│ ├── arangodb/ # ArangoDB benchmark (C++)
│ │ ├── Dockerfile
│ │ ├── CMakeLists.txt
│ │ └── src/
│ │ ├── arango_utils.hpp
│ │ ├── arangodb_graph_loader.hpp
│ │ ├── arangodb_benchmark_executor.hpp
│ │ ├── arangodb_property_benchmark_executor.hpp
│ │ └── arangodb_benchmark_server.cpp
│ └── aster/ # Aster embedded benchmark (C++)
│ ├── Dockerfile
│ ├── CMakeLists.txt
│ └── src/
│ ├── aster_graph_loader.hpp
│ ├── aster_benchmark_executor.hpp
│ ├── aster_property_benchmark_executor.hpp
│ └── aster_benchmark_server.cpp
├── config/ # Configuration files
│ ├── database-config.json
│ └── datasets.json
├── workloads/ # Workload templates
│ └── templates/
│ └── example_workload.json
├── visualize/ # Visualization scripts
│ ├── plot_batchsize_comparison.py
│ ├── plot_performance_comparison.py
│ ├── examples.sh
│ └── README.md
├── graph-datasets/ # Dataset files (submodule)
├── reports/ # Benchmark results (generated)
├── plots/ # Visualization outputs (generated)
├── requirements.txt # Python dependencies
├── visualize.sh # Visualization wrapper script
└── README.md
MIT