diff --git a/.trae-html-share-packages/geaflow-learning-map/geaflow-learning-map.html.zip b/.trae-html-share-packages/geaflow-learning-map/geaflow-learning-map.html.zip new file mode 100644 index 000000000..059b6f6b6 Binary files /dev/null and b/.trae-html-share-packages/geaflow-learning-map/geaflow-learning-map.html.zip differ diff --git a/geaflow-ai/GeaFlowBackend/Gearflow-backend-vertical-slice.md b/geaflow-ai/GeaFlowBackend/Gearflow-backend-vertical-slice.md new file mode 100644 index 000000000..74ba81bae --- /dev/null +++ b/geaflow-ai/GeaFlowBackend/Gearflow-backend-vertical-slice.md @@ -0,0 +1,286 @@ +# Minimal GeaFlow Backend Vertical Slice — Design + +Status: **Open / For review** · Scope: geaflow-ai · Issue: apache/geaflow#849 + +## 1. Context & Goals + +`geaflow-ai` today is a separate, in-process RAG / graph-memory library. It builds +a `MemoryGraph` in the JVM and serves retrieval through `GraphMemoryServer`. The main +GeaFlow engine (DSL + runtime) is **not** wired in as a Graph Memory backend — +`GraphComputeEngine` is an empty marker interface, and nothing consumes it. + +This document scopes the **minimal vertical slice** that lets a `GeaFlowBackend` +treat the real GeaFlow engine as a backend of geaflow-ai, so the two gain one shared +contract early instead of diverging. It deliberately does **not** implement anything +(see Issue 848 for the local persistent prototype, Issue 847 for the `GraphBackend` +SPI). + +> This is a scoping/design deliverable only. No production HA claim is made. + +## 2. Current State (two paths, one in-memory store) + +| Concern | Abstraction | Impl today | +| ---------------- | ------------------------------------------------------------------------------------------------- | ------------------------------------------ | +| Read / retrieval | `GraphAccessor` (`getVertex`, `getEdge`, `scanVertex`, `scanEdge`, `expand`, `copy`) | `LocalMemoryGraphAccessor` → `MemoryGraph` | +| Write / mutation | `MutableGraph` (`addVertex`, `updateVertex`, `removeVertex`, `addEdge`, `removeEdge`, schema ops) | `MemoryMutableGraph` → `MemoryGraph` | +| Backend seam | `GraphComputeEngine` | **empty interface, unused** | +| Server entry | `GraphMemoryServer` | iterates configured index stores | + +Observations: + +- Everything funnels to `MemoryGraph` (in-JVM maps), loaded once from graph files. + +- Writes are in-memory only; nothing survives restart, and there is no transaction / + checkpoint boundary. + +- The natural seam for a real backend already exists (`GraphComputeEngine`) but is empty. + +## 3. Goal of the Slice + +Prove that geaflow-ai can address GeaFlow-the-engine through one **shared, stable, +testable interface**, with the smallest surface that still exercises read + write + a +restart/checkpoint story end to end. + +## 4. Options for the First Slice + +Maintainers must decide among three candidates: + +### 4.1 Option A — Read-only slice + +- GeaFlow engine ingests the graph (already possible via GQL) and geaflow-ai reads it + only through `GraphAccessor` backed by a GeaFlow-backed accessor. + +- **Pros**: smallest surface; reuses existing read API; fastest to land. + +- **Cons**: does not exercise writes, so mutation/checkpoint risks surface later. + +- **Test scenario**: ingest a 3-vertex / 2-edge graph via GQL, then run + `getVertex("person","alice")` and `scanEdge(aliceVertex)`; expect the 2 edges back, + ordered deterministically. + +- **Existing tests that change**: may defer; the GeaFlow-backed accessor would need a + new parity suite (cf. Issue 850), but no `LocalMemoryGraphAccessor` test set is broken + because only the fake backend is swapped as a mirror. + +- **Effort**: \~3 developer-days (read-path accessor + local-run params). + +### 4.2 Option B — Write-through slice *(recommended)* + +- geaflow-ai writes via a GeaFlow-backed `MutableGraph`; GeaFlow materializes the graph + in its own store; reads come back through the same `GraphAccessor`. + +- **Pros**: covers both paths, gives a real write -> restart -> read loop, matches how + the extracted-facts pipeline (Issues 837–845) actually produces graphs. + +- **Cons**: needs write-path mapping (DSL table/state) which touches more runtime surface. + +- **Test scenario**: `addVertex(alice)` + `addEdge(alice→bob)` -> `flush/checkpoint` -> + new JVM -> `getVertex/getEdge` return identical `alice`,`bob`,`alice→bob`; then + `updateVertex(alice.age=30)` and re-checkpoint reads the updated value (idempotent + replay). + +- **Existing tests that change**: the shared contract suite (cf. Issue 850) now runs + against **both** the in-memory fake and the GeaFlow-backed backend; mutation tests + (`addVertexSchema` / `updateVertex` / `removeEdge`) must pass on both. + +- **Effort**: \~7–10 developer-days (write-path mapping is the bulk). + +### 4.3 Option C — Projector-based slice + +- A GeaFlow job projects/subgraphs the graph, and geaflow-ai consumes the projection. + +- **Pros**: cleanly uses GeaFlow's vertex-centric model. + +- **Cons**: higher conceptual overhead; more moving parts for a first slice. + +- **Test scenario**: run a projection job that emits the 2-hop neighborhood of `alice`; + expect a subgraph of exactly {alice, bob, carol} and their edges, before any query. + +- **Existing tests that change**: introduces a new job-orchestration surface (the + projector job) rather than touching the accessor; the geaflow-ai suite gains a + subgraph-contract test but no existing suite is rewritten. + +- **Effort**: \~10–12 developer-days (highest; projector job + handoff). + +### 4.4 Decision matrix + +| Criterion | A (read-only) | B (write-through) | C (projector) | +| --------------------------- | ------------- | ----------------- | --------------- | +| Exercises write path | no | yes | n/a (job side) | +| write → restart → read loop | no | **yes** | indirect | +| Runtime surface touched | minimal | moderate | largest | +| Effort | \~3 d | \~7–10 d | \~10–12 d | +| Fallback safety | — | degrades to A | degrades to B/A | + +> Means of comparison only; effort is a relative sizing for maintainers, not a commitment. + +## 5. Supported Operations (assuming Option B as the baseline) + +- **Write**: `addVertex/updateVertex/removeVertex`, `addEdge/removeEdge`, + `addVertexSchema/addEdgeSchema` (the existing `MutableGraph` method set). + +- **Read**: `getVertex`, `getEdge`, `scanVertex`, `scanEdge`, `expand` (existing + `GraphAccessor` method set). + +- **Lifecycle**: open/close backend; flush; a defined checkpoint/commit point; scan + normalization (sorted/deterministic where contract requires). + +## 6. Non-goals (explicitly out of scope for the slice) + +- No production HA / distributed consistency guarantees. + +- No ANN / vector-search integration work (separate line of issues). + +- No replacement of the existing `MemoryGraph` main path everywhere. + +- No schema evolution or cross-tenant semantics in this slice. + +- No bit-identical parity with HugeGraph Server. + +## 7. Mapping to Existing GeaFlow DSL/Runtime Components + +| GeaFlow component | Role in the slice | +| ----------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- | +| GQL / `geaflow-dsl` | Authoritative graph schema + ingest; express read projections | +| Vertex-centric API (`AlgorithmUserFunction`, `sendMessage`) | `scanVertex`/`scanEdge`/`expand` semantics via iteration/messaging (cf. `udf/graph/PageRank`, `SingleSourceShortestPath`) | +| State store (`geaflow-store`, rocksdb/memory) | Where upserted vertices/edges physically live; basis of checkpoint | +| File IO / persistent root (`geaflow.file.persistent.root`) | Export/import of graph snapshots for restart+reload tests | +| Local run (`LocalClusterManager` / GQL client) | Cheap way to run the slice in integration tests | + +## 8. Checkpoint Behavior + +- Define a **commit/flush boundary** that persists outstanding writes to the GeaFlow + store atomically (reuse engine checkpoint if available; otherwise a per-backend + flush marker). + +- On restart, reload from the last completed checkpoint; any write after the last + checkpoint must be either re-applied idempotently or fail closed — never silently lost. + +- Expose the checkpoint/lag state so a later `SearchableWatermark` (Issue 857) can read it. + +## 9. Missing APIs / Blockers to Resolve + +The slice must close these gaps before/when it lands. Where a blocker is "empty yet", +a minimum design is proposed here — as a decision, not a code sketch. + +1. **`GraphComputeEngine`** **lifecycle contract.** The interface is empty; that is the + *problem*, not the answer. The slice should agree a minimal lifecycle up-front: + + - `init(config)` — parse backend config (root dir, store type), open resources; + + - `close()` — release resources, no-op if never `init`ed; + + - capability flags exposed via a single read-only query (e.g. `capabilities()` + returning whether the backend is writable / restartable / in-process), so a + caller never assumes behavior. + This is the smallest contract that lets the runtime open/close a real engine + without leaking implementation details. + +2. **Cross-process read semantics for** **`scanVertex`/`scanEdge`.** Both return a Java + `Iterator` (`GraphAccessor`) — an in-process object graph that cannot cross a + process/job boundary. A GeaFlow-backed accessor must choose a strategy: + + - **Lazy fetch**: iterate a server-side cursor and page results (streaming, needs + a companion cursor API); or + + - **Materialized list**: run the iteration eagerly and return a bounded snapshot + (simple, memory-bounded, fits small test graphs); or + + - **Native traversal bypass**: express scan/expand natively in GeaFlow's + vertex-centric iteration and return a result set rather than an `Iterator`. + Recommendation for the slice: **materialized list**, because the vertical-slice + graphs are small and it keeps `scanEdge(GraphVertex)`'s signature stable as a + drop-in for the fake backend. Documenting this choice now avoids a silent + cross-process break later. + +3. **Write-path mapping.** No path exists from `MutableGraph` to a GeaFlow DSL/state + surface. Pick **table-based** for the slice (GQL-computable, introspection-friendly) + over raw state-store writes (faster but bypasses schema). This is Option B's bulk. + +4. **Identity semantics.** `scanEdge(GraphVertex)` and `MemoryGraph` key on in-memory + objects; a GeaFlow-backed impl must define equal `id`/`label` semantics + (`getVertex(label,id)` as the canonical key) so the two backends stay + interchangeable in the same contract suite. + +5. **Backend wiring / registration.** `GraphMemoryServer` does **not** discover + backends through `GraphComputeEngine` — it collects `GraphAccessor` via + `addGraphAccessor(...)` and `IndexStore` via `addIndexStore(...)`, and `search()` + only ever reads `graphAccessors.get(0)`. So the integration seam to settle is: + how does a newly built GeaFlow-backed accessor get **registered** so the server + picks it up — a build-time binding, a config key that selects the backend type, or + promotion of a `GraphBackend` SPI (Issue 847)? This doc should record the chosen + registration point so the wiring is explicit, not implicit in a test harness. + +6. **Startup trace signal.** No standard "backend started / graph loaded" signal for + retrieval-parity checks; define one log/metric line the parity suite can assert on. + +## 10. Test Strategy + +- **Contract tests** shared by a fake in-memory backend and the GeaFlow-backed backend + (same `GraphBackend`-level suite; cf. Issue 850). + +- **Round-trip**: write → flush/checkpoint → new JVM instance → read back equal graph. + +- **Failure injection**: truncated/partial write fails closed or is quarantined + (aligns with Issue 846/848). + +- **Idempotent replay**: re-applying the same upserts yields one logical state + (aligns with Issue 845). + +- Run the suite locally via the existing local-run path (fast, no external services). + +## 11. Rollback & Compatibility + +- Keep the existing `MemoryGraph` path untouched and default; the GeaFlow backend is an + opt-in implementation of the shared contract. + +- No breaking change to `GraphAccessor`/`MutableGraph` signatures in this slice; new + contracts are additive. + +- Public SPI additions require maintainer review before promotion. + +### Migration paths between approaches + +- **B → A (write path too risky)**: stop routing `MutableGraph` writes to GeaFlow; the + accessor keeps serving reads. No data migration needed — the GeaFlow store is + already materialized from the ingested graph, so reads stay valid. + +- **B → C (projection becomes the real need)**: add the projector job as a read-side + producer; writes remain as-is. A/C can coexist because both read from the same store. + +- **C → B (need write-back)**: the projector stays but geaflow-ai writes land on the + store directly, with the projector re-run to refresh projections — a refresh, not a + rewrite of the writer. + +- **Any → A**: A is the common floor; every option degrades to read-only without + touching the default `MemoryGraph` path. + +## 12. Open Questions + +- Write target: GeaFlow **DSL table** vs directly against **state store**? (affects §7 map) + +- Does a local GeaFlow instance run embedded in-process, or as a separate local job + the slice shells out to? + +- Who owns the `GraphComputeEngine` lifecycle contract — this slice or Issue 847 SPI? + +## 13. Recommendation + +Proceed with **Option B (write-through)** as the primary vertical slice, with **Option A +as a committed fallback**. After the B read/write loop proves in the local run, escalate +to the shared contract suite (Issue 850) before any broader rollout. + +**Rollback trigger & plan**: if the write-path mapping (Blockers 3/5) proves unstable +or over-scoped, revert to **A (read-only)** for the first landed slice and re-open B via +the §11 `B → A` migration path (reads stay valid, no data move). The decision record is +the record of why B was chosen and under what condition it degrades — so a later +maintainer can reverse it without re-deriving the context. + +**Decision inputs land first**: resolve §12 open questions (write target; embedded vs +separate local job; `GraphComputeEngine` lifecycle ownership) **in this design phase**, +because each one changes the blocker proposal above. A maintainer should be able to say +"yes, go with B" from this document alone, without inspecting Java code — that is the +acceptance bar for this design doc. + +Land this design first, then implement against 847 (`GraphBackend`) reusing these +decisions. diff --git a/geaflow-learning-map/_shared/fonts/JetBrainsMono-Bold.ttf b/geaflow-learning-map/_shared/fonts/JetBrainsMono-Bold.ttf new file mode 100644 index 000000000..1926c804b Binary files /dev/null and b/geaflow-learning-map/_shared/fonts/JetBrainsMono-Bold.ttf differ diff --git a/geaflow-learning-map/_shared/fonts/JetBrainsMono-Regular.ttf b/geaflow-learning-map/_shared/fonts/JetBrainsMono-Regular.ttf new file mode 100644 index 000000000..436c982ff Binary files /dev/null and b/geaflow-learning-map/_shared/fonts/JetBrainsMono-Regular.ttf differ diff --git a/geaflow-learning-map/_shared/fonts/WorkSans-Bold.ttf b/geaflow-learning-map/_shared/fonts/WorkSans-Bold.ttf new file mode 100644 index 000000000..5c9798929 Binary files /dev/null and b/geaflow-learning-map/_shared/fonts/WorkSans-Bold.ttf differ diff --git a/geaflow-learning-map/_shared/fonts/WorkSans-Regular.ttf b/geaflow-learning-map/_shared/fonts/WorkSans-Regular.ttf new file mode 100644 index 000000000..d24586cc0 Binary files /dev/null and b/geaflow-learning-map/_shared/fonts/WorkSans-Regular.ttf differ diff --git a/geaflow-learning-map/geaflow-learning-map.html b/geaflow-learning-map/geaflow-learning-map.html new file mode 100644 index 000000000..19aeb6c92 --- /dev/null +++ b/geaflow-learning-map/geaflow-learning-map.html @@ -0,0 +1,744 @@ + + + +
+ + +GeaFlow 完整学习路线:5 个阶段 · 7 个里程碑 · 逐模块技术映射
+ +路线分 5 个阶段:前 3 个阶段是"地基"(不看懂就寸步难行),阶段 3 是"灵魂"(GeaFlow 的领域知识),阶段 4 是"周边"(按兴趣选学)。每个技术点右侧标注优先级:P0 必学 · P1 重要 · P2 按需。
+ +目标:能顺利读懂任何一段 GeaFlow 源码——不要求会写复杂 Java,但接口、泛型、Lambda 必须一眼明白。
+<K, V>
+ P0集合 List / Map / Set
+ P0Lambda 与函数式接口
+ P1异常体系 try/catch
+ P1注解(≈Python 装饰器)
+ P1内部类 / 匿名类
+ geaflow-model(全是接口定义,最适合第一天读:IVertex、IEdge)
+ 目标:能把项目编译出来、跑起来、打断点调试。工具不熟,后面每一步都事倍功半。
+pom.xml / 依赖 / 多模块 / 打包
+ P0IntelliJ IDEA 断点调试
+ P0Git(应已掌握)
+ P1日志 SLF4J / Log4j2
+ P1JUnit 单测(≈ pytest)
+ pom.xml 管全部子模块,build.sh 只是 Maven 命令的封装。build.sh根 pom.xmltools/checkstyle.xml
+ 目标:理解引擎"为什么这样写"——并发模型、内存管理、插件加载,这三样撑起了整个分布式引擎。
+geaflow-plugins(SPI 发现)
+ geaflow-memory(堆外内存池 Chunk/Page)
+ geaflow-common
+ 目标:掌握 GeaFlow 的设计思想。这一阶段不是学"新语言",而是学"新世界观"——数据流、图迭代、SQL 编译、分布式协作。
+geaflow-coregeaflow-dslgeaflow-stategeaflow-deploygeaflow-examples
+ 目标:按想深入的子系统选方向。做平台开发学 Web 栈,做云原生学 K8s 栈,做 AI 结合学 MCP/LLM 栈。
+geaflow-consolegeaflow-kubernetes-operatorgeaflow-aigeaflow-mcpgeaflow-dashboard
+ 想知道"读懂某个目录需要什么技术",查这张表即可。按优先级从上往下读,前 5 行覆盖了核心引擎 80% 的代码量。
+ +| 项目模块(目录) | +在项目中的作用 | +读懂它需要的技术 | +优先级 / 建议 | +
|---|---|---|---|
| geaflow/geaflow-model数据模型与接口定义 | +定义 IVertex / IEdge / 消息 / Traversal 请求响应等"数据长什么样" | +Java 接口、泛型、继承体系 | +P0 · 第一天就读 | +
| geaflow/geaflow-core核心计算引擎(心脏) | +Pipeline API → 逻辑计划 → 物理执行图 → 调度器 → Worker/算子执行 | +并发编程、流/图计算概念、设计模式(工厂/模板方法/状态机) | +P0 · 主战场 | +
| geaflow/geaflow-examples示例与上手入口 | +PageRank、WordCount 等 Java API 用法示范 + GQL 脚本样例 | +Java 基础即可读懂;GQL 样例需 SQL 基础 | +P0 · 与 core 对照读 | +
| geaflow/geaflow-dslSQL+GQL 语言层 | +用 Apache Calcite 把 GQL 文本解析、优化、翻译成可执行计划 | +编译原理基础(词法/语法/AST)、Calcite、SQL | +P1 · 编译器方向 | +
| geaflow/geaflow-state图状态存储 | +GraphState 抽象:静态/动态图、快照、版本化查询,支撑增量计算 | +KV 存储概念、RocksDB/Redis、快照与版本 | +P1 · 存储方向 | +
| geaflow/geaflow-common公共基础库 | +配置 Configuration、序列化 Kryo、类型系统、异常、线程池 | +Java 基础、序列化概念 | +P1 · 随用随查 | +
| geaflow/geaflow-plugins插件体系 | +Store(内存/RocksDB/Redis/JDBC)、服务发现(Redis/ZK)、文件系统(OSS/S3/HDFS) | +Java SPI / ServiceLoader 机制 | +P1 · 学 SPI 必读 | +
| geaflow/geaflow-memory内存管理 | +堆外内存池:Chunk/Page 分配、LZ4/Snappy 压缩、引用清理 | +JVM 内存模型、直接内存(贴近 C 的 malloc 思维) | +P2 · 性能方向 | +
| geaflow/geaflow-deploy部署与集群 | +三种部署实现:Local(本地)/ Ray / Kubernetes,含 Master-Driver-Container 启动与容错 | +分布式架构概念、K8s、Ray 基础 | +P2 · 先读 Local | +
| geaflow/geaflow-metrics监控指标 | +Counter/Gauge/Histogram 指标 + Prometheus/InfluxDB 上报 | +可观测性概念(Python 里 ≈ prometheus_client) | +P2 · 简单 | +
| geaflow/geaflow-inferAI 推理桥接 | +Java 进程拉起 Python 推理环境,双向数据交换(含 pickle 协议实现) | +Java+Python 混合编程、进程间通信 | +P2 · Python 主场 | +
| geaflow-console一站式研发平台 | +Web 控制台:作业/图/表/函数元数据管理、白屏提交 | +Spring Boot、MyBatis、MySQL、React | +P2 · 平台方向 | +
| geaflow-kubernetes-operator云原生部署器 | +K8s Operator:监听 GeaflowJob CRD,自动创建/销毁/恢复作业 | +Kubernetes(CRD/Reconcile 模式)、fabric8 | +P2 · 云原生方向 | +
| geaflow-aiAI 图扩展 | +图记忆/语义搜索服务 + Python 图算法仿真工具 | +LLM、Embedding、NetworkX(Python) | +P2 · AI 方向 | +
| geaflow-mcpMCP 服务 | +把建图/查询/取 Schema 封装成 MCP 工具,供 AI 助手调用 | +MCP 协议概念 | +P2 · AI 方向 | +
| geaflow-analytics-service交互式分析服务 | +常驻查询服务:HTTP/RPC 接收 GQL,即时返回结果 | +HTTP/RPC 服务概念 | +P2 · 服务方向 | +
| geaflow-dashboard引擎监控面板 | +查看集群组件、Pipeline、日志、线程、火焰图的 Web UI | +React、TypeScript | +P2 · 前端方向 | +
技术学习必须落在代码上。每完成一步,说明对应阶段的知识已内化;走完 7 步,即达到"大概了解并掌握整个项目"的目标。
+ +build.sh 编译 + gql_submit.sh 提交环路检测示例,验证环境与工具链。