diff --git a/.trae-html-share-packages/geaflow-learning-map/geaflow-learning-map.html.zip b/.trae-html-share-packages/geaflow-learning-map/geaflow-learning-map.html.zip new file mode 100644 index 000000000..059b6f6b6 Binary files /dev/null and b/.trae-html-share-packages/geaflow-learning-map/geaflow-learning-map.html.zip differ diff --git a/geaflow-ai/GeaFlowBackend/Gearflow-backend-vertical-slice.md b/geaflow-ai/GeaFlowBackend/Gearflow-backend-vertical-slice.md new file mode 100644 index 000000000..74ba81bae --- /dev/null +++ b/geaflow-ai/GeaFlowBackend/Gearflow-backend-vertical-slice.md @@ -0,0 +1,286 @@ +# Minimal GeaFlow Backend Vertical Slice — Design + +Status: **Open / For review** · Scope: geaflow-ai · Issue: apache/geaflow#849 + +## 1. Context & Goals + +`geaflow-ai` today is a separate, in-process RAG / graph-memory library. It builds +a `MemoryGraph` in the JVM and serves retrieval through `GraphMemoryServer`. The main +GeaFlow engine (DSL + runtime) is **not** wired in as a Graph Memory backend — +`GraphComputeEngine` is an empty marker interface, and nothing consumes it. + +This document scopes the **minimal vertical slice** that lets a `GeaFlowBackend` +treat the real GeaFlow engine as a backend of geaflow-ai, so the two gain one shared +contract early instead of diverging. It deliberately does **not** implement anything +(see Issue 848 for the local persistent prototype, Issue 847 for the `GraphBackend` +SPI). + +> This is a scoping/design deliverable only. No production HA claim is made. + +## 2. Current State (two paths, one in-memory store) + +| Concern | Abstraction | Impl today | +| ---------------- | ------------------------------------------------------------------------------------------------- | ------------------------------------------ | +| Read / retrieval | `GraphAccessor` (`getVertex`, `getEdge`, `scanVertex`, `scanEdge`, `expand`, `copy`) | `LocalMemoryGraphAccessor` → `MemoryGraph` | +| Write / mutation | `MutableGraph` (`addVertex`, `updateVertex`, `removeVertex`, `addEdge`, `removeEdge`, schema ops) | `MemoryMutableGraph` → `MemoryGraph` | +| Backend seam | `GraphComputeEngine` | **empty interface, unused** | +| Server entry | `GraphMemoryServer` | iterates configured index stores | + +Observations: + +- Everything funnels to `MemoryGraph` (in-JVM maps), loaded once from graph files. + +- Writes are in-memory only; nothing survives restart, and there is no transaction / + checkpoint boundary. + +- The natural seam for a real backend already exists (`GraphComputeEngine`) but is empty. + +## 3. Goal of the Slice + +Prove that geaflow-ai can address GeaFlow-the-engine through one **shared, stable, +testable interface**, with the smallest surface that still exercises read + write + a +restart/checkpoint story end to end. + +## 4. Options for the First Slice + +Maintainers must decide among three candidates: + +### 4.1 Option A — Read-only slice + +- GeaFlow engine ingests the graph (already possible via GQL) and geaflow-ai reads it + only through `GraphAccessor` backed by a GeaFlow-backed accessor. + +- **Pros**: smallest surface; reuses existing read API; fastest to land. + +- **Cons**: does not exercise writes, so mutation/checkpoint risks surface later. + +- **Test scenario**: ingest a 3-vertex / 2-edge graph via GQL, then run + `getVertex("person","alice")` and `scanEdge(aliceVertex)`; expect the 2 edges back, + ordered deterministically. + +- **Existing tests that change**: may defer; the GeaFlow-backed accessor would need a + new parity suite (cf. Issue 850), but no `LocalMemoryGraphAccessor` test set is broken + because only the fake backend is swapped as a mirror. + +- **Effort**: \~3 developer-days (read-path accessor + local-run params). + +### 4.2 Option B — Write-through slice *(recommended)* + +- geaflow-ai writes via a GeaFlow-backed `MutableGraph`; GeaFlow materializes the graph + in its own store; reads come back through the same `GraphAccessor`. + +- **Pros**: covers both paths, gives a real write -> restart -> read loop, matches how + the extracted-facts pipeline (Issues 837–845) actually produces graphs. + +- **Cons**: needs write-path mapping (DSL table/state) which touches more runtime surface. + +- **Test scenario**: `addVertex(alice)` + `addEdge(alice→bob)` -> `flush/checkpoint` -> + new JVM -> `getVertex/getEdge` return identical `alice`,`bob`,`alice→bob`; then + `updateVertex(alice.age=30)` and re-checkpoint reads the updated value (idempotent + replay). + +- **Existing tests that change**: the shared contract suite (cf. Issue 850) now runs + against **both** the in-memory fake and the GeaFlow-backed backend; mutation tests + (`addVertexSchema` / `updateVertex` / `removeEdge`) must pass on both. + +- **Effort**: \~7–10 developer-days (write-path mapping is the bulk). + +### 4.3 Option C — Projector-based slice + +- A GeaFlow job projects/subgraphs the graph, and geaflow-ai consumes the projection. + +- **Pros**: cleanly uses GeaFlow's vertex-centric model. + +- **Cons**: higher conceptual overhead; more moving parts for a first slice. + +- **Test scenario**: run a projection job that emits the 2-hop neighborhood of `alice`; + expect a subgraph of exactly {alice, bob, carol} and their edges, before any query. + +- **Existing tests that change**: introduces a new job-orchestration surface (the + projector job) rather than touching the accessor; the geaflow-ai suite gains a + subgraph-contract test but no existing suite is rewritten. + +- **Effort**: \~10–12 developer-days (highest; projector job + handoff). + +### 4.4 Decision matrix + +| Criterion | A (read-only) | B (write-through) | C (projector) | +| --------------------------- | ------------- | ----------------- | --------------- | +| Exercises write path | no | yes | n/a (job side) | +| write → restart → read loop | no | **yes** | indirect | +| Runtime surface touched | minimal | moderate | largest | +| Effort | \~3 d | \~7–10 d | \~10–12 d | +| Fallback safety | — | degrades to A | degrades to B/A | + +> Means of comparison only; effort is a relative sizing for maintainers, not a commitment. + +## 5. Supported Operations (assuming Option B as the baseline) + +- **Write**: `addVertex/updateVertex/removeVertex`, `addEdge/removeEdge`, + `addVertexSchema/addEdgeSchema` (the existing `MutableGraph` method set). + +- **Read**: `getVertex`, `getEdge`, `scanVertex`, `scanEdge`, `expand` (existing + `GraphAccessor` method set). + +- **Lifecycle**: open/close backend; flush; a defined checkpoint/commit point; scan + normalization (sorted/deterministic where contract requires). + +## 6. Non-goals (explicitly out of scope for the slice) + +- No production HA / distributed consistency guarantees. + +- No ANN / vector-search integration work (separate line of issues). + +- No replacement of the existing `MemoryGraph` main path everywhere. + +- No schema evolution or cross-tenant semantics in this slice. + +- No bit-identical parity with HugeGraph Server. + +## 7. Mapping to Existing GeaFlow DSL/Runtime Components + +| GeaFlow component | Role in the slice | +| ----------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- | +| GQL / `geaflow-dsl` | Authoritative graph schema + ingest; express read projections | +| Vertex-centric API (`AlgorithmUserFunction`, `sendMessage`) | `scanVertex`/`scanEdge`/`expand` semantics via iteration/messaging (cf. `udf/graph/PageRank`, `SingleSourceShortestPath`) | +| State store (`geaflow-store`, rocksdb/memory) | Where upserted vertices/edges physically live; basis of checkpoint | +| File IO / persistent root (`geaflow.file.persistent.root`) | Export/import of graph snapshots for restart+reload tests | +| Local run (`LocalClusterManager` / GQL client) | Cheap way to run the slice in integration tests | + +## 8. Checkpoint Behavior + +- Define a **commit/flush boundary** that persists outstanding writes to the GeaFlow + store atomically (reuse engine checkpoint if available; otherwise a per-backend + flush marker). + +- On restart, reload from the last completed checkpoint; any write after the last + checkpoint must be either re-applied idempotently or fail closed — never silently lost. + +- Expose the checkpoint/lag state so a later `SearchableWatermark` (Issue 857) can read it. + +## 9. Missing APIs / Blockers to Resolve + +The slice must close these gaps before/when it lands. Where a blocker is "empty yet", +a minimum design is proposed here — as a decision, not a code sketch. + +1. **`GraphComputeEngine`** **lifecycle contract.** The interface is empty; that is the + *problem*, not the answer. The slice should agree a minimal lifecycle up-front: + + - `init(config)` — parse backend config (root dir, store type), open resources; + + - `close()` — release resources, no-op if never `init`ed; + + - capability flags exposed via a single read-only query (e.g. `capabilities()` + returning whether the backend is writable / restartable / in-process), so a + caller never assumes behavior. + This is the smallest contract that lets the runtime open/close a real engine + without leaking implementation details. + +2. **Cross-process read semantics for** **`scanVertex`/`scanEdge`.** Both return a Java + `Iterator` (`GraphAccessor`) — an in-process object graph that cannot cross a + process/job boundary. A GeaFlow-backed accessor must choose a strategy: + + - **Lazy fetch**: iterate a server-side cursor and page results (streaming, needs + a companion cursor API); or + + - **Materialized list**: run the iteration eagerly and return a bounded snapshot + (simple, memory-bounded, fits small test graphs); or + + - **Native traversal bypass**: express scan/expand natively in GeaFlow's + vertex-centric iteration and return a result set rather than an `Iterator`. + Recommendation for the slice: **materialized list**, because the vertical-slice + graphs are small and it keeps `scanEdge(GraphVertex)`'s signature stable as a + drop-in for the fake backend. Documenting this choice now avoids a silent + cross-process break later. + +3. **Write-path mapping.** No path exists from `MutableGraph` to a GeaFlow DSL/state + surface. Pick **table-based** for the slice (GQL-computable, introspection-friendly) + over raw state-store writes (faster but bypasses schema). This is Option B's bulk. + +4. **Identity semantics.** `scanEdge(GraphVertex)` and `MemoryGraph` key on in-memory + objects; a GeaFlow-backed impl must define equal `id`/`label` semantics + (`getVertex(label,id)` as the canonical key) so the two backends stay + interchangeable in the same contract suite. + +5. **Backend wiring / registration.** `GraphMemoryServer` does **not** discover + backends through `GraphComputeEngine` — it collects `GraphAccessor` via + `addGraphAccessor(...)` and `IndexStore` via `addIndexStore(...)`, and `search()` + only ever reads `graphAccessors.get(0)`. So the integration seam to settle is: + how does a newly built GeaFlow-backed accessor get **registered** so the server + picks it up — a build-time binding, a config key that selects the backend type, or + promotion of a `GraphBackend` SPI (Issue 847)? This doc should record the chosen + registration point so the wiring is explicit, not implicit in a test harness. + +6. **Startup trace signal.** No standard "backend started / graph loaded" signal for + retrieval-parity checks; define one log/metric line the parity suite can assert on. + +## 10. Test Strategy + +- **Contract tests** shared by a fake in-memory backend and the GeaFlow-backed backend + (same `GraphBackend`-level suite; cf. Issue 850). + +- **Round-trip**: write → flush/checkpoint → new JVM instance → read back equal graph. + +- **Failure injection**: truncated/partial write fails closed or is quarantined + (aligns with Issue 846/848). + +- **Idempotent replay**: re-applying the same upserts yields one logical state + (aligns with Issue 845). + +- Run the suite locally via the existing local-run path (fast, no external services). + +## 11. Rollback & Compatibility + +- Keep the existing `MemoryGraph` path untouched and default; the GeaFlow backend is an + opt-in implementation of the shared contract. + +- No breaking change to `GraphAccessor`/`MutableGraph` signatures in this slice; new + contracts are additive. + +- Public SPI additions require maintainer review before promotion. + +### Migration paths between approaches + +- **B → A (write path too risky)**: stop routing `MutableGraph` writes to GeaFlow; the + accessor keeps serving reads. No data migration needed — the GeaFlow store is + already materialized from the ingested graph, so reads stay valid. + +- **B → C (projection becomes the real need)**: add the projector job as a read-side + producer; writes remain as-is. A/C can coexist because both read from the same store. + +- **C → B (need write-back)**: the projector stays but geaflow-ai writes land on the + store directly, with the projector re-run to refresh projections — a refresh, not a + rewrite of the writer. + +- **Any → A**: A is the common floor; every option degrades to read-only without + touching the default `MemoryGraph` path. + +## 12. Open Questions + +- Write target: GeaFlow **DSL table** vs directly against **state store**? (affects §7 map) + +- Does a local GeaFlow instance run embedded in-process, or as a separate local job + the slice shells out to? + +- Who owns the `GraphComputeEngine` lifecycle contract — this slice or Issue 847 SPI? + +## 13. Recommendation + +Proceed with **Option B (write-through)** as the primary vertical slice, with **Option A +as a committed fallback**. After the B read/write loop proves in the local run, escalate +to the shared contract suite (Issue 850) before any broader rollout. + +**Rollback trigger & plan**: if the write-path mapping (Blockers 3/5) proves unstable +or over-scoped, revert to **A (read-only)** for the first landed slice and re-open B via +the §11 `B → A` migration path (reads stay valid, no data move). The decision record is +the record of why B was chosen and under what condition it degrades — so a later +maintainer can reverse it without re-deriving the context. + +**Decision inputs land first**: resolve §12 open questions (write target; embedded vs +separate local job; `GraphComputeEngine` lifecycle ownership) **in this design phase**, +because each one changes the blocker proposal above. A maintainer should be able to say +"yes, go with B" from this document alone, without inspecting Java code — that is the +acceptance bar for this design doc. + +Land this design first, then implement against 847 (`GraphBackend`) reusing these +decisions. diff --git a/geaflow-learning-map/_shared/fonts/JetBrainsMono-Bold.ttf b/geaflow-learning-map/_shared/fonts/JetBrainsMono-Bold.ttf new file mode 100644 index 000000000..1926c804b Binary files /dev/null and b/geaflow-learning-map/_shared/fonts/JetBrainsMono-Bold.ttf differ diff --git a/geaflow-learning-map/_shared/fonts/JetBrainsMono-Regular.ttf b/geaflow-learning-map/_shared/fonts/JetBrainsMono-Regular.ttf new file mode 100644 index 000000000..436c982ff Binary files /dev/null and b/geaflow-learning-map/_shared/fonts/JetBrainsMono-Regular.ttf differ diff --git a/geaflow-learning-map/_shared/fonts/WorkSans-Bold.ttf b/geaflow-learning-map/_shared/fonts/WorkSans-Bold.ttf new file mode 100644 index 000000000..5c9798929 Binary files /dev/null and b/geaflow-learning-map/_shared/fonts/WorkSans-Bold.ttf differ diff --git a/geaflow-learning-map/_shared/fonts/WorkSans-Regular.ttf b/geaflow-learning-map/_shared/fonts/WorkSans-Regular.ttf new file mode 100644 index 000000000..d24586cc0 Binary files /dev/null and b/geaflow-learning-map/_shared/fonts/WorkSans-Regular.ttf differ diff --git a/geaflow-learning-map/geaflow-learning-map.html b/geaflow-learning-map/geaflow-learning-map.html new file mode 100644 index 000000000..19aeb6c92 --- /dev/null +++ b/geaflow-learning-map/geaflow-learning-map.html @@ -0,0 +1,744 @@ + + + + + + +GeaFlow 技术学习一图流 + + + +
+ +
+ APACHE GEAFLOW · 技术学习一图流 +

从 Python/C 到读懂流图计算引擎

+

GeaFlow 完整学习路线:5 个阶段 · 7 个里程碑 · 逐模块技术映射

+
+ 背景:只会 Python / C,Java 仅了解 + 目标:大概了解并掌握整个项目 + 项目:Java 分布式流图计算引擎 +
+
+ + +
+ 01 +

学习总路线(自下而上,按顺序推进)

+
+

路线分 5 个阶段:前 3 个阶段是"地基"(不看懂就寸步难行),阶段 3 是"灵魂"(GeaFlow 的领域知识),阶段 4 是"周边"(按兴趣选学)。每个技术点右侧标注优先级:P0 必学 · P1 重要 · P2 按需。

+ +
+ + +
+
阶段0
+
+
+

Java 语言地基

+ 约 2~4 周 + P0 · 地基 +
+

目标:能顺利读懂任何一段 GeaFlow 源码——不要求会写复杂 Java,但接口、泛型、Lambda 必须一眼明白。

+
+ P0类 / 对象 / 继承 / 多态 + P0接口 interface / 抽象类 + P0泛型 <K, V> + P0集合 List / Map / Set + P0Lambda 与函数式接口 + P1异常体系 try/catch + P1注解(≈Python 装饰器) + P1内部类 / 匿名类 +
+
给你已有的知识做类比:接口 ≈ C 的头文件契约 + 函数指针表;泛型 ≈ C++ 模板 / Python 类型注解;Lambda ≈ Python 的 lambda;集合 ≈ list / dict / set。
+
+ 入门读码点: + geaflow-model(全是接口定义,最适合第一天读:IVertex、IEdge) +
+
+
+ + +
+
阶段1
+
+
+

工程工具链

+ 约 1 周 + P0 · 顺手的工具 +
+

目标:能把项目编译出来、跑起来、打断点调试。工具不熟,后面每一步都事倍功半。

+
+ P0Maven:pom.xml / 依赖 / 多模块 / 打包 + P0IntelliJ IDEA 断点调试 + P0Git(应已掌握) + P1日志 SLF4J / Log4j2 + P1JUnit 单测(≈ pytest) +
+
类比:Maven ≈ pip + requirements.txt + make 三合一。GeaFlow 是 Maven 多模块工程,根 pom.xml 管全部子模块,build.sh 只是 Maven 命令的封装。
+
+ 入门读码点: + build.sh根 pom.xmltools/checkstyle.xml +
+
+
+ + +
+
阶段2
+
+
+

Java 进阶与 JVM

+ 约 2~3 周 + P0/P1 · 引擎的底牌 +
+

目标:理解引擎"为什么这样写"——并发模型、内存管理、插件加载,这三样撑起了整个分布式引擎。

+
+ P0多线程:线程池 / 锁 / volatile / 并发容器 + P0Java SPI / ServiceLoader 插件机制 + P1JVM 内存:堆 / 栈 / 直接内存 + P1ClassLoader 与反射 + P1序列化:Serializable / Kryo + P2Netty 与 RPC 通信 +
+
类比:Java SPI ≈ Python 包的 entry_points 插件注册机制;volatile / 内存可见性 ≈ C 的内存屏障与 cache coherence 问题,比 Python threading 更贴近 C 的思维。
+
+ 入门读码点: + geaflow-plugins(SPI 发现) + geaflow-memory(堆外内存池 Chunk/Page) + geaflow-common +
+
+
+ + +
+
阶段3
+
+
+

领域概念:流 + 图 + SQL(项目灵魂)

+ 持续投入 · 核心中枢 + P0 · 灵魂 +
+

目标:掌握 GeaFlow 的设计思想。这一阶段不是学"新语言",而是学"新世界观"——数据流、图迭代、SQL 编译、分布式协作。

+
+ P0流计算:Source / Sink / 窗口 Window + P0图计算:顶点中心 Pregel / BSP 迭代 + P0经典算法:PageRank / SSSP / WCC + P1状态管理:KV State / 快照 Snapshot + P1Checkpoint 与 Exactly-Once + P1SQL 编译:AST → 逻辑计划 → 优化 → 物理 + P1Apache Calcite(DSL 解析框架) + P1分布式:Master-Driver-Container / Failover + P1Shuffle 与 KeyGroup 数据分区 + P2存储引擎:RocksDB / Redis +
+
+ 对应模块: + geaflow-coregeaflow-dslgeaflow-stategeaflow-deploygeaflow-examples +
+
+
+ + +
+
阶段4
+
+
+

周边生态(按兴趣与岗位选学)

+ 按需 · 无先后 + P2 · 外围 +
+

目标:按想深入的子系统选方向。做平台开发学 Web 栈,做云原生学 K8s 栈,做 AI 结合学 MCP/LLM 栈。

+
+ P2Spring Boot + MyBatis + MySQL(Console 后端) + P2React + TypeScript + Ant Design(前端) + P2Docker + Kubernetes + Operator 模式 + P2Ray 分布式框架 + P2MCP 协议(AI 工具接入) + P2LLM / Embedding 基础(geaflow-ai) +
+
+ 对应模块: + geaflow-consolegeaflow-kubernetes-operatorgeaflow-aigeaflow-mcpgeaflow-dashboard +
+
+
+
+ +
+ P0:必学,不学看不懂代码 + P1:重要,决定理解深度 + P2:按需,只对应特定子模块 + 时间估算:按每天 2~3 小时投入估算 +
+ + +
+ 02 +

项目模块 × 所需技术 映射表

+
+

想知道"读懂某个目录需要什么技术",查这张表即可。按优先级从上往下读,前 5 行覆盖了核心引擎 80% 的代码量。

+ +
+ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
项目模块(目录)在项目中的作用读懂它需要的技术优先级 / 建议
geaflow/geaflow-model数据模型与接口定义定义 IVertex / IEdge / 消息 / Traversal 请求响应等"数据长什么样"Java 接口、泛型、继承体系P0 · 第一天就读
geaflow/geaflow-core核心计算引擎(心脏)Pipeline API → 逻辑计划 → 物理执行图 → 调度器 → Worker/算子执行并发编程、流/图计算概念、设计模式(工厂/模板方法/状态机)P0 · 主战场
geaflow/geaflow-examples示例与上手入口PageRank、WordCount 等 Java API 用法示范 + GQL 脚本样例Java 基础即可读懂;GQL 样例需 SQL 基础P0 · 与 core 对照读
geaflow/geaflow-dslSQL+GQL 语言层用 Apache Calcite 把 GQL 文本解析、优化、翻译成可执行计划编译原理基础(词法/语法/AST)、Calcite、SQLP1 · 编译器方向
geaflow/geaflow-state图状态存储GraphState 抽象:静态/动态图、快照、版本化查询,支撑增量计算KV 存储概念、RocksDB/Redis、快照与版本P1 · 存储方向
geaflow/geaflow-common公共基础库配置 Configuration、序列化 Kryo、类型系统、异常、线程池Java 基础、序列化概念P1 · 随用随查
geaflow/geaflow-plugins插件体系Store(内存/RocksDB/Redis/JDBC)、服务发现(Redis/ZK)、文件系统(OSS/S3/HDFS)Java SPI / ServiceLoader 机制P1 · 学 SPI 必读
geaflow/geaflow-memory内存管理堆外内存池:Chunk/Page 分配、LZ4/Snappy 压缩、引用清理JVM 内存模型、直接内存(贴近 C 的 malloc 思维)P2 · 性能方向
geaflow/geaflow-deploy部署与集群三种部署实现:Local(本地)/ Ray / Kubernetes,含 Master-Driver-Container 启动与容错分布式架构概念、K8s、Ray 基础P2 · 先读 Local
geaflow/geaflow-metrics监控指标Counter/Gauge/Histogram 指标 + Prometheus/InfluxDB 上报可观测性概念(Python 里 ≈ prometheus_client)P2 · 简单
geaflow/geaflow-inferAI 推理桥接Java 进程拉起 Python 推理环境,双向数据交换(含 pickle 协议实现)Java+Python 混合编程、进程间通信P2 · Python 主场
geaflow-console一站式研发平台Web 控制台:作业/图/表/函数元数据管理、白屏提交Spring Boot、MyBatis、MySQL、ReactP2 · 平台方向
geaflow-kubernetes-operator云原生部署器K8s Operator:监听 GeaflowJob CRD,自动创建/销毁/恢复作业Kubernetes(CRD/Reconcile 模式)、fabric8P2 · 云原生方向
geaflow-aiAI 图扩展图记忆/语义搜索服务 + Python 图算法仿真工具LLM、Embedding、NetworkX(Python)P2 · AI 方向
geaflow-mcpMCP 服务把建图/查询/取 Schema 封装成 MCP 工具,供 AI 助手调用MCP 协议概念P2 · AI 方向
geaflow-analytics-service交互式分析服务常驻查询服务:HTTP/RPC 接收 GQL,即时返回结果HTTP/RPC 服务概念P2 · 服务方向
geaflow-dashboard引擎监控面板查看集群组件、Pipeline、日志、线程、火焰图的 Web UIReact、TypeScriptP2 · 前端方向
+
+ + +
+ 03 +

上手里程碑:7 步验证学习成果

+
+

技术学习必须落在代码上。每完成一步,说明对应阶段的知识已内化;走完 7 步,即达到"大概了解并掌握整个项目"的目标。

+ +
+
+
跑通 Quick Start
+
build.sh 编译 + gql_submit.sh 提交环路检测示例,验证环境与工具链。
+
+
+
读懂两个示例
+
StreamWordCountPipeline(流)与 PageRank(图),对着源码理解 Pipeline API 的用法。
+
+
+
改造一个算子
+
在 MapOperator 或 UDF 里加一行日志,重新打包跑通,体会"改码 → 构建 → 运行"闭环。
+
+
+
顺链读 core
+
从 PipelineTaskExecutor → PipelinePlanBuilder → ExecutionGraphBuilder → Scheduler → Worker 断点走一遍。
+
+
+
读懂 DSL 编译
+
跟踪一条 GQL 语句从 GeaFlowDSLParser 到逻辑计划优化规则的全过程。
+
+
+
读懂状态与部署
+
理解 GraphState 快照如何支撑增量计算;用 Local 模式跑通后看 K8s 部署差异。
+
+
+
写出图算法
+
实现 VertexCentricCompute 接口,写一个自己的图算法(如 KCore 变体)并跑通。
+
+
+ + + +
+ +