|
| 1 | +# Native Kev decision inference |
| 2 | + |
| 3 | +Embed [Kev](https://huggingface.co/jaredpalmer/kev-0.8b) in a C++ application |
| 4 | +with ExecuTorch's `Module` API, using XNNPACK on CPU or MLX on Apple GPUs. |
| 5 | +Prefill a shared text once, then ask batches of questions from that snapshot. |
| 6 | +Inference uses the native binary, model, and tokenizer; Python is needed for |
| 7 | +export. |
| 8 | + |
| 9 | +[main.cpp](main.cpp) is the complete application. It loads the model and |
| 10 | +tokenizer and asks three questions about a customer message in one |
| 11 | +`system_one` call. The output contains a Choice answer with its distribution |
| 12 | +and two Noul probabilities. [benchmark.cpp](benchmark.cpp) measures explicit |
| 13 | +prefix reuse across two evaluation calls. |
| 14 | + |
| 15 | +## Export |
| 16 | + |
| 17 | +Use an ExecuTorch development environment and a local `~/kev` checkout at |
| 18 | +[d3d2f73](https://github.com/jaredpalmer/kev/tree/d3d2f73cc5f7828ae9c114c3d221ef454801b903). |
| 19 | +This example uses Kev's loader to merge LoRA in FP32, then casts the backbone to |
| 20 | +the requested precision. The pointer head, temperature scaling, and DeltaNet |
| 21 | +recurrence stay in FP32. No weight quantization is applied. |
| 22 | + |
| 23 | +```bash |
| 24 | +python -m pip install transformers==5.17.0 peft==0.21.0 |
| 25 | + |
| 26 | +hf download jaredpalmer/kev-0.8b \ |
| 27 | + --revision 54f4f8777356cd5bbbb6c6919c657f26e6f2f6d8 \ |
| 28 | + --local-dir kev-checkpoint |
| 29 | + |
| 30 | +PYTHONPATH="$HOME/kev:$PYTHONPATH" python examples/kev/export.py \ |
| 31 | + --checkpoint kev-checkpoint --backend xnnpack --dtype fp32 --output kev-cpu |
| 32 | + |
| 33 | +PYTHONPATH="$HOME/kev:$PYTHONPATH" python examples/kev/export.py \ |
| 34 | + --checkpoint kev-checkpoint --backend mlx --dtype bf16 --output kev-mlx |
| 35 | +``` |
| 36 | + |
| 37 | +Both backends support `fp32` and `bf16`. Each export contains `model.pte` and its |
| 38 | +tokenizer files. The program has two methods: `prefill` returns convolution |
| 39 | +history, DeltaNet state, and attention KV; `score` fans that prefix out across |
| 40 | +question rows and applies the pointer head to option-boundary hidden states. |
| 41 | +It preserves the checkpoint's fitted temperature and has no generation loop. |
| 42 | + |
| 43 | +The default limits are 384 prefix tokens and 1,024 tokens for the prefix plus |
| 44 | +one question. Set larger export bounds for longer documents: |
| 45 | + |
| 46 | +```bash |
| 47 | +PYTHONPATH="$HOME/kev:$PYTHONPATH" python examples/kev/export.py \ |
| 48 | + --checkpoint kev-checkpoint --backend mlx --dtype bf16 --output kev-mlx-long \ |
| 49 | + --max-prefix 2048 --max-context 4096 |
| 50 | +``` |
| 51 | + |
| 52 | +Token counts include delimiters. Larger bounds increase planned memory; |
| 53 | +exceeding a bound at inference returns an error without truncating the input. |
| 54 | +Longer contexts were not part of the checkpoint's training configuration. |
| 55 | + |
| 56 | +Constant methods record these bounds and the tokenizer IDs. `get_temperature` |
| 57 | +returns the fitted temperature as a float; `get_checkpoint_id` returns |
| 58 | +`sha256:<digest>` of the exported checkpoint's `head.pt` bytes. Call them through |
| 59 | +`Module::execute` to check the program against the checkpoint used by the app. |
| 60 | + |
| 61 | +## Build and run |
| 62 | + |
| 63 | +From the ExecuTorch root, with submodules initialized: |
| 64 | + |
| 65 | +```bash |
| 66 | +cmake -S examples/kev -B cmake-out/kev-cpu \ |
| 67 | + -DCMAKE_BUILD_TYPE=Release -DPYTHON_EXECUTABLE="$(command -v python)" |
| 68 | +cmake --build cmake-out/kev-cpu --target kev_runner --parallel 8 |
| 69 | + |
| 70 | +cmake-out/kev-cpu/kev_runner kev-cpu/model.pte kev-cpu/tokenizer.json |
| 71 | +``` |
| 72 | + |
| 73 | +For MLX, use macOS 14+ with Xcode's Metal compiler installed: |
| 74 | + |
| 75 | +```bash |
| 76 | +cmake -S examples/kev -B cmake-out/kev-mlx \ |
| 77 | + -DCMAKE_BUILD_TYPE=Release -DPYTHON_EXECUTABLE="$(command -v python)" \ |
| 78 | + -DEXECUTORCH_BUILD_MLX=ON |
| 79 | +cmake --build cmake-out/kev-mlx --target kev_runner --parallel 8 |
| 80 | + |
| 81 | +cmake-out/kev-mlx/kev_runner kev-mlx/model.pte kev-mlx/tokenizer.json |
| 82 | +``` |
| 83 | + |
| 84 | +Fresh MLX builds default to a 26.2 deployment target on macOS 26.2+ and 14.0 |
| 85 | +on older systems. With macOS SDK 26.2+ and Metal 4, the 26.2 target includes |
| 86 | +MLX's NAX matrix kernels for M5; the same binary uses standard kernels on M1–M4. |
| 87 | +To ship a build from a newer Mac to macOS 14+, set |
| 88 | +`-DCMAKE_OSX_DEPLOYMENT_TARGET=14.0`; this excludes NAX kernels. |
| 89 | +An explicit target, including `MACOSX_DEPLOYMENT_TARGET` in the environment, |
| 90 | +takes precedence. Existing CMake build directories retain their cached target; |
| 91 | +pass `-DCMAKE_OSX_DEPLOYMENT_TARGET=26.2` to update a previous 14.0 build. |
| 92 | + |
| 93 | +CMake copies `mlx.metallib` beside each MLX executable; ship it with the binary. |
| 94 | +The optional `TEXT` argument replaces the bundled customer message. |
| 95 | + |
| 96 | +The current tokenizer revision can drop accents during NFC normalization. |
| 97 | +The bundled ASCII request is unaffected; the NFC correction is being handled |
| 98 | +separately in the tokenizer library. |
| 99 | + |
| 100 | +## Benchmark |
| 101 | + |
| 102 | +Build and run `kev_benchmark` to measure the same three questions split across |
| 103 | +two evaluation calls sharing one prefix: |
| 104 | + |
| 105 | +```bash |
| 106 | +cmake --build cmake-out/kev-cpu --target kev_benchmark --parallel 8 |
| 107 | +cmake-out/kev-cpu/kev_benchmark kev-cpu/model.pte kev-cpu/tokenizer.json |
| 108 | + |
| 109 | +cmake --build cmake-out/kev-mlx --target kev_benchmark --parallel 8 |
| 110 | +cmake-out/kev-mlx/kev_benchmark kev-mlx/model.pte kev-mlx/tokenizer.json |
| 111 | +``` |
| 112 | + |
| 113 | +The benchmark loads the model once, performs two warmups, then reports median |
| 114 | +milliseconds over five runs for prefill, cached evaluation, and prefill plus |
| 115 | +evaluation. Each run creates a prefix and reuses it for both evaluation calls. |
| 116 | +Timing includes tokenization and prefix snapshot creation, excludes loading and |
| 117 | +printing, and waits for GPU results. An optional third argument replaces the |
| 118 | +bundled text. |
| 119 | + |
| 120 | +Measurements on an Apple M1 Pro (8 performance cores, 2 efficiency cores, |
| 121 | +32 GiB RAM), macOS 26.6.2, using a Release build with Apple Clang 17 and the |
| 122 | +default thread pool. Measured on September 22, 2026, with the pinned checkpoint |
| 123 | +above and the bundled customer message (19 prefix tokens including the state |
| 124 | +delimiter). Cached evaluation covers all three questions across both calls. |
| 125 | + |
| 126 | +| Median latency (ms) | XNNPACK FP32 | XNNPACK BF16 | MLX BF16 | |
| 127 | +|---|---:|---:|---:| |
| 128 | +| Prefill | 100.28 | 176.45 | 31.55 | |
| 129 | +| Cached evaluation | 457.59 | 804.16 | 86.55 | |
| 130 | +| Prefill + evaluation | 562.93 | 980.17 | 118.56 | |
| 131 | + |
| 132 | +The combined value is the median of complete runs, so it need not equal the sum |
| 133 | +of the other medians. FP32 is faster on this CPU; all exports are unquantized. |
| 134 | +To reproduce the XNNPACK BF16 column with the CPU benchmark: |
| 135 | + |
| 136 | +```bash |
| 137 | +PYTHONPATH="$HOME/kev:$PYTHONPATH" python examples/kev/export.py \ |
| 138 | + --checkpoint kev-checkpoint --backend xnnpack --dtype bf16 --output kev-cpu-bf16 |
| 139 | + |
| 140 | +cmake-out/kev-cpu/kev_benchmark kev-cpu-bf16/model.pte kev-cpu-bf16/tokenizer.json |
| 141 | +``` |
| 142 | + |
| 143 | +## C++ API |
| 144 | + |
| 145 | +[api.h](api.h) defines the question and answer types and the virtual |
| 146 | +`SystemOne::system_one(state, questions)` interface, following |
| 147 | +[TypeSafe's SDK operation](https://docs.typesafe.ai/sdk/python/api/clients/sync). |
| 148 | +It returns ExecuTorch's `Result<Answers>`. [kev.h](kev.h) provides `Kev`, |
| 149 | +which implements this interface and borrows a `Module` and tokenizer. |
| 150 | + |
| 151 | +For explicit prefix reuse, `Kev` also provides `prefill(state)` and |
| 152 | +`evaluate(prefix, questions)`. `Prefix` owns its snapshot and must be used with |
| 153 | +the `Kev` instance that created it. The instance must outlive its prefixes, and |
| 154 | +the `Module` and tokenizer must outlive the instance. Multiple prefixes may |
| 155 | +coexist; all calls must be serialized per `Module`. |
| 156 | + |
| 157 | +Each question starts from the same snapshot. Evaluation leaves it unchanged, |
| 158 | +so later calls can ask different questions. Answers own their labels and scores. |
| 159 | + |
| 160 | +The question and answer types follow [TypeSafe's typed SDK](https://docs.typesafe.ai/sdk/python/api/types/questions). |
| 161 | +`kev.cpp` handles Kev's input encoding and answer mapping. `Question` is a |
| 162 | +`std::variant<Choice, Noul, Score>`; `Answer` holds the matching answer type. |
| 163 | +`Questions` and `Answers` are ordered `(id, value)` pairs. IDs identify answers |
| 164 | +and are not sent to the model. |
| 165 | + |
| 166 | +```cpp |
| 167 | +kev::Questions questions{ |
| 168 | + {"department", kev::Choice{"Which team should handle this?", |
| 169 | + {{"billing", "Invoices and refunds"}, {"technical", "Bugs and outages"}}}}, |
| 170 | + {"refund", kev::Noul{"Is a refund requested?", |
| 171 | + {{kev::NoulOutcome::False, "No refund requested"}, |
| 172 | + {kev::NoulOutcome::True, "Explicitly asks for a refund"}}}}, |
| 173 | + {"urgency", kev::Score{"How urgent is this?", |
| 174 | + {"Can wait", "This week", "Today"}}}}; |
| 175 | +``` |
| 176 | +
|
| 177 | +With a loaded module and tokenizer, call through the interface: |
| 178 | +
|
| 179 | +```cpp |
| 180 | +kev::Kev model(module, tokenizer); |
| 181 | +kev::SystemOne& api = model; |
| 182 | +auto answers = api.system_one(state, questions); |
| 183 | +``` |
| 184 | + |
| 185 | +Each `system_one` call prefills the supplied state and evaluates its questions. |
| 186 | +For multiple requests about the same state, reuse a prefix: |
| 187 | + |
| 188 | +```cpp |
| 189 | +auto prefix = model.prefill(state); |
| 190 | +if (!prefix.ok()) { |
| 191 | + return 1; |
| 192 | +} |
| 193 | +auto answers = model.evaluate(*prefix, questions); |
| 194 | +if (!answers.ok()) { |
| 195 | + return 1; |
| 196 | +} |
| 197 | +auto followup = model.evaluate(*prefix, { |
| 198 | + {"duplicate", kev::Noul{"Was the customer charged more than once?", {}}}}); |
| 199 | +if (!followup.ok()) { |
| 200 | + return 1; |
| 201 | +} |
| 202 | +``` |
| 203 | + |
| 204 | +Both calls use the same unchanged snapshot. See [benchmark.cpp](benchmark.cpp) |
| 205 | +for a complete example with timing. |
| 206 | + |
| 207 | +Text fields accept UTF-8 strings; structured content can be rendered to text |
| 208 | +before calling this native interface. Instructions may be empty. Choice |
| 209 | +criteria preserve caller order and have unique names with optional descriptions |
| 210 | +(`std::nullopt` for none). Noul criteria are optional descriptions keyed by |
| 211 | +`NoulOutcome::False` and `NoulOutcome::True`, corresponding to no and yes. |
| 212 | +Score criteria are ordered descriptions: repeated and empty levels are allowed |
| 213 | +and retain their separate indices. |
| 214 | + |
| 215 | +`ChoiceAnswer` contains `choice`, `probabilities` by name, and `confidence`. |
| 216 | +`NoulAnswer::noul` is the probability of yes. `ScoreAnswer` contains the expected |
| 217 | +zero-based `score`, ordered `legend` and `probabilities` vectors, and |
| 218 | +`confidence`. Each vector index identifies the corresponding level in |
| 219 | +`Score::criteria`; for example, `answer.probabilities.at(0)` gives the first |
| 220 | +level's probability. Values retain full precision. Confidence uses Kev's |
| 221 | +formulas; exact equivalence to TypeSafe is not established. |
| 222 | + |
| 223 | +`evaluate` accepts any nonempty list of questions with unique IDs and splits it |
| 224 | +into batches of up to eight, reusing the same prefix. Answers retain request |
| 225 | +order. Each batch is padded to its own longest question and criterion count. |
| 226 | +Choice and Score accept 1–255 criteria, following upstream Kev; TypeSafe's |
| 227 | +hosted API documents 2–10 Score levels. |
0 commit comments