Skip to content

Commit e5acd53

Browse files
authored
Add native Kev decision inference on CPU and GPU (pytorch#23023)
Part of pytorch#22985 Run Kev prefill and pointer-head scoring natively through ExecuTorch Module, using XNNPACK on CPU or MLX on Apple GPUs with FP32 or unquantized BF16 weights. Applications submit named Choice, Noul, and Score questions through system_one; Kev also exposes an owned prefix snapshot for reuse across evaluation calls. Keep the example self-contained in examples/kev. Review api.h and main.cpp for the typed application contract, then kev.h and kev.cpp for the Kev adapter and prefix API. model.py and export.py contain model-specific export code, including configurable context bounds and constant temperature and checkpoint metadata. Requests exceeding the exported question batch limit are processed in chunks. A separate kev_benchmark executable measures prefill, cached evaluation, and combined latency. The README documents CPU and MLX builds and measurements. Validated native XNNPACK and MLX FP32/BF16 on 1,024 records (1,264 questions) against upstream Kev's FP32 PyTorch model. Probability checks pass upstream thresholds on matching-token inputs; targeted CPU checks also cover the 29 questions with tokenizer differences. Upstream's 13 CPU tests pass in each precision, alongside native API, prefix lifecycle, bounds, and metadata checks. These runs used the prior tokenizer pin. End-to-end tokenizer parity needs revalidation after rebuilding with the landed tokenizer update (3fc1df7). Authored with OpenAI Codex.
1 parent 2c07d2b commit e5acd53

9 files changed

Lines changed: 1354 additions & 0 deletions

File tree

‎examples/kev/CMakeLists.txt‎

Lines changed: 81 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,81 @@
1+
# Copyright (c) Meta Platforms, Inc. and affiliates.
2+
# All rights reserved.
3+
#
4+
# This source code is licensed under the BSD-style license found in the
5+
# LICENSE file in the root directory of this source tree.
6+
7+
cmake_minimum_required(VERSION 3.24)
8+
if(EXECUTORCH_BUILD_MLX
9+
AND CMAKE_HOST_APPLE
10+
AND NOT DEFINED CMAKE_OSX_DEPLOYMENT_TARGET
11+
AND NOT DEFINED ENV{MACOSX_DEPLOYMENT_TARGET}
12+
)
13+
cmake_host_system_information(RESULT _kev_macos_version QUERY OS_RELEASE)
14+
set(_kev_deployment_target 14.0)
15+
if(_kev_macos_version VERSION_GREATER_EQUAL 26.2)
16+
set(_kev_deployment_target 26.2)
17+
endif()
18+
set(CMAKE_OSX_DEPLOYMENT_TARGET
19+
${_kev_deployment_target}
20+
CACHE STRING "Minimum macOS version"
21+
)
22+
endif()
23+
project(kev LANGUAGES C CXX)
24+
25+
set(CMAKE_CXX_STANDARD 20)
26+
set(CMAKE_CXX_STANDARD_REQUIRED ON)
27+
get_filename_component(EXECUTORCH_ROOT "../.." ABSOLUTE)
28+
29+
option(BUILD_SHARED_LIBS "" OFF)
30+
option(EXECUTORCH_BUILD_EXECUTOR_RUNNER "" OFF)
31+
option(EXECUTORCH_BUILD_EXTENSION_DATA_LOADER "" ON)
32+
option(EXECUTORCH_BUILD_EXTENSION_MODULE "" ON)
33+
option(EXECUTORCH_BUILD_EXTENSION_FLAT_TENSOR "" ON)
34+
option(EXECUTORCH_BUILD_EXTENSION_NAMED_DATA_MAP "" ON)
35+
option(EXECUTORCH_BUILD_EXTENSION_TENSOR "" ON)
36+
option(EXECUTORCH_BUILD_KERNELS_OPTIMIZED "" ON)
37+
option(EXECUTORCH_BUILD_KERNELS_LLM "" ON)
38+
option(EXECUTORCH_BUILD_XNNPACK "" ON)
39+
option(SUPPORT_REGEX_LOOKAHEAD "" ON)
40+
if(EXECUTORCH_BUILD_MLX)
41+
set(EXECUTORCH_BUILD_EXTENSION_LLM
42+
ON
43+
CACHE BOOL "" FORCE
44+
)
45+
endif()
46+
47+
add_subdirectory(${EXECUTORCH_ROOT} executorch)
48+
if(NOT TARGET tokenizers::tokenizers)
49+
add_subdirectory(${EXECUTORCH_ROOT}/extension/llm/tokenizers tokenizers)
50+
endif()
51+
52+
add_library(kev STATIC kev.cpp)
53+
target_link_libraries(
54+
kev
55+
PUBLIC extension_module_static extension_tensor tokenizers::tokenizers
56+
PRIVATE executorch optimized_native_cpu_ops_lib custom_ops
57+
)
58+
include(${EXECUTORCH_ROOT}/tools/cmake/Utils.cmake)
59+
executorch_target_link_options_shared_lib(optimized_native_cpu_ops_lib)
60+
executorch_target_link_options_shared_lib(custom_ops)
61+
if(EXECUTORCH_BUILD_XNNPACK)
62+
target_link_libraries(kev PRIVATE xnnpack_backend)
63+
executorch_target_link_options_shared_lib(xnnpack_backend)
64+
endif()
65+
if(EXECUTORCH_BUILD_MLX)
66+
target_link_libraries(kev PRIVATE mlxdelegate)
67+
endif()
68+
69+
add_executable(kev_runner main.cpp)
70+
add_executable(kev_benchmark benchmark.cpp)
71+
foreach(target kev_runner kev_benchmark)
72+
target_link_libraries(${target} PRIVATE kev)
73+
if(EXECUTORCH_BUILD_MLX)
74+
add_custom_command(
75+
TARGET ${target}
76+
POST_BUILD
77+
COMMAND ${CMAKE_COMMAND} -E copy_if_different ${MLX_METALLIB_PATH}
78+
$<TARGET_FILE_DIR:${target}>/mlx.metallib
79+
)
80+
endif()
81+
endforeach()

‎examples/kev/README.md‎

Lines changed: 227 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,227 @@
1+
# Native Kev decision inference
2+
3+
Embed [Kev](https://huggingface.co/jaredpalmer/kev-0.8b) in a C++ application
4+
with ExecuTorch's `Module` API, using XNNPACK on CPU or MLX on Apple GPUs.
5+
Prefill a shared text once, then ask batches of questions from that snapshot.
6+
Inference uses the native binary, model, and tokenizer; Python is needed for
7+
export.
8+
9+
[main.cpp](main.cpp) is the complete application. It loads the model and
10+
tokenizer and asks three questions about a customer message in one
11+
`system_one` call. The output contains a Choice answer with its distribution
12+
and two Noul probabilities. [benchmark.cpp](benchmark.cpp) measures explicit
13+
prefix reuse across two evaluation calls.
14+
15+
## Export
16+
17+
Use an ExecuTorch development environment and a local `~/kev` checkout at
18+
[d3d2f73](https://github.com/jaredpalmer/kev/tree/d3d2f73cc5f7828ae9c114c3d221ef454801b903).
19+
This example uses Kev's loader to merge LoRA in FP32, then casts the backbone to
20+
the requested precision. The pointer head, temperature scaling, and DeltaNet
21+
recurrence stay in FP32. No weight quantization is applied.
22+
23+
```bash
24+
python -m pip install transformers==5.17.0 peft==0.21.0
25+
26+
hf download jaredpalmer/kev-0.8b \
27+
--revision 54f4f8777356cd5bbbb6c6919c657f26e6f2f6d8 \
28+
--local-dir kev-checkpoint
29+
30+
PYTHONPATH="$HOME/kev:$PYTHONPATH" python examples/kev/export.py \
31+
--checkpoint kev-checkpoint --backend xnnpack --dtype fp32 --output kev-cpu
32+
33+
PYTHONPATH="$HOME/kev:$PYTHONPATH" python examples/kev/export.py \
34+
--checkpoint kev-checkpoint --backend mlx --dtype bf16 --output kev-mlx
35+
```
36+
37+
Both backends support `fp32` and `bf16`. Each export contains `model.pte` and its
38+
tokenizer files. The program has two methods: `prefill` returns convolution
39+
history, DeltaNet state, and attention KV; `score` fans that prefix out across
40+
question rows and applies the pointer head to option-boundary hidden states.
41+
It preserves the checkpoint's fitted temperature and has no generation loop.
42+
43+
The default limits are 384 prefix tokens and 1,024 tokens for the prefix plus
44+
one question. Set larger export bounds for longer documents:
45+
46+
```bash
47+
PYTHONPATH="$HOME/kev:$PYTHONPATH" python examples/kev/export.py \
48+
--checkpoint kev-checkpoint --backend mlx --dtype bf16 --output kev-mlx-long \
49+
--max-prefix 2048 --max-context 4096
50+
```
51+
52+
Token counts include delimiters. Larger bounds increase planned memory;
53+
exceeding a bound at inference returns an error without truncating the input.
54+
Longer contexts were not part of the checkpoint's training configuration.
55+
56+
Constant methods record these bounds and the tokenizer IDs. `get_temperature`
57+
returns the fitted temperature as a float; `get_checkpoint_id` returns
58+
`sha256:<digest>` of the exported checkpoint's `head.pt` bytes. Call them through
59+
`Module::execute` to check the program against the checkpoint used by the app.
60+
61+
## Build and run
62+
63+
From the ExecuTorch root, with submodules initialized:
64+
65+
```bash
66+
cmake -S examples/kev -B cmake-out/kev-cpu \
67+
-DCMAKE_BUILD_TYPE=Release -DPYTHON_EXECUTABLE="$(command -v python)"
68+
cmake --build cmake-out/kev-cpu --target kev_runner --parallel 8
69+
70+
cmake-out/kev-cpu/kev_runner kev-cpu/model.pte kev-cpu/tokenizer.json
71+
```
72+
73+
For MLX, use macOS 14+ with Xcode's Metal compiler installed:
74+
75+
```bash
76+
cmake -S examples/kev -B cmake-out/kev-mlx \
77+
-DCMAKE_BUILD_TYPE=Release -DPYTHON_EXECUTABLE="$(command -v python)" \
78+
-DEXECUTORCH_BUILD_MLX=ON
79+
cmake --build cmake-out/kev-mlx --target kev_runner --parallel 8
80+
81+
cmake-out/kev-mlx/kev_runner kev-mlx/model.pte kev-mlx/tokenizer.json
82+
```
83+
84+
Fresh MLX builds default to a 26.2 deployment target on macOS 26.2+ and 14.0
85+
on older systems. With macOS SDK 26.2+ and Metal 4, the 26.2 target includes
86+
MLX's NAX matrix kernels for M5; the same binary uses standard kernels on M1–M4.
87+
To ship a build from a newer Mac to macOS 14+, set
88+
`-DCMAKE_OSX_DEPLOYMENT_TARGET=14.0`; this excludes NAX kernels.
89+
An explicit target, including `MACOSX_DEPLOYMENT_TARGET` in the environment,
90+
takes precedence. Existing CMake build directories retain their cached target;
91+
pass `-DCMAKE_OSX_DEPLOYMENT_TARGET=26.2` to update a previous 14.0 build.
92+
93+
CMake copies `mlx.metallib` beside each MLX executable; ship it with the binary.
94+
The optional `TEXT` argument replaces the bundled customer message.
95+
96+
The current tokenizer revision can drop accents during NFC normalization.
97+
The bundled ASCII request is unaffected; the NFC correction is being handled
98+
separately in the tokenizer library.
99+
100+
## Benchmark
101+
102+
Build and run `kev_benchmark` to measure the same three questions split across
103+
two evaluation calls sharing one prefix:
104+
105+
```bash
106+
cmake --build cmake-out/kev-cpu --target kev_benchmark --parallel 8
107+
cmake-out/kev-cpu/kev_benchmark kev-cpu/model.pte kev-cpu/tokenizer.json
108+
109+
cmake --build cmake-out/kev-mlx --target kev_benchmark --parallel 8
110+
cmake-out/kev-mlx/kev_benchmark kev-mlx/model.pte kev-mlx/tokenizer.json
111+
```
112+
113+
The benchmark loads the model once, performs two warmups, then reports median
114+
milliseconds over five runs for prefill, cached evaluation, and prefill plus
115+
evaluation. Each run creates a prefix and reuses it for both evaluation calls.
116+
Timing includes tokenization and prefix snapshot creation, excludes loading and
117+
printing, and waits for GPU results. An optional third argument replaces the
118+
bundled text.
119+
120+
Measurements on an Apple M1 Pro (8 performance cores, 2 efficiency cores,
121+
32 GiB RAM), macOS 26.6.2, using a Release build with Apple Clang 17 and the
122+
default thread pool. Measured on September 22, 2026, with the pinned checkpoint
123+
above and the bundled customer message (19 prefix tokens including the state
124+
delimiter). Cached evaluation covers all three questions across both calls.
125+
126+
| Median latency (ms) | XNNPACK FP32 | XNNPACK BF16 | MLX BF16 |
127+
|---|---:|---:|---:|
128+
| Prefill | 100.28 | 176.45 | 31.55 |
129+
| Cached evaluation | 457.59 | 804.16 | 86.55 |
130+
| Prefill + evaluation | 562.93 | 980.17 | 118.56 |
131+
132+
The combined value is the median of complete runs, so it need not equal the sum
133+
of the other medians. FP32 is faster on this CPU; all exports are unquantized.
134+
To reproduce the XNNPACK BF16 column with the CPU benchmark:
135+
136+
```bash
137+
PYTHONPATH="$HOME/kev:$PYTHONPATH" python examples/kev/export.py \
138+
--checkpoint kev-checkpoint --backend xnnpack --dtype bf16 --output kev-cpu-bf16
139+
140+
cmake-out/kev-cpu/kev_benchmark kev-cpu-bf16/model.pte kev-cpu-bf16/tokenizer.json
141+
```
142+
143+
## C++ API
144+
145+
[api.h](api.h) defines the question and answer types and the virtual
146+
`SystemOne::system_one(state, questions)` interface, following
147+
[TypeSafe's SDK operation](https://docs.typesafe.ai/sdk/python/api/clients/sync).
148+
It returns ExecuTorch's `Result<Answers>`. [kev.h](kev.h) provides `Kev`,
149+
which implements this interface and borrows a `Module` and tokenizer.
150+
151+
For explicit prefix reuse, `Kev` also provides `prefill(state)` and
152+
`evaluate(prefix, questions)`. `Prefix` owns its snapshot and must be used with
153+
the `Kev` instance that created it. The instance must outlive its prefixes, and
154+
the `Module` and tokenizer must outlive the instance. Multiple prefixes may
155+
coexist; all calls must be serialized per `Module`.
156+
157+
Each question starts from the same snapshot. Evaluation leaves it unchanged,
158+
so later calls can ask different questions. Answers own their labels and scores.
159+
160+
The question and answer types follow [TypeSafe's typed SDK](https://docs.typesafe.ai/sdk/python/api/types/questions).
161+
`kev.cpp` handles Kev's input encoding and answer mapping. `Question` is a
162+
`std::variant<Choice, Noul, Score>`; `Answer` holds the matching answer type.
163+
`Questions` and `Answers` are ordered `(id, value)` pairs. IDs identify answers
164+
and are not sent to the model.
165+
166+
```cpp
167+
kev::Questions questions{
168+
{"department", kev::Choice{"Which team should handle this?",
169+
{{"billing", "Invoices and refunds"}, {"technical", "Bugs and outages"}}}},
170+
{"refund", kev::Noul{"Is a refund requested?",
171+
{{kev::NoulOutcome::False, "No refund requested"},
172+
{kev::NoulOutcome::True, "Explicitly asks for a refund"}}}},
173+
{"urgency", kev::Score{"How urgent is this?",
174+
{"Can wait", "This week", "Today"}}}};
175+
```
176+
177+
With a loaded module and tokenizer, call through the interface:
178+
179+
```cpp
180+
kev::Kev model(module, tokenizer);
181+
kev::SystemOne& api = model;
182+
auto answers = api.system_one(state, questions);
183+
```
184+
185+
Each `system_one` call prefills the supplied state and evaluates its questions.
186+
For multiple requests about the same state, reuse a prefix:
187+
188+
```cpp
189+
auto prefix = model.prefill(state);
190+
if (!prefix.ok()) {
191+
return 1;
192+
}
193+
auto answers = model.evaluate(*prefix, questions);
194+
if (!answers.ok()) {
195+
return 1;
196+
}
197+
auto followup = model.evaluate(*prefix, {
198+
{"duplicate", kev::Noul{"Was the customer charged more than once?", {}}}});
199+
if (!followup.ok()) {
200+
return 1;
201+
}
202+
```
203+
204+
Both calls use the same unchanged snapshot. See [benchmark.cpp](benchmark.cpp)
205+
for a complete example with timing.
206+
207+
Text fields accept UTF-8 strings; structured content can be rendered to text
208+
before calling this native interface. Instructions may be empty. Choice
209+
criteria preserve caller order and have unique names with optional descriptions
210+
(`std::nullopt` for none). Noul criteria are optional descriptions keyed by
211+
`NoulOutcome::False` and `NoulOutcome::True`, corresponding to no and yes.
212+
Score criteria are ordered descriptions: repeated and empty levels are allowed
213+
and retain their separate indices.
214+
215+
`ChoiceAnswer` contains `choice`, `probabilities` by name, and `confidence`.
216+
`NoulAnswer::noul` is the probability of yes. `ScoreAnswer` contains the expected
217+
zero-based `score`, ordered `legend` and `probabilities` vectors, and
218+
`confidence`. Each vector index identifies the corresponding level in
219+
`Score::criteria`; for example, `answer.probabilities.at(0)` gives the first
220+
level's probability. Values retain full precision. Confidence uses Kev's
221+
formulas; exact equivalence to TypeSafe is not established.
222+
223+
`evaluate` accepts any nonempty list of questions with unique IDs and splits it
224+
into batches of up to eight, reusing the same prefix. Answers retain request
225+
order. Each batch is padded to its own longest question and criterion count.
226+
Choice and Score accept 1–255 criteria, following upstream Kev; TypeSafe's
227+
hosted API documents 2–10 Score levels.

‎examples/kev/api.h‎

Lines changed: 73 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,73 @@
1+
/*
2+
* Copyright (c) Meta Platforms, Inc. and affiliates.
3+
* All rights reserved.
4+
*
5+
* This source code is licensed under the BSD-style license found in the
6+
* LICENSE file in the root directory of this source tree.
7+
*/
8+
9+
#pragma once
10+
11+
#include <executorch/runtime/core/result.h>
12+
#include <map>
13+
#include <optional>
14+
#include <string>
15+
#include <utility>
16+
#include <variant>
17+
#include <vector>
18+
19+
namespace kev {
20+
21+
struct Choice {
22+
std::string instructions;
23+
std::vector<std::pair<std::string, std::optional<std::string>>> criteria;
24+
};
25+
26+
enum class NoulOutcome {
27+
False,
28+
True,
29+
};
30+
31+
struct Noul {
32+
std::string instructions;
33+
std::map<NoulOutcome, std::string> criteria;
34+
};
35+
36+
struct Score {
37+
std::string instructions;
38+
std::vector<std::string> criteria;
39+
};
40+
41+
using Question = std::variant<Choice, Noul, Score>;
42+
using Questions = std::vector<std::pair<std::string, Question>>;
43+
44+
struct ChoiceAnswer {
45+
std::string choice;
46+
std::map<std::string, double> probabilities;
47+
double confidence;
48+
};
49+
50+
struct NoulAnswer {
51+
double noul;
52+
};
53+
54+
struct ScoreAnswer {
55+
double score;
56+
std::vector<std::string> legend;
57+
std::vector<double> probabilities;
58+
double confidence;
59+
};
60+
61+
using Answer = std::variant<ChoiceAnswer, NoulAnswer, ScoreAnswer>;
62+
using Answers = std::vector<std::pair<std::string, Answer>>;
63+
64+
class SystemOne {
65+
public:
66+
virtual ~SystemOne() = default;
67+
68+
virtual executorch::runtime::Result<Answers> system_one(
69+
const std::string& state,
70+
const Questions& questions) = 0;
71+
};
72+
73+
} // namespace kev

0 commit comments

Comments
 (0)