feat(llma): add typed offline evaluations SDK - #1005
Radu-Raicea wants to merge 2 commits into
Conversation
posthog-python Compliance ReportDate: 2026-10-02T17:09:24.431900+00:00 ✅ All Tests Passed!121/121 tests passed Capture_V1 Tests✅ 95/95 tests passed View Details
Capture_Ai Tests✅ 5/5 tests passed View Details
Feature_Flags Tests✅ 17/17 tests passed View Details
Feature_Flags_Local_Evaluation Tests✅ 4/4 tests passed View Details
|
|
[Medium risk] Adds a new offline evaluations SDK module. The PR appears safe to merge; the remaining unchecked scorer-version return type is non-blocking. Reviews (2) · Last reviewed commit: "fix(llma): reject integer scores that lo..." |
| def create_version( | ||
| self, | ||
| scorer_id: str | UUID, | ||
| *, | ||
| config: _ConfigT, | ||
| base_version: int | None = None, |
There was a problem hiding this comment.
Version return type is unchecked
create_version() infers its return type from the supplied config without checking the target scorer’s kind. Passing a numeric config for a boolean scorer can leave callers with a value typed as Scorer[NumericScorerConfig] even if the operation is rejected or returns a boolean scorer. The async method has the same issue. Tie the config type to the scorer kind instead of relying on a cast.
Prompt To Fix With AI
This is a comment left during a code review.
Path: posthog/ai/evaluations/_scorers.py
Line: 367-372
Comment:
**Version return type is unchecked** `create_version()` infers its return type from the supplied config without checking the target scorer’s kind. Passing a numeric config for a boolean scorer can leave callers with a value typed as `Scorer[NumericScorerConfig]` even if the operation is rejected or returns a boolean scorer. The async method has the same issue. Tie the config type to the scorer kind instead of relying on a cast.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.
💡 Motivation and Context
Add
posthog.ai.evaluationsso Python callers can upload offline evaluation results and manage experiments through typed synchronous and asynchronous clients. Includes acknowledged bulk uploads, stable identities for safe retries and recovery, and explicit completion/failure.Scorer management includes typed boolean, numeric, and categorical configurations with polarity, metadata updates, immutable version creation, and programmatic version-ID discovery. Includes usage examples and a minor changeset.
💚 How did you test it?
📝 Checklist
If releasing new changes
.sampo/changesets/.🤖 Agent context
Autonomy: Human-driven (agent-assisted)
Implemented with Codex and collaborating agents using shell/Python tooling, Git, and the GitHub CLI. Local Codex session.
The API follows the current backend contract. Configuration types expose supported fields; scorer mutations avoid automatic retries, while replay-safe offline operations retain identities and content.