Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 5 additions & 3 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -171,9 +171,11 @@ claim.
The main eval set should stay focused on realistic tasks where context should improve stream
quality. It must include natural activation prompts and explicit invocation prompts. Natural
scenarios must not mention `$java-streams` or ask to use the skill. Explicit scenarios may name the
skill and must be labeled as explicit in `criteria.json`.
skill and must be labeled as explicit in `criteria-meta.json`.

Every scenario directory must contain `task.md`, `criteria.json`, and `capability.txt`. Main eval
Every scenario directory must contain `task.md`, `criteria.json`, `criteria-meta.json`, and
`capability.txt` (`criteria.json` keeps Tessl's schema; `criteria-meta.json` carries this
repository's metadata and per-item categories). Main eval
implementation criteria must include compile/artifact checks and behavior correctness checks as
safety checks, but the main score should mainly measure stream-specific quality. Each main eval
criterion must set `category` to `safety`, `stream_quality`, or `maintainability`.
Expand Down Expand Up @@ -221,7 +223,7 @@ current benchmark claims until they are rerun against the current active suite m
denominator, commit/ref, natural/explicit split, and pinned CLI behavior. Main eval weights should
stay evidence-weighted: put more points on scenario families with larger observed missed-point
reduction, keep ordinary 100-point main scenarios around 15 safety / 80 stream-quality / 5
maintainability points, and document any `main_eval_weight_multiplier` in `criteria.json` metadata.
maintainability points, and document any `main_eval_weight_multiplier` in `criteria-meta.json`.

When with-context is below 100%, keep the scenario wherever it already lives. Fix the skill or eval
there, then rerun only that targeted scenario until it is clean before running broader suites. After
Expand Down
13 changes: 9 additions & 4 deletions docs/agents/evals.md
Original file line number Diff line number Diff line change
Expand Up @@ -51,8 +51,12 @@ benchmark claims, or scoring rules.
with-context and without-context, plus skill-context-dependent checks that are only fair as
with-context regression coverage. These protect against regressions but should not be part of
normal lift discovery runs.
- Every scenario directory must contain `task.md`, `criteria.json`, and `capability.txt`.
- Every `criteria.json` must classify `metadata.invocation` and `metadata.task_type`.
- Every scenario directory must contain `task.md`, `criteria.json`, `criteria-meta.json`, and
`capability.txt`. `criteria.json` holds only what Tessl's schema knows (`context`, `type`, and
checklist items with `name`, `description`, `max_score`), so `tessl eval lint` stays clean;
`criteria-meta.json` holds this repository's `metadata` object and a `categories` map from
checklist name to category. The validators read the merged view.
- Every `criteria-meta.json` must classify `metadata.invocation` and `metadata.task_type`.
- Use `metadata.evidence_type` when scenario placement needs to be explicit:
- `ordinary_lift`: an ordinary main or reference scenario where both variants are fair to compare.
This value is invalid in `evals-regression/`, and it must not be used when the task overlaps
Expand Down Expand Up @@ -138,7 +142,7 @@ benchmark claims, or scoring rules.
- Normalize ordinary 100-point main scenarios around 15 safety, 80 stream-quality, and 5
maintainability points unless the scenario has a documented reason to differ.
- Use `main_eval_weight_multiplier` only when a scenario family has stronger hosted delta or higher
benchmark importance; document why in `criteria.json` metadata and this file.
benchmark importance; document why in `criteria-meta.json` and this file.
- Do not add or inflate weak-delta scenarios only to make coverage look balanced.
- A 2x raw score ratio is useful only when earned by honest, realistic eval design. Don't suppress
legitimate coverage just to improve lift.
Expand Down Expand Up @@ -184,7 +188,8 @@ benchmark claims, or scoring rules.
not the final all-suite requirement itself; it is rerunning required evidence after later edits to
`skills/java-streams/SKILL.md` or bundled runtime references change the skill fingerprint. Do the
local scenario/criteria crosswalk and obvious skill wording fixes before starting hosted runs.
- A pure suite move does not require a hosted rerun when `task.md`, `criteria.json`, and
- A pure suite move does not require a hosted rerun when `task.md`, `criteria.json`,
`criteria-meta.json`, and
`capability.txt` content are unchanged except for suite-placement metadata or numbering notes.
Run local validators and update suite totals/numbering instead. If the move also changes task
wording, scoring criteria, capability text, runtime skill behavior, or benchmark claims, follow
Expand Down
16 changes: 16 additions & 0 deletions evals-reference/05-parallel-cpu-review/criteria-meta.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
{
"metadata": {
"invocation": "explicit",
"task_type": "review",
"evidence_type": "focused_reference",
"reference_selection": "Focused reference coverage for CPU-heavy stateless parallel stream advice.",
"runtime_reference_overlap_rationale": "Allowed only because this is reference-suite focused coverage, not ordinary broad lift."
},
"categories": {
"Creates review artifact": "safety",
"Allows parallelism for this stream chain": "stream_quality",
"Mentions overhead and measurement": "stream_quality",
"Mentions common pool": "stream_quality",
"Avoids multi-line lambda examples": "maintainability"
}
}
14 changes: 1 addition & 13 deletions evals-reference/05-parallel-cpu-review/criteria.json
Original file line number Diff line number Diff line change
Expand Up @@ -4,40 +4,28 @@
"checklist": [
{
"name": "Creates review artifact",
"category": "safety",
"max_score": 5,
"description": "Creates review.md with a clear review decision."
},
{
"name": "Allows parallelism for this stream chain",
"category": "stream_quality",
"max_score": 20,
"description": "Recognizes the work is CPU-heavy, stateless, and has a primitive sum reduction, so parallelism can be reasonable."
},
{
"name": "Mentions overhead and measurement",
"category": "stream_quality",
"max_score": 15,
"description": "Says the performance should be measured because parallel stream overhead can outweigh benefits."
},
{
"name": "Mentions common pool",
"category": "stream_quality",
"max_score": 10,
"description": "Notes that parallel streams use the common fork-join pool by default and may contend with other work."
},
{
"name": "Avoids multi-line lambda examples",
"category": "maintainability",
"max_score": 10,
"description": "If the review includes replacement or optional example code, keeps lambda bodies on the same line as `->` or uses named helpers/method references instead of showing block lambdas, arrows whose body starts on the next line, or callback bodies that continue on later lines."
}
],
"metadata": {
"invocation": "explicit",
"task_type": "review",
"evidence_type": "focused_reference",
"reference_selection": "Focused reference coverage for CPU-heavy stateless parallel stream advice.",
"runtime_reference_overlap_rationale": "Allowed only because this is reference-suite focused coverage, not ordinary broad lift."
}
]
}
17 changes: 17 additions & 0 deletions evals-reference/08-primary-contact-review/criteria-meta.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
{
"metadata": {
"invocation": "natural",
"task_type": "review",
"evidence_type": "focused_reference",
"reference_selection": "Focused reference coverage for preserving priority ordering when reviewing findFirst versus findAny changes.",
"runtime_reference_overlap_rationale": "Allowed only because this is reference-suite focused coverage, not ordinary broad lift."
},
"categories": {
"Creates review artifact": "safety",
"Rejects behavior-changing proposal": "safety",
"Identifies priority ordering contract": "stream_quality",
"Flags unjustified parallelStream": "stream_quality",
"Suggests safe stream chain": "stream_quality",
"Review is concise and focused": "maintainability"
}
}
15 changes: 1 addition & 14 deletions evals-reference/08-primary-contact-review/criteria.json
Original file line number Diff line number Diff line change
Expand Up @@ -4,46 +4,33 @@
"checklist": [
{
"name": "Creates review artifact",
"category": "safety",
"max_score": 5,
"description": "Creates review.md with a clear review decision about the proposed change."
},
{
"name": "Rejects behavior-changing proposal",
"category": "safety",
"max_score": 10,
"description": "Clearly says the proposed change should not be accepted because it can return a different contact."
},
{
"name": "Identifies priority ordering contract",
"category": "stream_quality",
"max_score": 35,
"description": "Explains that sorted by Contact::priority followed by findFirst means the lowest-priority verified contact wins, and findAny does not preserve that contract."
},
{
"name": "Flags unjustified parallelStream",
"category": "stream_quality",
"max_score": 25,
"description": "Explains that parallelStream is not justified here and makes ordered selection less predictable or more expensive without evidence of CPU-bound work."
},
{
"name": "Suggests safe stream chain",
"category": "stream_quality",
"max_score": 20,
"description": "Suggests keeping sorted(...).filter(...).findFirst(), or filtering before sorting only if that preserves the same lowest-priority verified contact behavior."
},
{
"name": "Review is concise and focused",
"category": "maintainability",
"max_score": 5,
"description": "Keeps the review focused on stream ordering and parallelism without internal workflow labels such as hard stop, marker, scan, checklist, skill, rubric, or criteria. Award full credit for a brief note that the original chain can also be simplified only if it supports the decision and does not distract from it."
}
],
"metadata": {
"invocation": "natural",
"task_type": "review",
"evidence_type": "focused_reference",
"reference_selection": "Focused reference coverage for preserving priority ordering when reviewing findFirst versus findAny changes.",
"runtime_reference_overlap_rationale": "Allowed only because this is reference-suite focused coverage, not ordinary broad lift."
}
]
}
19 changes: 19 additions & 0 deletions evals-reference/15-session-roster-indexes/criteria-meta.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
{
"metadata": {
"invocation": "natural",
"task_type": "implementation",
"evidence_type": "focused_reference",
"reference_selection": "Focused reference coverage for nested session-registration collector behavior and helper extraction.",
"runtime_reference_overlap_rationale": "Allowed only because this is reference-suite focused coverage, not ordinary broad lift."
},
"categories": {
"Creates coherent Java 17 artifact": "safety",
"Preserves requested roster and index behavior": "safety",
"Flattens nested conference data directly": "stream_quality",
"Uses collector semantics for track email lists": "stream_quality",
"Avoids multi-line lambdas": "maintainability",
"Handles duplicate rooms explicitly": "stream_quality",
"Short-circuits waitlisted existence": "stream_quality",
"Keeps implementation focused": "maintainability"
}
}
17 changes: 1 addition & 16 deletions evals-reference/15-session-roster-indexes/criteria.json
Original file line number Diff line number Diff line change
Expand Up @@ -4,58 +4,43 @@
"checklist": [
{
"name": "Creates coherent Java 17 artifact",
"category": "safety",
"max_score": 5,
"description": "Creates SessionRosterIndexes.java with the requested methods, nested records, imports, and Java 17-compatible code."
},
{
"name": "Preserves requested roster and index behavior",
"category": "safety",
"max_score": 10,
"description": "Correctly handles null tracks, null rooms, null emails, opted-in filtering, duplicate emails, sorted email lists, room duplicate handling, encounter-order ties, and waitlisted existence."
},
{
"name": "Flattens nested conference data directly",
"category": "stream_quality",
"max_score": 20,
"description": "Traverses conferences to sessions to registrations with direct flatMap-style composition or an equivalently clear stream chain, without building nested temporary collections or mutating shared result maps inside forEach."
},
{
"name": "Uses collector semantics for track email lists",
"category": "stream_quality",
"max_score": 20,
"description": "Builds opted-in emails by track with collector semantics such as groupingBy plus downstream mapping/filtering/collection, producing sorted unique email lists without a post-hoc shared mutable accumulation pass."
},
{
"name": "Avoids multi-line lambdas",
"category": "maintainability",
"max_score": 10,
"description": "Keeps stream and collector lambdas as short same-line glue. Extracts nested registration streams, duplicate-room comparison logic, or other multi-step callback bodies into named helpers or method references instead of using block lambdas, arrows whose body starts on the next line, or lambda bodies that continue their stream chain on later lines."
},
{
"name": "Handles duplicate rooms explicitly",
"category": "stream_quality",
"max_score": 20,
"description": "Builds longestSessionByRoom with an explicit duplicate-room merge or grouping reduction that keeps the longer session and preserves the earlier encounter-order session on ties."
},
{
"name": "Short-circuits waitlisted existence",
"category": "stream_quality",
"max_score": 10,
"description": "Uses anyMatch or an equivalent early-exit check for hasWaitlistedRegistration rather than collecting or counting all waitlisted registrations first."
},
{
"name": "Keeps implementation focused",
"category": "maintainability",
"max_score": 5,
"description": "Keeps the implementation limited to the requested helpers without unrelated caching, concurrency, dependencies, API redesign, or broad style rewrites."
}
],
"metadata": {
"invocation": "natural",
"task_type": "implementation",
"evidence_type": "focused_reference",
"reference_selection": "Focused reference coverage for nested session-registration collector behavior and helper extraction.",
"runtime_reference_overlap_rationale": "Allowed only because this is reference-suite focused coverage, not ordinary broad lift."
}
]
}
16 changes: 16 additions & 0 deletions evals-reference/26-uppercase-side-effect-review/criteria-meta.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
{
"metadata": {
"invocation": "explicit",
"task_type": "review",
"evidence_type": "ordinary_lift",
"issue": "https://github.com/martinfrancois/java-streams-skill/issues/4"
},
"categories": {
"Creates review artifact": "safety",
"Identifies external mutation": "stream_quality",
"Warns against unsafe parallelStream shortcut": "stream_quality",
"Suggests direct result-producing replacement": "stream_quality",
"Addresses performance honestly": "stream_quality",
"Keeps review focused": "maintainability"
}
}
14 changes: 1 addition & 13 deletions evals-reference/26-uppercase-side-effect-review/criteria.json
Original file line number Diff line number Diff line change
Expand Up @@ -4,45 +4,33 @@
"checklist": [
{
"name": "Creates review artifact",
"category": "safety",
"max_score": 5,
"description": "Creates review.md with a clear recommendation."
},
{
"name": "Identifies external mutation",
"category": "stream_quality",
"max_score": 25,
"description": "Explains that forEach mutates the external ArrayList and that the stream should produce the result list directly instead."
},
{
"name": "Warns against unsafe parallelStream shortcut",
"category": "stream_quality",
"max_score": 25,
"description": "Does not recommend simply changing stream() to parallelStream() while keeping the external ArrayList mutation. Full credit may still be awarded when the review first makes the code side-effect-free, then prominently recommends benchmarking a pure parallelStream map plus collector/toList version for the 10-million-name case before relying on it for throughput."
},
{
"name": "Suggests direct result-producing replacement",
"category": "stream_quality",
"max_score": 25,
"description": "Shows or clearly recommends a replacement such as names.stream().map(String::toUpperCase).toList(), or collect(Collectors.toList()) when Java-version compatibility or mutability requires it."
},
{
"name": "Addresses performance honestly",
"category": "stream_quality",
"max_score": 15,
"description": "Distinguishes the correctness fix from the performance decision: does not claim direct collection alone is a guaranteed throughput win, strongly recommends benchmarking a pure parallel stream for the 10-million-name case, and warns that parallelStream can be slower for small lists or mostly-small call paths."
},
{
"name": "Keeps review focused",
"category": "maintainability",
"max_score": 5,
"description": "Keeps the review concise and focused on the stream/performance issue without unrelated rewrites, broad style commentary, or internal workflow labels such as hard stop, marker, scan, checklist, skill, rubric, or criteria."
}
],
"metadata": {
"invocation": "explicit",
"task_type": "review",
"evidence_type": "ordinary_lift",
"issue": "https://github.com/martinfrancois/java-streams-skill/issues/4"
}
]
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
{
"metadata": {
"invocation": "natural",
"task_type": "implementation",
"evidence_type": "ordinary_lift",
"issue": "https://github.com/martinfrancois/java-streams-skill/issues/4"
},
"categories": {
"Creates coherent Java 17 artifact": "safety",
"Preserves requested behavior": "safety",
"Uses direct stream result production": "stream_quality",
"Avoids unsafe parallel shortcut": "stream_quality",
"Keeps performance scope focused": "stream_quality",
"Keeps API focused": "maintainability"
}
}
14 changes: 1 addition & 13 deletions evals-reference/27-uppercase-names-implementation/criteria.json
Original file line number Diff line number Diff line change
Expand Up @@ -4,45 +4,33 @@
"checklist": [
{
"name": "Creates coherent Java 17 artifact",
"category": "safety",
"max_score": 10,
"description": "Creates UppercaseNames.java with the requested uppercaseNames(List<String> names) method, necessary imports, and Java 17-compatible code."
},
{
"name": "Preserves requested behavior",
"category": "safety",
"max_score": 15,
"description": "Returns every input name converted with String::toUpperCase in the original encounter order and does not mutate the input list."
},
{
"name": "Uses direct stream result production",
"category": "stream_quality",
"max_score": 35,
"description": "Uses a direct stream chain such as names.stream().map(String::toUpperCase).toList(), or an equivalent direct collector, instead of creating a list and mutating it from forEach."
},
{
"name": "Avoids unsafe parallel shortcut",
"category": "stream_quality",
"max_score": 25,
"description": "Does not use parallelStream(), .parallel(), or manual concurrent workers together with shared mutable accumulation. Full credit may be awarded for either sequential direct collection or a pure side-effect-free parallelStream map plus ordered collector/toList implementation that preserves behavior for the large-list performance task."
},
{
"name": "Keeps performance scope focused",
"category": "stream_quality",
"max_score": 10,
"description": "Keeps the implementation O(n), avoids extra unnecessary passes, and does not add caching, batching frameworks, or unrelated memory-heavy transformations."
},
{
"name": "Keeps API focused",
"category": "maintainability",
"max_score": 5,
"description": "Keeps the implementation limited to the requested helper without unrelated methods, dependencies, or broad API redesign."
}
],
"metadata": {
"invocation": "natural",
"task_type": "implementation",
"evidence_type": "ordinary_lift",
"issue": "https://github.com/martinfrancois/java-streams-skill/issues/4"
}
]
}
Loading
Loading