Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
25 changes: 23 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

LuaPyre is a clean-slate Lua runtime written in Python. It targets **Lua 5.5.1** semantics, a sandbox-first embedding model, and optional gradual type annotations that feed runtime optimization without creating a second language/runtime.

**Python 3.13+** · **current pre-alpha: 0.38.0a1**
**Python 3.13+** · **current pre-alpha: 0.40.0a1**

LuaPyre implements Lua 5.5.1 language semantics for its supported sandboxed embedding profile. The runtime is built around a register VM and explicit Lua frames, with a guarded tiered JIT that specializes proven hot paths and deoptimizes back to the same interpreter.

Expand Down Expand Up @@ -194,6 +194,11 @@ proved dense primitive-write regions, and cheaper fresh record construction.
See the [0.38 performance record](docs/speed-0.38.md) and
[implementation roadmap](docs/performance-roadmap-0.38.md).

0.39 enters cached numeric loops immediately after `FORPREP` and specializes
proved hash-only integer-to-boolean regions without tagged keys or primitive GC
barriers. The typed Sieve workload is 64–67% faster than 0.38. See the
[0.39 performance record](docs/speed-0.39.md).

The [0.36 performance roadmap](docs/performance-roadmap-0.36.md) adds fresh
Python-headroom measurements, 11 focused speed probes, and the next ordered
work on scalar entry, cross-block facts, nested regions, tables, and recursion.
Expand Down Expand Up @@ -363,11 +368,19 @@ It compares:
- PUC Lua 5.5 through `lupa.lua55`
- LuaJIT through `lupa.luajit21` (falling back to `lupa.luajit20` when appropriate)

The corpus contains `micro`, `typed`, and `algorithm` groups. Algorithm workloads include recursive Fibonacci, Sieve, binary trees, table mixing, string construction, and spectral norm. Where LuaPyre uses typed annotations, the native engines receive an equivalent standard-Lua spelling and all backends must produce the same result.
The corpus contains `micro`, `typed`, `algorithm`, and `language` groups. The
language group adds bounded n-body, Mandelbrot, spectral-norm, fannkuch-redux,
binary-trees, FASTA, k-nucleotide, and reverse-complement workloads. Output-heavy
cases use deterministic in-memory checksums. Where LuaPyre uses typed
annotations, the native engines receive an equivalent standard-Lua spelling
and all backends must produce the same result. The exact parameters,
adaptations, and pidigits portability decision are documented in
[`docs/language-benchmarks.md`](docs/language-benchmarks.md).

```bash
python benchmarks/compare_runtimes.py --group typed --require-all
python benchmarks/compare_runtimes.py --group algorithm --require-all
python benchmarks/compare_runtimes.py --group language --require-all
```

Use `--json PATH` for a versioned machine-readable report. The permanent **Four-way runtime benchmark** Actions workflow records text and JSON artifacts. See [`benchmarks/README.md`](benchmarks/README.md) for methodology and comparison guidance.
Expand All @@ -385,6 +398,14 @@ LuaPyre-native `string.dump` output, cross-runtime PUC emission, and additional

## Performance roadmap

The corrected [0.40 performance roadmap](docs/performance-roadmap-0.40.md)
removes three whole-algorithm matrix/tree compilers and retains only reusable
typed-IR, scalar-entry, record, GC/accounting, and sparse-set improvements.
Published benchmark reports now include admission telemetry and classify paths
seen in fewer than three workloads as experimental. See the
[0.40 corrective record](docs/speed-0.40.md) and permanent
[performance generality policy](docs/performance-generality-policy.md).

0.17 establishes the intended optimizer architecture: **a small statically typed IR compiler targeting optimized Python AST today, with the IR remaining reusable by future backends.**

1. exact table-dispatched interpreter as Tier 0 and universal deoptimization target
Expand Down
15 changes: 12 additions & 3 deletions benchmarks/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,7 @@ python benchmarks/compare_runtimes.py --warmups 5 --repeats 15 --require-all --j
python benchmarks/compare_runtimes.py --group micro --require-all
python benchmarks/compare_runtimes.py --group typed --require-all
python benchmarks/compare_runtimes.py --group algorithm --require-all
python benchmarks/compare_runtimes.py --group language --require-all
```

Groups may be supplied more than once. With no `--group`, the complete corpus runs.
Expand All @@ -33,7 +34,8 @@ Groups may be supplied more than once. With no `--group`, the complete corpus ru

- **micro** keeps the small VM/JIT kernels: arithmetic, tables, calls, branches, and coroutines. These are useful for locating dispatch and specialization overhead.
- **typed** contains direct typed-vs-dynamic optimization probes. LuaPyre receives a source-equivalent file headed by `-- luapyre: typed`; Lua 5.5 and LuaJIT receive ordinary Lua because LuaPyre's annotations are intentionally a source extension.
- **algorithm** contains larger end-to-end programs: recursive Fibonacci, Sieve of Eratosthenes, binary trees, a table-mixing kernel, string construction, and spectral norm. Their LuaPyre variants use the fully typed source contract wherever useful, while the reference engines execute equivalent standard Lua.
- **algorithm** contains the project-specific larger programs: recursive Fibonacci, Sieve of Eratosthenes, a table-mixing kernel, and string construction. Their LuaPyre variants use the fully typed source contract wherever useful, while the reference engines execute equivalent standard Lua.
- **language** contains bounded, deterministic ports of n-body, Mandelbrot, spectral-norm, fannkuch-redux, binary-trees, FASTA, k-nucleotide, and reverse-complement. Output-heavy cases use in-memory checksums so every backend runs under the same sandbox. See [`docs/language-benchmarks.md`](../docs/language-benchmarks.md) for parameters, adaptations, and the pidigits disposition.

A workload can therefore carry two spellings of the same algorithm: standard Lua for the reference engines and a fully typed LuaPyre spelling for the optimization target. Both must produce the same expected result. Floating-point workloads may declare a tight absolute tolerance; integer/string workloads remain exact.

Expand Down Expand Up @@ -101,7 +103,7 @@ For performance comparisons, compare runs on the same runner class and Python ve
timings, raw samples, and per-phase tier counters. The committed baseline
uses three processes on each Python version; it does not claim a 0.36 speedup.
See [`performance-roadmap-0.36.md`](../docs/performance-roadmap-0.36.md).
- `python_headroom.py` compares nine algorithms against reduced-contract direct
- `python_headroom.py` compares the headroom corpus against reduced-contract direct
Python implementations. Use `PYTHONPATH=src PYTHONHASHSEED=0`, `--revision`,
`--json`, and optionally `--python-first` or repeated `--case` selections.
Every result is checked. Ratios include Lua semantic and representation
Expand All @@ -111,7 +113,14 @@ For performance comparisons, compare runs on the same runner class and Python ve
earlier comparison retained in the 0.34 roadmap.
- `native_headroom.py` compares that same corpus three ways: LuaPyre, native
Lua 5.5 through `lupa.lua55`, and the reduced-contract Python lower bounds.
It runs isolated processes and rotates implementation order.
It runs isolated processes and rotates implementation order. Both headroom
tools report per-workload fast-path admission counts; the native aggregate
labels paths seen in fewer than three distinct workloads as experimental.
- `speed_040_ab.py` is the corrected 0.40 generality corpus. It contains three
structurally different scalar DAGs, nested reductions, and sparse-set uses.
Run the unchanged file on baseline and candidate revisions, then use
`check_speed_generality.py`; a group passes only when all three cases admit
the intended path and improve over baseline.
- `inspect_codegen.py` captures final generated AST/source, selects hot
generated functions using checked profiled executions, and records generic
and warmed adaptive opcode counts plus pooled frame/register-slot counts.
Expand Down
71 changes: 71 additions & 0 deletions benchmarks/check_speed_generality.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,71 @@
"""Enforce the three-program performance/admission gate on two A/B reports."""
from __future__ import annotations

import argparse
import json
from pathlib import Path

from speed_040_ab import GENERALITY_GROUPS


def evaluate(
baseline: dict, candidate: dict, *, minimum_improvement: float = 0.02
) -> dict:
groups = {}
for path, names in GENERALITY_GROUPS.items():
cases = {}
for name in names:
old = baseline["workloads"][name]["steady"]["luapyre"]["median_ms"]
new = candidate["workloads"][name]["steady"]["luapyre"]["median_ms"]
admissions = candidate["workloads"][name].get(
"fast_path_admissions", {}
).get(path, 0)
cases[name] = {
"baseline_ms": old,
"candidate_ms": new,
"improvement_fraction": old / new - 1.0,
"admissions": admissions,
"passes": admissions > 0 and old / new - 1.0 >= minimum_improvement,
}
passed = sum(case["passes"] for case in cases.values())
groups[path] = {
"cases": cases,
"passing_cases": passed,
"status": "accepted" if passed >= 3 else "experimental",
}
return {
"schema_version": 1,
"minimum_improvement_fraction": minimum_improvement,
"groups": groups,
"accepted": all(group["status"] == "accepted" for group in groups.values()),
}


def main() -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("baseline", type=Path)
parser.add_argument("candidate", type=Path)
parser.add_argument("--json", type=Path)
parser.add_argument(
"--minimum-improvement",
type=float,
default=0.02,
help="minimum improvement required in each case (default: 0.02)",
)
args = parser.parse_args()
if args.minimum_improvement <= 0:
parser.error("--minimum-improvement must be positive")
report = evaluate(
json.loads(args.baseline.read_text()),
json.loads(args.candidate.read_text()),
minimum_improvement=args.minimum_improvement,
)
rendered = json.dumps(report, indent=2) + "\n"
if args.json is not None:
args.json.write_text(rendered)
print(rendered, end="")
return 0 if report["accepted"] else 1


if __name__ == "__main__":
raise SystemExit(main())
34 changes: 34 additions & 0 deletions benchmarks/native_headroom.py
Original file line number Diff line number Diff line change
Expand Up @@ -103,6 +103,9 @@ def run_worker(names, order, warmups, repeats):
for key, value in before.items()
if type(value) is int and after[key] != value
}
row["fast_path_admissions"] = dict(
luapyre_runtime.jit_stats.fast_path_admissions
)
results[name] = row
return {
"python": platform.python_version(),
Expand Down Expand Up @@ -134,7 +137,37 @@ def aggregate(revision, runs, warmups, repeats):
row["native_lua55"]["median_ms"]
/ row["python_lower_bound"]["median_ms"]
)
row["fast_path_admissions"] = {
path: max(
run["workloads"][name].get("fast_path_admissions", {}).get(path, 0)
for run in runs
)
for path in sorted({
path
for run in runs
for path in run["workloads"][name].get("fast_path_admissions", {})
})
}
results[name] = row
coverage = {}
paths = sorted({
path
for run in runs
for workload in run["workloads"].values()
for path in workload.get("fast_path_admissions", {})
})
for path in paths:
workloads = sorted({
name
for run in runs
for name, row in run["workloads"].items()
if row.get("fast_path_admissions", {}).get(path, 0) > 0
})
coverage[path] = {
"workloads": workloads,
"workload_count": len(workloads),
"status": "established" if len(workloads) >= 3 else "experimental",
}
return {
"schema_version": 1,
"revision": revision,
Expand All @@ -155,6 +188,7 @@ def aggregate(revision, runs, warmups, repeats):
"attainable-speed guarantees or mathematical lower bounds."
),
"runs": runs,
"fast_path_coverage": coverage,
"workloads": results,
}

Expand Down
Loading
Loading