Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
160 changes: 160 additions & 0 deletions docs/development/ADRs/next/0028-Plain-Builders-Instead-of-Factories.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,160 @@
---

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think some of the information here should not be in the ADR, but part of the PR description. Everything that is a statement about the code in other places rods very quick and will confuse agents as they take this for granted.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, done in 0e472ed. The ADR now only has the context, the decision, the consequences in abstract terms and the alternatives. The statements about specific code (suppression counts, the run_gtfn_no_transforms rename and its cache/metrics analysis, the verification of the pre-built toolchains, the migration list) now live in the PR description.

tags: [backend, otf, toolchain, workflows, dependencies]
---

# Plain Builders Instead of factory-boy Factories

- **Status**: valid
- **Authors**: Enrique González Paredes (@egparedes)
- **Created**: 2026-08-20
- **Updated**: 2026-09-30

In the context of composing the GTFN and DaCe toolchains and their OTF compile
workflows, facing a production dependency on `factory-boy` — a test-data
library — whose `Trait` / `SubFactory` / `SelfAttribute` / `LazyAttribute`
machinery and stringly-typed `__`-path overrides are invisible to the type
checker and fail silently, we decided to replace the factory classes with
plain builder functions driven by one configuration object per toolchain
family and replaceable step builders, to achieve statically checked
composition, steps that agree on shared settings by construction, and one
fewer runtime dependency.

## Context

Every object the toolchain factories built was already a frozen dataclass.
`factory-boy` added a second, parallel construction language on top of them:

- **Untyped.** The declarations are class attributes of a `Params` block, so
`mypy` cannot check them, and the factories needed `type: ignore`
suppressions to type-check at all.
- **Silently wrong.** Overrides are `__`-delimited strings resolved at
runtime. When a path does not resolve, nothing happens. This was not
hypothetical: an override meant to switch a backend to imperative code
generation (`otf_workflow__translation__use_imperative_backend=True`) never
reached the translation step, because a caching trait had wrapped it, so the
backend was the declarative one under another name for its whole life.
- **A runtime dependency for a test-time concern.** `factory-boy` was shipped
to every user to compose a handful of backends.

Whatever replaces the factories must still meet two concerns ADR 0017 lists
for toolchain configuration: some settings must be configured in **several
components in sync** (the target device reaches the translation step, the
compiler and the allocator), and users need to **switch out or tweak nested
steps** (a translation step without the GTIR transforms, a different build
system). `factory-boy` met both, untyped: `SelfAttribute("..device_type")` for
the first, `SubFactory` plus `__` paths for the second.

## Decision

Factory classes are replaced by **plain builder functions**; `factory-boy`
moves to the test dependencies, where the IR test-data factories keep using it
for what it is designed for. The builders follow three rules.

1. **Shared settings live in one configuration.** Each toolchain family has a
frozen config dataclass extending `backend.ToolchainConfig`. It holds the
settings that describe what the toolchain builds for — the device, the
build type, the cache lifetime, the data layout, translation caching — and
the family-specific settings that several steps must agree on. Its defaults
are read from `gt4py.next.config` when the config is created, which gives
ADR 0017's precedence: an explicit argument wins over the user
configuration, which wins over the builder default. Values derived from
shared settings are derived in exactly one place.

2. **Every step is created by a step builder that receives the config.** A
step builder takes the config as its only positional argument and
**step-local settings only** as keyword arguments, so a shared setting
cannot be set for one step alone. A step is customized by passing a
different builder to the toolchain or compile-workflow builder: a
`functools.partial` of the default builder to change a step-local setting,
or any `Callable[[Config], Step]` to replace the step. Nested steps follow
the same pattern.

3. **The toolchain builder owns the composition.** It calls the step builders
and then applies the wrappers, such as the translation cache, so a
customization always lands on the bare step. A fully custom step builder is
responsible for configuring its step consistently with the config it
receives; the builders do not validate what it returns.

```python
make_gtfn_toolchain(
GTFNConfig(gpu=True),
name_postfix="_no_transforms",
translation=functools.partial(make_gtfn_translation, enable_itir_transforms=False),
)
```

The pre-existing flat-keyword `make_dace_backend` is kept as a deprecated
front end over the config-based builder, for existing callers.

## Consequences

- Composition is ordinary, statically checked Python: a misspelled step-local
setting, a value of the wrong type, or an attempt to set a shared setting
through a step builder is a type error, and a `TypeError` when the toolchain
is built.
- Default and partially customized steps agree on the shared settings by
construction, because they read them from the same config. Steps from fully
custom step builders are not checked.
- A step field that must agree with other steps should have no default, so a
builder that forgets to pass it fails instead of silently using the default.
Fields read by a single step, such as the build type of the build system,
may keep a default; a custom step builder that creates such a component
must pass the config value itself.
- A new shared setting is one config field, read where it is needed, instead
of a keyword argument threaded through every builder layer.
- Step builders run at build time and are not stored, so a `lambda` step
builder does not make the toolchain unpicklable.
- There is one flat config per toolchain family. It is not composable: a
toolchain assembled from sub-toolchains with different shared settings would
need a different structure. Nothing needs that today.
- Configuration is pulled by each step builder from the config rather than
pushed down by the parent, so a step builder's signature does not show which
shared settings it reads, and a setting a step gains later silently keeps its
step default until its builder reads it from the config.
- The construction logic is spread over many small step builders, which makes
the consistency between steps harder to see and requires unit tests per
builder.
- The price is also two concepts instead of one (the config and the step
builders), a `partial` is less discoverable than a keyword argument, the step
builders re-list the step-local fields of their step, and the config must
stay limited to shared settings or it turns into a grab-bag.

## Alternatives Considered

### Typed forwarding of per-step options

Builders could take the step-local settings of each inner step as a
`TypedDict` and forward them (`translation={"enable_itir_transforms": False}`).
That keeps `factory-boy`'s one-call ergonomics, keeps the configuration flow
from parent to child visible, and is statically checked: `mypy` checks the
keys when a `TypedDict` is unpacked into the step constructor, so a stale key
is an error, and leaving the shared settings out of the `TypedDict`s keeps
them in sync. Like the step builders chosen here, a `TypedDict` only exposes
the step settings someone added to it.

It was not chosen because replacing a whole step needs a second mechanism next
to the options, typically injecting a pre-built step, which brings back the
problems of the next alternative; and because every shared setting has to be
forwarded by hand through each builder layer.

### Inject pre-built steps, used verbatim

Builders could take shared settings as keyword arguments and accept a
pre-built step, used verbatim. Changing one setting of an inner step then
means building the whole step and repeating the shared settings in it, which
the builder already knew; a GPU toolchain with a CPU translation step is only
caught if the builder checks for it.

### Stamp shared settings onto injected steps

A `with_changes(step, **changes)` helper would stamp the shared settings onto
whichever step is present, applying only the fields the target declares.
Silently ignoring the fields a target does not declare reproduces the failure
mode that motivated this ADR.

### Edit the built toolchain

`dataclasses.replace` on a built toolchain keeps working, but the caller must
know the nesting of wrappers — the path the silently dropped override above
never reached — and the values the builder derived (cache folders, name,
allocator) are not recomputed.
1 change: 1 addition & 0 deletions docs/development/ADRs/next/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,7 @@ Writing a new ADR is simple:
- [0016 - Multiple Backends and Build Systems](0016-Multiple-Backends-and-Build-Systems.md)
- [0017 - Toolchain Configuration](0017-Toolchain-Configuration.md)
- [0027 - External Workspace Memory for DaCe Transients](0027-External_Workspace_Memory.md)
- [0028 - Plain Builders Instead of factory-boy Factories](0028-Plain-Builders-Instead-of-Factories.md)

### Python Integration

Expand Down
41 changes: 30 additions & 11 deletions docs/user/next/advanced/HackTheToolchain.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,25 +46,44 @@ skip_linting_transforms = SkipLinting(**same_steps)
skip_linting_transforms.step_order(DUMMY_FOP)
```

## Alternative Factory
## Alternative Workflow

A toolchain is built from one configuration, `GTFNConfig` (or `DaCeConfig`),
which holds the settings all steps must agree on: the device, the build type,
the cache lifetime, the data layout. Each step is created by a step builder
that receives that configuration. To change a single setting of one step, pass
a `functools.partial` of its default builder; the other settings still come
from the configuration.

```python
class MyCodeGen: ...
import functools

gtfn = gtx.program_processors.runners.gtfn

class Cpp2BindingsGen: ...
debug_gpu_no_transforms = gtfn.make_gtfn_toolchain(
gtfn.GTFNConfig(gpu=True, cmake_build_type=gtx.config.CMakeBuildType.DEBUG),
name_postfix="_debug_no_transforms",
translation=functools.partial(gtfn.make_gtfn_translation, enable_itir_transforms=False),
)
```

To replace a step, pass any callable that takes the configuration and returns
the step. It is still wrapped in the translation cache. Configuring it
consistently with the configuration it receives (the device, for instance) is
up to the callable.

```python
class MyCodeGen: ...

class PureCpp2WorkflowFactory(gtx.program_processors.runners.gtfn.GTFNCompileWorkflowFactory):
translation: workflow.Workflow[
gtx.otf.stages.CompilableProgramDef, gtx.otf.artifacts.ProgramSource
] = MyCodeGen()
bindings: workflow.Workflow[
gtx.otf.artifacts.ProgramSource, gtx.otf.artifacts.ExtensionSource
] = Cpp2BindingsGen()

class Cpp2BindingsGen: ...

PureCpp2WorkflowFactory(cmake_build_type=gtx.config.CMAKE_BUILD_TYPE.DEBUG)

pure_cpp2_workflow = gtfn.make_gtfn_compile_workflow(
gtfn.GTFNConfig(cmake_build_type=gtx.config.CMakeBuildType.DEBUG, cached_translation=False),
translation=lambda cfg: MyCodeGen(),
bindings=lambda cfg: Cpp2BindingsGen(),
)
```

## Invent new Workflow Types
Expand Down
36 changes: 13 additions & 23 deletions docs/user/next/advanced/WorkflowPatterns.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,6 @@ jupyter:
import dataclasses
import re

import factory

import gt4py.next as gtx

Expand Down Expand Up @@ -199,7 +198,7 @@ Let's say we want to make our calculation workflow compatible with string input.

```python editable=true slideshow={"slide_type": ""}
# A plain conversion step turning a string into an int, chained into the
# workflow below and reused by `StrToIntFactory(cached=True)`.
# workflow below and reused by `make_str_to_int(cached=True)`.
def to_int(inp: str) -> int:
assert isinstance(inp, str), "Can not work with 'int'!" # yes, this is horribly contrived
return int(inp)
Expand All @@ -214,9 +213,9 @@ str_calc("1")

<!-- #region editable=true slideshow={"slide_type": ""} -->

### Step with factory (builder)
### Step with a builder

If a step can be useful with different combinations of parameters and wrappers, it should have a factory. In this case we will add a neutral wrapper around it, so we can put any combination of wrappers into that:
If a step is useful with different combinations of parameters and wrappers, give it a **builder function**: a plain function taking the cross-cutting options and returning the assembled step. Steps are frozen dataclasses, so the builder is ordinary code — no factory framework involved, and the result is fully type-checked.

<!-- #endregion -->

Expand All @@ -229,32 +228,23 @@ class AnyStrToInt(gtx.otf.workflow.ChainableWorkflowMixin[str | int, int]):
return self.inner_step(inp)


class StrToIntFactory(factory.Factory):
class Meta:
model = AnyStrToInt
def make_str_to_int(
*, cached: bool = False, step: gtx.otf.workflow.Workflow[str, int] = to_int
) -> AnyStrToInt:
if cached:
step = gtx.otf.workflow.CachedStep.in_memory(step=step, input_fingerprinter=str)
return AnyStrToInt(inner_step=step)

class Params:
default_step = to_int
cached = factory.Trait(
inner_step=factory.LazyAttribute(
lambda o: gtx.otf.workflow.CachedStep.in_memory(
step=o.default_step, input_fingerprinter=str
)
)
)

inner_step = factory.LazyAttribute(lambda o: o.default_step)


cached = StrToIntFactory(cached=True)
uncached = StrToIntFactory()
cached = make_str_to_int(cached=True)
uncached = make_str_to_int()
uncached.inner_step
```

### Example in the Wild

```python
gtx.ffront.past_passes.linters.LinterFactory??
gtx.ffront.past_passes.linters.linter_factory??
```

<!-- #region editable=true slideshow={"slide_type": ""} tags=["skip-execution"] -->
Expand Down Expand Up @@ -413,5 +403,5 @@ gtx.program_processors.runners.gtfn.run_gtfn_gpu.executor.otf_workflow??
```

```python
gtx.program_processors.runners.gtfn.GTFNBackendFactory??
gtx.program_processors.runners.gtfn.make_gtfn_toolchain??
```
8 changes: 1 addition & 7 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -43,6 +43,7 @@ profiling = [
]
scripts = ["pyyaml>=6.0.1", "typer>=0.16.0", "packaging"]
test = [
'factory-boy>=3.3.3',
'hypothesis>=6.0.0',
'nbmake>=1.4.6',
'nox>=2025.02.09',
Expand Down Expand Up @@ -99,7 +100,6 @@ dependencies = [
'dace==2.0.0a9',
'deepdiff>=8.1.0',
'devtools>=0.6',
'factory-boy>=3.3.3',
"filelock>=3.18.0",
'frozendict>=2.3',
'gridtools-cpp>=2.3.9,==2.*',
Expand Down Expand Up @@ -231,12 +231,6 @@ module = 'gt4py.next.iterator.*'
ignore_errors = true
module = 'gt4py.next.iterator.runtime'

[[tool.mypy.overrides]]
ignore_missing_imports = true
implicit_reexport = true
# factory-boy is broken, see https://github.com/FactoryBoy/factory_boy/pull/1114
module = "factory.*"

# -- pytest --
[tool.pytest]

Expand Down
50 changes: 49 additions & 1 deletion src/gt4py/next/backend.py
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@
from typing import Generic

from gt4py._core import definitions as core_defs
from gt4py.next import custom_layout_allocators as next_allocators
from gt4py.next import config, custom_layout_allocators as next_allocators
from gt4py.next.ffront import (
foast_to_gtir,
foast_to_past,
Expand Down Expand Up @@ -172,3 +172,51 @@ def __gt_allocator__(
self,
) -> next_allocators.FieldBufferAllocatorProtocol[core_defs.DeviceTypeT]:
return self.allocator


@dataclasses.dataclass(frozen=True)
class ToolchainConfig:
"""
Settings that describe what a compiled toolchain builds for.

A toolchain builder creates every step from one config, so steps that must
agree on a setting (the target device, the build type, the cache lifetime,
the data layout) read it from the same place and cannot drift apart.
Settings that only tune how a single step does its job are not part of the
config: they are keyword arguments of that step's builder.

Defaults are read from `gt4py.next.config` when the config is created, not
when this module is imported, so the toolchains built from a default config
follow the current user configuration.
"""

#: Target a GPU (the one CuPy was built for) instead of the CPU.
gpu: bool = False
#: Wrap the translation step in a persistent cache.
cached_translation: bool = True
cmake_build_type: config.CMakeBuildType = dataclasses.field(
default_factory=lambda: config.CMAKE_BUILD_TYPE
)
cache_lifetime: config.BuildCacheLifetime = dataclasses.field(
default_factory=lambda: config.BUILD_CACHE_LIFETIME
)
#: Assume unit stride in the horizontal dimension of unstructured fields.
unstructured_horizontal_has_unit_stride: bool = dataclasses.field(
default_factory=lambda: config.UNSTRUCTURED_HORIZONTAL_HAS_UNIT_STRIDE
)

@property
def device_type(self) -> core_defs.DeviceType:
if self.gpu:
return core_defs.CUPY_DEVICE_TYPE or core_defs.DeviceType.CUDA
return core_defs.DeviceType.CPU

@property
def device_name(self) -> str:
"""Device part of the toolchain name."""
return "gpu" if self.gpu else "cpu"

def make_allocator(self) -> next_allocators.FieldBufferAllocatorProtocol:
if self.gpu:
return next_allocators.StandardGPUFieldBufferAllocator()
return next_allocators.StandardCPUFieldBufferAllocator()
Loading
Loading