Note
pnpm setup updated by a LLM-based AI tool (Codex/GPT-6).
Conciseness edits by a LLM-based AI tool (Codex/GPT-6).
Note
Font requirements added by a LLM-based AI tool (Claude Code/Opus 5.5).
Warning
This is work in progress. All data and models are preliminary.
This repository contains an exploratory study of factors associated with language and reading outcomes in children with Down syndrome.
Children with Down syndrome experience delays in language and reading development. Understanding the factors associated with better outcomes may help families and practitioners choose teaching strategies and interventions.
We analyse language, reading and related measures from the Reading and Language Intervention (RLI) trial. Earlier publications report the randomised trial, speech production accuracy and teaching of blending skills.
We first use gradient boosting, which combines decision trees, to identify variables that help predict gains and achievement levels. Bayesian models then estimate treatment contrasts and adjusted associations, with probability distributions that describe their uncertainty. METHODS.md explains the methods and limits on causal interpretation.
Start with the documentation guide. It links the methods, model catalogue, worked example and refit instructions. The notes guide separates dated findings from decisions that still govern the analysis. The integrated report is a draft.
We share source code and anonymised data under open licences. We welcome partners to develop models, interpret findings, contribute data and explore further datasets or methods.
git clone https://github.com/dseinternational/language-reading-predictors.git
cd language-reading-predictorsuv sync fetches the pinned dse-research-utils dependency from GitHub. A sibling research checkout is needed only for local development of that library; see pyproject.toml for the path-source option.
For a worked introduction to the Bayesian code, follow one statistical model from synthetic data to a reported difference.
Install uv, which also provides Python. From the repository root, create the environment:
uv sync --lockedRun commands with uv run, for example uv run pytest or uv run python scripts/fit_model.py lrp-rli-gbg-001. Activating .venv also works.
Supported platforms are Linux (x86-64 and arm64), Apple Silicon macOS, and Windows (x86-64). Intel macOS is not supported, because numba no longer publishes macOS x86-64 wheels.
Plotting model graphs additionally requires the system Graphviz dot binary, which is not a Python package: brew install graphviz, apt install graphviz or winget install Graphviz.Graphviz.
Figures, model graphs and reports set text in Noto Sans and equations in Noto Sans Math. These are system fonts: brew install --cask font-noto-sans font-noto-sans-math, sudo apt install fonts-noto-core, or install both from Google Fonts on Windows. matplotlib caches its font list, so after installing them delete fontlist-*.json from the directory that uv run python -c "import matplotlib; print(matplotlib.get_cachedir())" prints. Without the fonts, text falls back to another installed font, such as Arial or DejaVu Sans.
Install Quarto to create reports and the Node.js version declared in package.json to run spelling and formatting checks.
Install pnpm. The packageManager field in package.json pins its version. From the repository root, install the Node dependencies from pnpm-lock.yaml:
pnpm install --frozen-lockfile
pnpm run spellcheck
pnpm run format:check- Code uses the GNU Affero General Public License, version 3 or later (
AGPL-3.0-or-later). See the package metadata and source-file headers. - Documentation, reports and papers use Creative Commons Attribution 4.0 International (
CC BY 4.0). See docs/LICENSE. - Data use Creative Commons Attribution 4.0 International (
CC BY 4.0). See data/LICENSE.