From 28eb35108306e382fd749e210ebdb404ab0775fe Mon Sep 17 00:00:00 2001 From: Amit Bahree Date: Sat, 26 Sep 2026 17:07:11 -0700 Subject: [PATCH] MEAP review feedback - Ch6-9 --- .gitignore | 21 +- CHANGELOG.md | 14 +- .../eval/test_split/eval_report_test.log | 34 ++++ .../eval/test_split/it_qlora_rerun_train.log | 192 ++++++++++++++++++ .../eval/test_split/format_adherence_test.log | 35 ++++ 5 files changed, 293 insertions(+), 3 deletions(-) create mode 100644 code/chapter05/eval/test_split/eval_report_test.log create mode 100644 code/chapter05/eval/test_split/it_qlora_rerun_train.log create mode 100644 code/chapter07/eval/test_split/format_adherence_test.log diff --git a/.gitignore b/.gitignore index fc284fe..5dd6d5f 100644 --- a/.gitignore +++ b/.gitignore @@ -50,6 +50,9 @@ cover/ # Environments .env +.env.* +!.env.example +!code/.env.example .envrc .venv env/ @@ -79,4 +82,20 @@ cython_debug/ code/**/runs/ code/**/wandb/ code/**/__pycache__/ -code/**/.pytest_cache/ +code/**/.pytest_cache/ + +# SSH keys / credentials (never commit private keys or secrets) +*.pem +*.key +id_rsa* +id_ed25519* +id_ecdsa* +*_rsa +*_ed25519 +*_ecdsa +*_key +*_key.txt +known_hosts +*.ppk +secrets.* +credentials.* diff --git a/CHANGELOG.md b/CHANGELOG.md index 46ad763..15d5132 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -8,6 +8,16 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/) and +### Added + +- **Chapters 6 through 9** (2026-09-23) — the remaining hands-on code: full supervised fine-tuning (`chapter06`), black-box distillation (`chapter07`), DPO with GRPO/RFT examples (`chapter08`), and the operations toolkit (`chapter09`: JSON model registry, TF-IDF drift detector, rollback demo, safety monitor). All run on the shared IT-support dataset in `data/it_support/` with the held-out test split (`--split test`). + +### Changed + +- **Chapter 6** (2026-09-26) — `eval/safety/` now holds the run the chapter prints (base 62%, fine-tuned 62%, harmful-request refusal 100% to 50%, regression flagged); the README's safety summary matches it. +- **Chapter 8** (2026-09-26) — README training-time row reflects the measured runs (about 3.5 min full DPO on three 24 GB cards, about 2.8 min LoRA-DPO on one). +- **Chapter 9** (2026-09-26) — README safety-monitor compare command writes `safety_report_dpo.json` against `safety_report.json`; `model_registry.py` docstring shows `--registry_dir` before the subcommand; the retired `data/drift_report_same_domain.json` removed. + ## [MEAP-v0.1] — TBD First public release of the code repository alongside Manning's MEAP launch. @@ -16,7 +26,7 @@ First public release of the code repository alongside Manning's MEAP launch. - Initial release of code for Chapters 1 through 5. - **Chapter 1** — reproducibility script for the §1.6 sidebar (`run_sidebar_example.py`). Runs the chapter's prompt through base Qwen3-4B, the Chapter 5 LoRA adapter, and the Chapter 6 SFT model side by side; degrades gracefully when later-chapter artifacts are not yet built. -- **Chapter 2** — Unsloth-based fine-tuning quickstart reproducing the Dragon LLM open-finance recipe on Qwen3-0.6B end to end (data preparation across four HF datasets, LoRA via TRL's `SFTTrainer`, five evaluation tests, model export). +- **Chapter 2** — five-step LoRA quickstart (`quickstart.py`) on Qwen3-4B-Instruct-2507 with a 40-example slice of the IT-support dataset: prepare data, load the base model with a LoRA config, train 20 steps with TRL's `SFTTrainer`, compare outputs before and after, save the adapter with a manifest. Plus `run_chapter5_adapter.py`, which previews the chapter 5 adapter (local or from the Hub) on the same prompts. - **Chapter 3** — data-quality experiment, six-step synthetic-data-generation pipeline using a frontier teacher, and a standalone `DatasetManifest` module for content hashing and lineage tracking. - **Chapter 4** — few-shot ticket classifier, many-shot prompt assembly, prompt validator with run-to-run variability measurement, minimal RAG pipeline (50 lines), Precision@k / Recall@k / Hit@1 retrieval evaluator. - **Chapter 5** — LoRA and QLoRA training, evaluation, and inference on a 400-example Dolly subset of Qwen3-4B-Instruct-2507; published adapter on Hugging Face Hub at `bahree/qwen3-4b-dolly-lora-ch5`. @@ -27,7 +37,7 @@ First public release of the code repository alongside Manning's MEAP launch. ### Notes -- Chapters 6 through 9 (Full SFT, Distillation, DPO, Operations) are written and will be released to this repo as they reach MEAP in subsequent drops. +- Chapters 6 through 9 (Full SFT, Distillation, DPO, Operations) landed on `main` on 2026-09-23 (see Unreleased above). - All hands-on chapters use Qwen3-4B-Instruct-2507 as the base model and Databricks Dolly-15K (filtered subsets) as the dataset, so the chapters compose into a single coherent example pipeline. --- diff --git a/code/chapter05/eval/test_split/eval_report_test.log b/code/chapter05/eval/test_split/eval_report_test.log new file mode 100644 index 0000000..1f037ef --- /dev/null +++ b/code/chapter05/eval/test_split/eval_report_test.log @@ -0,0 +1,34 @@ + +Step 1/4: Loading base model... +Failed to load /home/amit/FTBook-pvt/code/.venv/lib/python3.12/site-packages/torchao/_C_mxfp8.cpython-310-x86_64-linux-gnu.so: Could not load this library: /home/amit/FTBook-pvt/code/.venv/lib/python3.12/site-packages/torchao/_C_mxfp8.cpython-310-x86_64-linux-gnu.so +Failed to load /home/amit/FTBook-pvt/code/.venv/lib/python3.12/site-packages/torchao/_C_cutlass_90a.abi3.so: Could not load this library: /home/amit/FTBook-pvt/code/.venv/lib/python3.12/site-packages/torchao/_C_cutlass_90a.abi3.so +/home/amit/FTBook-pvt/code/.venv/lib/python3.12/site-packages/torch/cuda/__init__.py:1061: UserWarning: Can't initialize NVML + raw_cnt = _raw_device_count_nvml() + Loading checkpoint shards: 0%| | 0/3 [00:00= 0) instead. + torch._check_is_size(blocksize) + Loading checkpoint shards: 33%|███▎ | 1/3 [00:13<00:27, 13.72s/it] Loading checkpoint shards: 67%|██████▋ | 2/3 [00:27<00:13, 13.72s/it] Loading checkpoint shards: 100%|██████████| 3/3 [00:27<00:00, 7.62s/it] Loading checkpoint shards: 100%|██████████| 3/3 [00:27<00:00, 9.27s/it] + Tokenizing train dataset: 0%| | 0/450 [00:00 loading base + PeftModel + Loading checkpoint shards: 0%| | 0/3 [00:00