Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 20 additions & 1 deletion .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -50,6 +50,9 @@ cover/

# Environments
.env
.env.*
!.env.example
!code/.env.example
.envrc
.venv
env/
Expand Down Expand Up @@ -79,4 +82,20 @@ cython_debug/
code/**/runs/
code/**/wandb/
code/**/__pycache__/
code/**/.pytest_cache/
code/**/.pytest_cache/

# SSH keys / credentials (never commit private keys or secrets)
*.pem
*.key
id_rsa*
id_ed25519*
id_ecdsa*
*_rsa
*_ed25519
*_ecdsa
*_key
*_key.txt
known_hosts
*.ppk
secrets.*
credentials.*
14 changes: 12 additions & 2 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,16 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/) and

<!-- Changes landing on `main` between releases go here. -->

### Added

- **Chapters 6 through 9** (2026-09-23) — the remaining hands-on code: full supervised fine-tuning (`chapter06`), black-box distillation (`chapter07`), DPO with GRPO/RFT examples (`chapter08`), and the operations toolkit (`chapter09`: JSON model registry, TF-IDF drift detector, rollback demo, safety monitor). All run on the shared IT-support dataset in `data/it_support/` with the held-out test split (`--split test`).

### Changed

- **Chapter 6** (2026-09-26) — `eval/safety/` now holds the run the chapter prints (base 62%, fine-tuned 62%, harmful-request refusal 100% to 50%, regression flagged); the README's safety summary matches it.
- **Chapter 8** (2026-09-26) — README training-time row reflects the measured runs (about 3.5 min full DPO on three 24 GB cards, about 2.8 min LoRA-DPO on one).
- **Chapter 9** (2026-09-26) — README safety-monitor compare command writes `safety_report_dpo.json` against `safety_report.json`; `model_registry.py` docstring shows `--registry_dir` before the subcommand; the retired `data/drift_report_same_domain.json` removed.

## [MEAP-v0.1] — TBD

First public release of the code repository alongside Manning's MEAP launch.
Expand All @@ -16,7 +26,7 @@ First public release of the code repository alongside Manning's MEAP launch.

- Initial release of code for Chapters 1 through 5.
- **Chapter 1** — reproducibility script for the §1.6 sidebar (`run_sidebar_example.py`). Runs the chapter's prompt through base Qwen3-4B, the Chapter 5 LoRA adapter, and the Chapter 6 SFT model side by side; degrades gracefully when later-chapter artifacts are not yet built.
- **Chapter 2** — Unsloth-based fine-tuning quickstart reproducing the Dragon LLM open-finance recipe on Qwen3-0.6B end to end (data preparation across four HF datasets, LoRA via TRL's `SFTTrainer`, five evaluation tests, model export).
- **Chapter 2** — five-step LoRA quickstart (`quickstart.py`) on Qwen3-4B-Instruct-2507 with a 40-example slice of the IT-support dataset: prepare data, load the base model with a LoRA config, train 20 steps with TRL's `SFTTrainer`, compare outputs before and after, save the adapter with a manifest. Plus `run_chapter5_adapter.py`, which previews the chapter 5 adapter (local or from the Hub) on the same prompts.
- **Chapter 3** — data-quality experiment, six-step synthetic-data-generation pipeline using a frontier teacher, and a standalone `DatasetManifest` module for content hashing and lineage tracking.
- **Chapter 4** — few-shot ticket classifier, many-shot prompt assembly, prompt validator with run-to-run variability measurement, minimal RAG pipeline (50 lines), Precision@k / Recall@k / Hit@1 retrieval evaluator.
- **Chapter 5** — LoRA and QLoRA training, evaluation, and inference on a 400-example Dolly subset of Qwen3-4B-Instruct-2507; published adapter on Hugging Face Hub at `bahree/qwen3-4b-dolly-lora-ch5`.
Expand All @@ -27,7 +37,7 @@ First public release of the code repository alongside Manning's MEAP launch.

### Notes

- Chapters 6 through 9 (Full SFT, Distillation, DPO, Operations) are written and will be released to this repo as they reach MEAP in subsequent drops.
- Chapters 6 through 9 (Full SFT, Distillation, DPO, Operations) landed on `main` on 2026-09-23 (see Unreleased above).
- All hands-on chapters use Qwen3-4B-Instruct-2507 as the base model and Databricks Dolly-15K (filtered subsets) as the dataset, so the chapters compose into a single coherent example pipeline.

---
Expand Down
34 changes: 34 additions & 0 deletions code/chapter05/eval/test_split/eval_report_test.log
Original file line number Diff line number Diff line change
@@ -0,0 +1,34 @@

Step 1/4: Loading base model...
Failed to load /home/amit/FTBook-pvt/code/.venv/lib/python3.12/site-packages/torchao/_C_mxfp8.cpython-310-x86_64-linux-gnu.so: Could not load this library: /home/amit/FTBook-pvt/code/.venv/lib/python3.12/site-packages/torchao/_C_mxfp8.cpython-310-x86_64-linux-gnu.so
Failed to load /home/amit/FTBook-pvt/code/.venv/lib/python3.12/site-packages/torchao/_C_cutlass_90a.abi3.so: Could not load this library: /home/amit/FTBook-pvt/code/.venv/lib/python3.12/site-packages/torchao/_C_cutlass_90a.abi3.so
/home/amit/FTBook-pvt/code/.venv/lib/python3.12/site-packages/torch/cuda/__init__.py:1061: UserWarning: Can't initialize NVML
raw_cnt = _raw_device_count_nvml()
Loading checkpoint shards: 0%| | 0/3 [00:00<?, ?it/s]Loading checkpoint shards: 33%|███▎ | 1/3 [00:01<00:02, 1.17s/it]Loading checkpoint shards: 67%|██████▋ | 2/3 [00:02<00:01, 1.14s/it]Loading checkpoint shards: 100%|██████████| 3/3 [00:02<00:00, 1.29it/s]
✓ Base model loaded

Step 2/4: Evaluating base model...
The following generation flags are not valid and may be ignored: ['temperature', 'top_p', 'top_k']. Set `TRANSFORMERS_VERBOSITY=info` for more details.
Evaluating examples... ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 100%
Evaluating toy test set... ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 100%
Running safety checks... ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 100%
✓ Base evaluation complete

Step 3/4: Loading adapter from chapter05/runs/it_lora...
Loading checkpoint shards: 0%| | 0/3 [00:00<?, ?it/s]Loading checkpoint shards: 33%|███▎ | 1/3 [00:00<00:01, 1.64it/s]Loading checkpoint shards: 67%|██████▋ | 2/3 [00:01<00:00, 1.60it/s]Loading checkpoint shards: 100%|██████████| 3/3 [00:01<00:00, 2.37it/s]
✓ Adapter loaded

Step 4/4: Evaluating fine-tuned model...
Evaluating examples... ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 100%
Evaluating toy test set... ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 100%
Running safety checks... ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 100%
✓ Fine-tuned evaluation complete


Writing evaluation reports...

✓ Evaluation complete!
✓ JSON report: chapter05/runs/eval_report/report.json
✓ Markdown summary: chapter05/runs/eval_report/report.md

→ View the markdown report for a human-readable summary
Expand Down
Loading
Loading