Day-2 of PrimeOdin’s daily public builds — RAG over a folder of Markdown notes.
Retrieve → cite → answer. No vector-DB soup on day one. Bag-of-words cosine keeps the machine small enough to read.
git clone https://github.com/primeodin/notes-rag.git
cd notes-rag
pip install -e ".[dev]"
pytest
python -m notes_rag --mock "What is a git remote?"
python -m notes_rag --mock --json "What is a git remote?" # machine-readable answer + sourcesExpected stdout (deterministic on the bundled notes/ corpus + --mock — yours should match):
# pytest
....... [100%]
7 passed
# mock ask
[mock] Based on Git remotes, What RAG is: A remote is a shared copy of your repo, usually on GitHub. `git push` sends your commits to the remote. `git pull` brings remote commits into your local branch. Always `git status`… (question: 'What is a git remote?')
Sources:
- Git remotes (git-remotes.md, score=0.537)
- What RAG is (rag-idea.md, score=0.241)
That [mock] answer plus Sources means retrieval works before you spend a token. If titles or scores drift, the note corpus or scorer changed — open an issue before "fixing" ranking by eye.
export OPENAI_API_KEY=sk-...
# optional: export OPENAI_BASE_URL=http://localhost:11434/v1
python -m notes_rag "What is RAG?"Drop your own .md files in notes/ and ask again.
| Piece | Job |
|---|---|
retrieve.py |
Load Markdown, score with bag-of-words cosine |
answer.py |
Mock or OpenAI-compatible completion grounded in hits |
cli.py |
Question in → answer + citations out |
notes/ |
Sample teaching notes (git, Ollama, RAG, pytest) |
docs/how-scoring-works.md |
Hand-worked bag-of-words cosine walkthrough |
docs/why-cite.md |
Why Sources matter — pytest ranking trap + empty shelf |
- Add
notes/my-topic.mdand ask a question only that file can answer - Raise
--kto pull more context - Swap bag-of-words for real embeddings later — keep the same CLI
See CONTRIBUTING.md for fork → install → mock → PR. Scoped tickets live in Issues.
Open (good first issue):
- #8 —
--list-notesflag (print indexed titles, no question) - #11 —
docs/add-your-own-note.md(prove retrieval picks your file) - #12 — empty-hits notice when nothing retrieves
Shipped:
#7docs/how-scoring-works.md— hand-worked bag-of-words cosinedocs/why-cite.md— trust the Sources line (pytest ranking trap)#3CONTRIBUTING.mdfor first-timers#1--jsonCLI output — merged from community PR #6. Thanks!
New to pull requests? Start at first-commit-ai, then come back.
Tiny, tested teaching repos — starter → mid. Ship one, read it, then climb:
| Lane | Repo | Why open it |
|---|---|---|
| Starter chat | first-commit-ai | Mock-first chat CLI + pytest |
| Starter RAG (this) | notes-rag | Retrieve, cite, answer over Markdown notes |
| Starter tokenizer | tiny-bpe-tokenizer | Watch text become token IDs — train, encode, decode |
| Mid tool agent | tiny-tool-agent | ReAct: Thought, Action, Observation, Final Answer |
| Mid prompt lab | prompt-lab | A/B eval: two prompts, fixed cases, score, winner |
| Starter embeddings | embedding-playground | Cosine nearest neighbors you can see |
| Mid vision | vision-caption-loop | Caption → verify → refine on mock shop scenes |
| Attention mid | attention-warrior | Transformer attention you can hold in one hand |
| Shop skills | mister-jay | Interactive DIY drills (vehicle, electrical, plumbing) — live |
| Literacy (Sinhala) | jay-ai-sinhala | Friends 70+ learning GitHub + AI — live |
| Systems DIY | camera-selector | NVR/Frigate camera planning — live |
Weekday cadence, in order: chat CLI → RAG (this) → tokenizer → tool agent → prompt lab → embedding playground → vision caption loop (shipped) → memory → shop-skill explainer.
Profile forge: github.com/primeodin
MIT