feat(facade): check repo size and enforce clone limit before cloning (#458) - #459
Dipro-cyber wants to merge 9 commits into
Conversation
9ab30b8 to
b0cb8fe
Compare
|
im a little confused on the theory of how your safety margin system is meant to work. The original idea for that is because the size as github reports it is very likely not accurate to the actual size of the whole directory that you get after running I have other branches that are very work-in-progress for refactoring a lot of CollectOSS to use Happy to chat on slack about this too, im very curious what you find |
|
@MoralCode thank u for the context sir, sorry i was a bit inactive. the github api size field is from a bare repo so it can differ significantly from the actual clone size. happy to revisit the safety margin approach once the pygit2 refactor lands, since that library likely has a way to get the actual disk usage after clone. for now, should i simplify this to just use the raw github-reported size without a safety margin, and let users set the limit knowing it's an approximation? that would make the behavior more predictable while still solving the core disk space problem. |
|
I think the main idea is to figure out the disk space before the clone so we can avoid cloning if its too big. Maybe thats where the conversation in #49 could help (I.e. if we migrated CollectOSS to use bare clones) |
…haoss#458) Signed-off-by: Diptesh Roy <droy88333@gmail.com>
…ent E2E worker timeout Signed-off-by: Diptesh Roy <droy88333@gmail.com>
Per MoralCode's review feedback, the safety margin concept was confusing because the GitHub/GitLab API size field is already an approximation (measured from a bare repo, not an actual clone). Adding a multiplier on top of an already-inaccurate value made the limit unpredictable. Simplified to compare the raw reported size directly against max_clone_size_kb. Users set the limit knowing it's an approximation of the bare repo size, which is the most honest and predictable behavior. Also removes clone_size_safety_margin from config and FacadeHelper. Signed-off-by: Diptesh Roy <droy88333@gmail.com>
4af2829 to
4cc4d18
Compare
|
@MoralCode sir can u check this once, the new feature by @drkrillo has been a help here! |
|
The underlying issue has had additional notes added since this PR was filed. The additional notes Include a way to estimate the size of a checkout (rather than using a fudge factor). This, combined with the bare size of a git repo should give a decent estimate of repo size Could you update this PR to account for that more precise method of checking the size? |
Oh yes sure sir, I will check out the notes. |
Per MoralCode's research, the GitHub API 'size' field alone underestimates by ~4% because it only measures the bare repo. Adding the working tree file size (sum of all blob sizes from /git/trees/HEAD?recursive=1) gives a much more accurate estimate of actual on-disk clone size. If the tree is truncated (very large repo), falls back to bare size only. Updated tests to cover the combined estimation logic. Signed-off-by: Diptesh Roy <droy88333@gmail.com>
|
@MoralCode the new commit now uses the combined bare repo size + file tree blob sizes for github repos. falls back to bare size only if the tree is truncated. also updated tests to cover the combined logic. |
|
Wdym if the tree is truncated? |
sir the github git tree API (/git/trees/HEAD?recursive=1) returns truncated: true for very large repos where the tree has too many entries to return in a single response. in that case we can't sum all blob sizes, so we fall back to just the bare repo size from the metadata api. for most repos this won't happen, github truncates at ~100,000 tree entries. |
- Default max_clone_size_kb changed from 0 (disabled) to 5242880 KB (5 GB) as suggested by MoralCode — provides a sensible out-of-the-box limit - Reverted contributor_interface.py and tasks.py to upstream — those changes are unrelated to the repo size limit feature Signed-off-by: Diptesh Roy <droy88333@gmail.com>
Signed-off-by: Diptesh Roy <droy88333@gmail.com>
|
@MoralCode sir can you give this a check once? |
MoralCode
left a comment
There was a problem hiding this comment.
Took another look and found some additional design related things.
Its going to take me a while to get around to testing this as well, but thanks for the reminders to keep this moving!
Address MoralCode's design review: - Split check_repo_size_limit into three functions: _get_github_repo_size_kb: raises on API error or non-dict response _get_gitlab_repo_size_kb: calls raise_for_status, raises on missing field check_repo_size_limit: policy-only, calls forge-specific fetchers - Fail CLOSED: API error, malformed response, or unsupported forge with limit>0 returns (False, None) to block the clone rather than allow it - Truncated GitHub tree: partial blob data is counted even when truncated=True, since the endpoint is not paginated and partial is better than zero (per MoralCode's feedback) - Fix TypeError crash in git_repo_initialize: estimated_kb can be None when fail-closed, so the error message now handles both cases - Tests: 12 tests covering disabled limit, under/over/exact, truncated tree (blocked and allowed), API error, malformed response, unsupported forge, and GitLab path Signed-off-by: Diptesh Roy <droy88333@gmail.com>
Pre-review audit fixes: - Correct _get_github_repo_size_kb docstring: GitHub's non-recursive tree API does allow full traversal via sub-tree requests; the previous claim that 'partial is all we can get' was inaccurate. Document the actual trade-off: we intentionally use partial data as a lower-bound estimate to avoid N round-trips on large repos. - Update inline comment and truncation warning message to reflect the lower-bound framing consistently. - Add test_gitlab_repo_under_limit: GitLab repo below the configured limit should allow the clone (was previously untested). - Add test_negative_limit_treated_as_disabled: negative limit values are treated the same as 0 (disabled); make this explicit in tests. - Remove unused GitCloneError import from test file. Signed-off-by: Diptesh Roy <droy88333@gmail.com>
dfeb967 to
22037ae
Compare
|
@MoralCode hopefully sir no more reviews are needed after this. |
| "run_analysis": 1, | ||
| "run_facade_contributors": 1, | ||
| "commit_messages": 1, | ||
| "max_clone_size_kb": 5242880, |
There was a problem hiding this comment.
We should probably include the presence of this setting in the docs. people may be surprised if their repos stop cloning if we dont document it (and ideally link to the docs page in the error message)
This is a relatively major new feature so I suspect there may end up being more. I still have yet to actually run this code/test it with a live instance, Ive just been giving quick code reviews in response to your requests. including github and gitlab in this also makes it trickier to test since we dont have full gitlab support yet |
Per MoralCode's request to document this setting so operators aren't surprised when repos stop cloning. - Add max_clone_size_kb section to configuration-file-reference.rst describing default (5 GB), how to disable (set to 0), and the truncated-tree lower-bound caveat - Include the docs URL in the GitCloneError message so the error log points operators directly to the configuration reference Signed-off-by: Diptesh Roy <droy88333@gmail.com>
|
@MoralCode sir did the necessary changes. |
| docs_ref = ( | ||
| "See max_clone_size_kb in the configuration reference: " | ||
| "https://github.com/chaoss/CollectOSS/blob/main/docs/source" | ||
| "/development-guide/configuration-file-reference.rst" |
There was a problem hiding this comment.
can you link to the compiled docs on docs.collectoss.org?
Description
Some git repositories are exceptionally large, which can rapidly exhaust local disk space when cloned. This PR introduces a configurable maximum repository clone size limit (
max_clone_size_kb) and safety margin (clone_size_safety_margin) in CollectOSS.Before running
git clone, CollectOSS queries platform APIs (GitHub REST API or GitLab API) to obtain reported repository size stats and computesestimated_size_kb = reported_size_kb * (1 + safety_margin). If the estimated clone size exceedsmax_clone_size_kb, cloning is skipped and the repository collection status is set toFailed Clone.This PR fixes #458
Notes for Reviewers
max_clone_size_kb(default0, disabled) andclone_size_safety_margin(default0.5, 50% margin) toFacadeconfig incollectoss/application/config.pyandFacadeHelper.check_repo_size_limit()and pre-clone enforcement logic incollectoss/tasks/git/util/facade_worker/facade_worker/repofetch.py.tests/test_tasks/test_git/test_repo_size_limit.py.Signed commits
Generative AI disclosure