Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
38 changes: 28 additions & 10 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,6 +58,8 @@ slices rather than nil slices.
| `testdata/package-lock.json` | `test`, `fixture`, `generated` |
| `packages/api/uv.lock` | `generated` |
| `src/Form.Designer.cs` | `source`, `generated` |
| `src/app.min.js` | `source`, `generated`, `minified` |
| `service.pb.go` | `generated` |
| `.pytest_cache/v/cache/nodeids` | `generated`, `cache` |
| `CMakeFiles/app.dir/main.o` | `generated`, `build-output` |
| `Pods/library/main.swift` | `vendor` |
Expand All @@ -76,7 +78,7 @@ slices rather than nil slices.

The vocabulary is `source`, `test`, `fixture`, `example`, `benchmark`, `fuzz`,
`vendor`, `generated`, `build-output`, `cache`, `documentation`, `legal`,
`build`, `ci`, `packaging`, `tooling` and `configuration`. `build` identifies
`build`, `ci`, `packaging`, `tooling`, `configuration` and `minified`. `build` identifies
build instructions and wrappers. `build-output` identifies specific emitted
build-system files and directories. These labels describe path conventions;
they do not establish that a tracked file is safe to delete. Coverage is based
Expand Down Expand Up @@ -210,21 +212,35 @@ The caller supplies a rooted filesystem; `os.Root.FS` provides containment
for a directory on disk. No files are read for content classification during
`Walk`, and no directory names are automatically excluded.

`ClassifyReader(path, reader)` adds optional generated Go header detection and
`ClassifyReader(path, reader)` adds optional generated and minified detection and
reads at most 8 KiB plus one byte used to detect truncation. `ClassifyBlob`
provides the same result for callers that already hold the complete contents.
Both have classifier methods that retain vendor-root context. Inspection is
limited to the first 8 KiB and 40 lines, ending at the first non-comment token.
Go tokenization prevents markers in strings or block comments from becoming
generated evidence.
limited to the first 8 KiB and 40 lines. Go retains its strict generated marker
check, ending at the first non-comment token. Go tokenization prevents markers
in strings or block comments from becoming generated evidence.

Other supported source files match `Code generated by`, `DO NOT EDIT` or
`@generated` in leading comments, ignoring case. Hash comments are checked for
`.py`, `.rb` and `.sh`; line and block comments are checked for `.js`, `.mjs`,
`.cjs`, `.jsx`, `.ts`, `.tsx`, `.c`, `.h`, `.cc`, `.cpp`, `.hpp`, `.cs`, `.java`,
`.rs`, `.swift`, `.kt` and `.dart`. CSS uses block comments. Inspection stops at
code, so a marker in a string or a later comment is not a generated header.

JavaScript, TypeScript and CSS also receive `generated` for source-map comment
directives within the inspected prefix. A line with at least 1,000 code bytes
and fewer than one space or tab per ten code bytes receives both `minified`
and `generated`. Strings and comments do not count toward that threshold.
These are heuristics: short minified files and source-map directives beyond
the inspection limits can be missed. No tail reads are performed.

The result includes `HeaderChecked`, `BytesExamined` and `HeaderLimited`.
Other file types receive path-only classification. Missing generated
Unsupported file types receive path-only classification. Missing generated
evidence does not prove a file was handwritten, and a generated lockfile
may still be essential dependency evidence.

Content caches should include the blob ID, `ContentVersion` and whether the
occurrence name has a `.go` suffix. The same blob can receive path-only
Content caches should include the blob ID, `ContentVersion` and the occurrence
filename extension, including its casing. The same blob can receive path-only
classification at one occurrence and generated-header evidence at another.

## Corpus
Expand Down Expand Up @@ -262,8 +278,10 @@ any roles matched inside them.
Named minified JavaScript/CSS files, source maps and .NET designer files
also receive `generated`. Source-map and designer suffixes accept mixed
case; minified suffixes use lowercase `.min.js`, `-min.js`, `.min.css` and
`-min.css`. A plain `.d.ts` filename or a name such as `jquery.js` does not
establish generated or vendored content.
`-min.css` and also receive `minified`. Protobuf output names ending in `.pb.go`,
`_pb2.py`, `_pb2_grpc.py`, `_pb2.pyi` or `_pb2_grpc.pyi` receive `generated`.
A plain `.d.ts` filename or a name such as `jquery.js` does not establish
generated or vendored content.

Named community documents include support, governance, maintainers,
authors and roadmaps, using case-insensitive stems and selected document
Expand Down
27 changes: 17 additions & 10 deletions blob.go
Original file line number Diff line number Diff line change
Expand Up @@ -13,11 +13,11 @@ import (
const (
MaxHeaderBytes = 8192
MaxHeaderLines = 40
ContentVersion = "1"
ContentVersion = "2"
headerReadLimit = MaxHeaderBytes + 1
)

// BlobResult describes the bounded Go header inspection, where applicable.
// BlobResult describes bounded content inspection, where applicable.
// Absence of Generated is not proof that a file was handwritten.
type BlobResult struct {
Result
Expand All @@ -26,14 +26,14 @@ type BlobResult struct {
HeaderLimited bool `json:"header_limited"`
}

// ClassifyBlob adds generated Go header evidence to path classification.
// ClassifyBlob adds generated and minified content evidence to path classification.
// Contents must be the complete file. At most MaxHeaderBytes and MaxHeaderLines
// are inspected. Other languages receive path-only classification.
// are inspected. Unsupported file types receive path-only classification.
func ClassifyBlob(name string, contents []byte) (BlobResult, error) {
return defaults.ClassifyBlob(name, contents)
}

// ClassifyReader adds generated Go header evidence using bounded input.
// ClassifyReader adds content evidence using bounded input.
// It reads at most MaxHeaderBytes plus one byte used to detect truncation.
func ClassifyReader(name string, reader io.Reader) (BlobResult, error) {
return defaults.ClassifyReader(name, reader)
Expand All @@ -45,7 +45,7 @@ func (c *Classifier) ClassifyBlob(name string, contents []byte) (BlobResult, err
if err != nil || !inspect {
return blob, err
}
return inspectGoHeader(blob, name, contents), nil
return inspectContent(blob, name, contents), nil
}

// ClassifyReader preserves this classifier's vendor-root context while
Expand All @@ -63,9 +63,9 @@ func (c *Classifier) ClassifyReader(name string, reader io.Reader) (BlobResult,
}
contents, err := io.ReadAll(io.LimitReader(reader, headerReadLimit))
if err != nil {
return BlobResult{}, fmt.Errorf("read Go header: %w", err)
return BlobResult{}, fmt.Errorf("read content header: %w", err)
}
return inspectGoHeader(blob, name, contents), nil
return inspectContent(blob, name, contents), nil
}

func (c *Classifier) blobResult(name string) (BlobResult, bool, error) {
Expand All @@ -74,10 +74,10 @@ func (c *Classifier) blobResult(name string) (BlobResult, bool, error) {
return BlobResult{}, false, err
}
blob := BlobResult{Result: result}
return blob, strings.HasSuffix(name, ".go"), nil
return blob, contentType(name) != "", nil
}

func inspectGoHeader(blob BlobResult, name string, contents []byte) BlobResult {
func inspectContent(blob BlobResult, name string, contents []byte) BlobResult {
blob.HeaderChecked = true
header := contents[:min(len(contents), MaxHeaderBytes)]
lines := 0
Expand All @@ -92,6 +92,13 @@ func inspectGoHeader(blob BlobResult, name string, contents []byte) BlobResult {
}
blob.HeaderLimited = len(header) < len(contents)
blob.BytesExamined = len(header)
if contentType(name) == "go" {
return inspectGoHeader(blob, name, header)
}
return inspectSourceHeader(blob, name, header)
}

func inspectGoHeader(blob BlobResult, name string, header []byte) BlobResult {
file := token.NewFileSet().AddFile(name, -1, len(header))
var lexer scanner.Scanner
lexer.Init(file, header, nil, scanner.ScanComments)
Expand Down
16 changes: 9 additions & 7 deletions blob_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -98,13 +98,15 @@ func (r *countingReader) Read(buffer []byte) (int, error) {
func FuzzBlob(f *testing.F) {
f.Add([]byte("// Code generated by test. DO NOT EDIT.\n"))
f.Fuzz(func(t *testing.T, data []byte) {
got, err := roles.ClassifyBlob("vendor/client.go", data)
if err != nil || !got.Has(roles.Vendor) || got.BytesExamined > min(len(data), roles.MaxHeaderBytes) {
t.Fatal(got, err)
}
fromReader, err := roles.ClassifyReader("vendor/client.go", bytes.NewReader(data))
if err != nil || !reflect.DeepEqual(fromReader, got) {
t.Fatalf("reader differs: %+v, %v", fromReader, err)
for _, name := range []string{"vendor/client.go", "vendor/client.py", "vendor/client.js", "vendor/client.css"} {
got, err := roles.ClassifyBlob(name, data)
if err != nil || !got.Has(roles.Vendor) || got.BytesExamined > min(len(data), roles.MaxHeaderBytes) {
t.Fatal(got, err)
}
fromReader, err := roles.ClassifyReader(name, bytes.NewReader(data))
if err != nil || !reflect.DeepEqual(fromReader, got) {
t.Fatalf("reader differs: %+v, %v", fromReader, err)
}
}
})
}
26 changes: 26 additions & 0 deletions cmd/roles/main_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,32 @@ func TestRun(t *testing.T) {
}
}

func TestRunGeneratedAndMinifiedPaths(t *testing.T) {
for _, name := range []string{"src/app.min.js", "src/app-min.css", "src/service.pb.go", "src/service_pb2.py"} {
for _, labelsOnly := range []bool{false, true} {
args := []string{name}
if labelsOnly {
args = append([]string{labelsOnlyFlag}, args...)
}
var out bytes.Buffer
if err := run(args, &out); err != nil {
t.Fatal(err)
}
var got roles.Result
if err := json.Unmarshal(out.Bytes(), &got); err != nil {
t.Fatal(err)
}
want := []roles.Role{roles.Source, roles.Generated}
if strings.Contains(name, "app") {
want = append(want, roles.Minified)
}
if !slices.Equal(got.Roles, want) {
t.Fatalf("run(%v) = %v, want %v", args, got.Roles, want)
}
}
}
}

func TestRunJSONSchema(t *testing.T) {
var out bytes.Buffer
if err := run([]string{"LICENSE"}, &out); err != nil {
Expand Down
182 changes: 182 additions & 0 deletions content.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,182 @@
package roles

import (
"bytes"
"path"
"regexp"
"strings"
"unicode/utf8"
)

const (
minifiedLineBytes = 1000
cssContent = "css"
)

var generatedMarker = regexp.MustCompile(`(?i)\bcode generated by\b|\bdo not edit\b|(?:^|\s)@generated(?:\s|$)`)

func contentType(name string) string {
switch path.Ext(name) {
case ".go":
return "go"
case ".py", ".rb", ".sh":
return "hash"
case ".js", ".mjs", ".cjs", ".jsx", ".ts", ".tsx":
return "javascript"
case ".css":
return cssContent
case ".c", ".h", ".cc", ".cpp", ".hpp", ".cs", ".java", ".rs", ".swift", ".kt", ".dart":
return "c"
default:
return ""
}
}

func inspectSourceHeader(blob BlobResult, name string, header []byte) BlobResult {
if blob.HeaderLimited {
header = trimPartialRune(header)
}
if bytes.IndexByte(header, 0) >= 0 || !utf8.Valid(header) {
return blob
}
kind := contentType(name)
if generatedHeader(header, kind, blob.HeaderLimited) {
blob.addContentEvidence(name, "generated.header", Generated, "header")
}
if kind == "javascript" || kind == cssContent {
minified, sourceMap := inspectWebContent(header, kind, blob.HeaderLimited)
if sourceMap {
blob.addContentEvidence(name, "generated.source-map", Generated, "source-map")
}
if minified {
blob.addContentEvidence(name, "generated.minified-content", Generated, "minified")
blob.addContentEvidence(name, "minified.content", Minified, "long-line")
}
}
return blob
}

func (b *BlobResult) addContentEvidence(name, rule string, role Role, subtype string) {
b.Roles = (setOf(b.Roles) | roleBit(role)).List()
b.Evidence = append(b.Evidence, Evidence{Rule: rule, Role: role, Path: name, Subtype: subtype, Source: "roles"})
}

func generatedHeader(header []byte, kind string, limited bool) bool {
text := strings.TrimPrefix(string(header), "\ufeff")
if strings.HasPrefix(text, "#!") {
_, text, _ = strings.Cut(text, "\n")
}
for {
text = strings.TrimLeft(text, " \t\r\n\v\f")
comment, rest, ok := leadingComment(text, kind, limited)
if !ok {
return false
}
if generatedMarker.MatchString(comment) {
return true
}
text = rest
}
}

func leadingComment(text, kind string, limited bool) (comment, rest string, ok bool) {
switch {
case kind == "hash":
if !strings.HasPrefix(text, "#") {
return "", "", false
}
text = text[1:]
case strings.HasPrefix(text, "/*"):
comment, rest, ok = strings.Cut(text[2:], "*/")
return comment, rest, ok
case kind != cssContent && strings.HasPrefix(text, "//"):
text = text[2:]
default:
return "", "", false
}
comment, rest, ok = strings.Cut(text, "\n")
return comment, rest, ok || !limited
}

func inspectWebContent(header []byte, kind string, limited bool) (minified, sourceMap bool) {
text := string(header)
code, spaces := 0, 0
finishLine := func() {
minified = minified || minifiedLine(code, spaces)
code, spaces = 0, 0
}
for len(text) > 0 {
comment, rest, ok := leadingComment(text, kind, limited)
if ok {
sourceMap = sourceMap || sourceMapComment(comment)
if strings.Contains(text[:len(text)-len(rest)], "\n") {
finishLine()
}
text = rest
continue
}
if strings.HasPrefix(text, "/*") || (kind != cssContent && strings.HasPrefix(text, "//")) {
break
}
if text[0] == '\'' || text[0] == '"' || text[0] == '`' {
end := quotedEnd(text)
if strings.Contains(text[:end], "\n") {
finishLine()
}
text = text[end:]
continue
}
switch text[0] {
case '\n', '\r':
finishLine()
case ' ', '\t':
spaces++
default:
code++
}
text = text[1:]
}
finishLine()
return minified, sourceMap
}

func minifiedLine(code, spaces int) bool {
const whitespaceRatio = 10
return code >= minifiedLineBytes && spaces*whitespaceRatio < code
}

func quotedEnd(text string) int {
quote := text[0]
for i := 1; i < len(text); i++ {
switch text[i] {
case '\\':
i++
case quote:
return i + 1
}
}
return len(text)
}

// trimPartialRune drops a multibyte rune split by the header byte limit so
// truncation does not disqualify an otherwise valid UTF-8 header.
func trimPartialRune(header []byte) []byte {
for i := 1; i < utf8.UTFMax && i <= len(header); i++ {
if !utf8.RuneStart(header[len(header)-i]) {
continue
}
if !utf8.Valid(header[len(header)-i:]) {
return header[:len(header)-i]
}
break
}
return header
}

func sourceMapComment(comment string) bool {
if len(comment) == 0 || (comment[0] != '#' && comment[0] != '@') {
return false
}
value, ok := strings.CutPrefix(strings.TrimSpace(comment[1:]), "sourceMappingURL=")
return ok && len(strings.Fields(value)) == 1
}
Loading
Loading