Repository navigation
Conversation
…or change permissions Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… and keep plib::acl small and std-like Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… -r and -q Makefiles, libtool and automake's archiver-interface probe write the traditional form `ar cr lib.a x.o`; it now means the same letters with a leading '-'. libselinux's `ar rcs` also needed -s with -r; ar always writes the symbol table, so -s with -r or -q changes nothing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
gcc-defaults' rules run `cp -l debian/substvars.native debian/$p.substvars`. Each non-directory is linked with linkat from the directories the walk and the copy hold open, and the new name must be the file the walk examined. With -R, directories are made and their files linked. A link found in the walk is linked itself unless -H/-L say to follow it (GNU follows every link under -l). An existing destination is replaced only under -f or an affirmed -i, by linking under a temporary name and renaming over it, so a link that cannot be made leaves the destination intact; one that already is the source is left alone. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
po4a checks every PO file with `msgfmt --check-format --check-domain -o /dev/null FILE`, which failed the base-passwd and debianutils man-page builds. --check-format runs the c-format comparison that -c -v makes, on its own; --check-domain reports each domain directive that -o ignores as an error. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
XBD 12.2, Guideline 7 lets an option-argument begin with '-'. clap read
the word after -e as an option whenever it did, so autoconf's
AC_PROG_GREP probe, `grep -e 'GREP$' -e '-(cannot match)-'`, failed
with "unexpected argument '-('" and configure gave up.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… option XBD 12.2, Guideline 7 lets an option-argument begin with '-', but clap reads such a word as an option unless the option allows hyphen values. A survey of every utility's value-taking options (each run as `util -X -zq`) found the word refused in 57 utilities; each such option now has `allow_hyphen_values = true`, the idiom grep, tail and pax use: awk -F -f -v; cflow -i -D -I -U; ctags -f; cxref -o -w -D -I -U; at -f -q -t; nm -t; strings -t -n; strip -R -N; more -p -t -n; ed -p; file -m -M -e; od -A -j -N -t; split -a -l -b; gettext -d; iconv -f -t; localedef -f -i -u; msgfmt -D -o; ngettext -d; xgettext -d -K -p -x; m4 -D -U; make -C -f -j; man -C -M -m -S -s; pax --tape-length --new-volume-script; lp -d -n -o -t; xargs -L -n -s -E -I; delta -r -g; get -r -c -i -x; prs -c -d; rmdel -r; unget -r; tabs -T; tput -T; getconf -v; ipcrm -s -S -m -M -q -Q; ps -g -G -p -t -u -U -o -n; csplit -f -n; cut -b -c -f -d; diff -C -U -L --label2; expand -t; fold -w; head -n -c; join -a -e -o -t -v -1 -2; nl -b -d -f -h -i -l -n -s -v -w; paste -d; patch -B -z -V -F -d -D -i -o -p -r; pr --pages --columns -h -l -N -o -w; sed -e -f; sort -o -t -k; unexpand -t; uniq -f -s; mkdir -m; touch -d -t -r; logger -f -p -t; talkd -s --invite-timeout; uucp -n; uustat -k -r -s -u; compress/uncompress/zcat -b -m; uudecode -o. Options whose value is optional (pr -e -i -n -s, attached only) or that take several values (man -l) are left alone: there the next word is not necessarily the option's. pr also rewrites its argv before clap sees it, and turned the header in `pr -h -3` into `--columns=3`; the word after an option that requires a value, and every word after `--`, now passes through unrewritten. plib::testing::assert_hyphen_option_argument runs a utility and fails if the parser refused a '-'-leading option-argument; each utility above has a probe built on it, and cut, nl, awk and pr a test of the result. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…/b/` work The -e and -f sources were found by a second scan of argv for the words "-e" and "-f" alone. A cluster ending in e (`sed -ne p`), a script attached to -e (`-es/a/b/`), a file attached to -f, all of which clap had already accepted, contributed nothing, and sed then failed on the first file operand as a script; a word "-e" after `--` was taken as an option. The order now comes from clap's ArgMatches indices for both options. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…eadable one
POSIX sed reads its input files as one stream, but every file was
processed afresh: line numbers restarted, `$` matched each file's last
line, and the hold space was cleared, so `sed -n '$=' f1 f2` printed
each file's count instead of the total. read_line now runs on from one
file into the next, and a last line without a <newline> gets one when
another file's line follows, as in GNU sed.
An operand that cannot be opened or read is reported in GNU's words
("can't read FILE: ...", "read error on FILE: ...") and skipped; sed
then exits 2 where it used to exit 0.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Debian rules edit files in place, `sed -i '/^misc:Depends=/s/...//' f` and `sed -ri '/^POT-Creation-Date:/ d' po/*.po`; the user approved both extensions. -r and --regexp-extended are aliases of -E. -i's suffix is only ever attached, as in GNU: `-i`, `-i.bak`, `-ni~` (-n and a suffix of ~). The short form is rewritten to --in-place[=SUFFIX] before clap parses the line, leaving alone the option-argument of -e or -f and everything after `--`. A `*` in the suffix stands for the name. Each file is edited as a stream of its own (its own line numbers and `$`; the hold space carries over), with all output -- `=`, `i`, `a`, `r`, `l`, `p` -- going into the new version through the one emit() path. `q` ends the run once its file is written; later files are untouched. The replacement is race-free: the operand is opened without blocking and judged by fstat of what was opened, so only a regular file is edited and a FIFO is refused instead of waited on; the new version is created O_CREAT|O_EXCL 0600 in the same directory (plib::tmp), given the original's owner (when root) or group and then its mode, set-ID bits only where owner/group really match, and renamed over the name, after the original is renamed to the backup name when there is a suffix. A symbolic link operand is read through and replaced by a regular file, as GNU does without --follow-symlinks. A failure leaves the original untouched and removes the new file. Both extensions are listed in NONPOSIX.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… forever A read error other than invalid UTF-8 (EISDIR for a directory operand, EIO) does not advance the input, so the read loop printed "error reading line N" without end; `grep -i GNU pkg ../*` hung a Debian build. Report it once as "NAME: strerror", as GNU grep does, and go on with the next operand; the exit status is 2. An invalid UTF-8 line is still reported and skipped, since read_line has consumed it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`val -` skipped every error from its line iterator with `continue`, so an error that does not advance the input (EISDIR when standard input is a directory, EIO) looped for ever with no output. Report it once as "val: -: strerror", set the cannot-open bit (0x10) in the exit status and stop reading. A line that is not valid UTF-8 is still skipped. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
debhelper's dh_strip runs `rmdir -p --ignore-fail-on-non-empty`, and four Debian packages failed on the unknown option. A removal that fails only because the directory is not empty (ENOTEMPTY or EEXIST, or as in GNU an EACCES/EPERM/EROFS/EBUSY on a directory that holds an entry) is now silent and does not affect the exit status; with -p the walk up the parents stops there. Other errors are still reported. Exit status and resulting tree match GNU coreutils 9.4 over a 14-case matrix including -p. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ment The `ere` grammar rule was not atomic, so pest inserted its implicit WHITESPACE and COMMENT rules between the characters of a regular expression literal: `/#/` swallowed the rest of the line as a comment and failed with "expected program", and blanks inside a literal were skipped by the parser. autoconf's config.status relies on `/^[\t ]*#[\t ]*(define|undef)[\t ]+/` and `sub(/#.*/, "")`. Make the rule atomic, as `string` already is. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…) as a file An operand was classified by "is it a regular file?", so /dev/null or a FIFO took the file-versus-directory path: `diff -u /dev/null e` looked for "/dev/null/e" and exited 2. Only an actual directory now gets the directory-operand treatment; everything else is compared as a file. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The group execute position only ever printed x or -, so a set-group-ID file or directory looked the same as one without the bit. POSIX gives it s (executable) or S (not executable), as for set-user-ID in the owner position. The test also pins set-user-ID and the directory t/T forms. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The alternate-access probe used getxattr, which follows a symbolic link, so a link to a file with an ACL was listed as "lrwxrwxrwx+". A link listed as itself is now probed with lgetxattr (acl_get_link_np on macOS/BSD); when -L or -H follows the link, the line describes the target and the probe follows too (acl_get_file on macOS/BSD, which previously never followed). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ites one byte
`LC_ALL=C awk 'BEGIN { printf("%c", 200) }'` wrote C3 88, the UTF-8
encoding of U+00C8, where a single-byte locale needs the byte C8. awk
holds every string as a Rust String and always wrote it out as UTF-8,
and it always decoded input as UTF-8 first, so in the C locale a
character was still a UTF-8 sequence: length("é") was 1, and `.` in a
regular expression matched half of a two-byte character and then
panicked slicing the String at that offset.
Add awk/charset.rs, which decides once from LC_CTYPE whether characters
are bytes, and route every place bytes enter or leave awk through it:
program text, operands, -v assignments and -F, ENVIRON, input records,
standard output, output files and pipes, file and command names, and
the pattern and subject of every regular expression. In a single-byte
locale byte b is the char U+00bb and is written back as byte b, so
input reaches the output unchanged; in a UTF-8 locale nothing changes.
`%c` with a numeric argument writes the byte with that value modulo 256
there, as gawk and mawk do. A record separator regex now matches the
raw input bytes, which also stops input that is not valid UTF-8 from
failing with a UTF-8 error.
The two multibyte tests (character counts, gsub with an empty match)
relied on UTF-8 semantics in the C locale the harness defaults to; they
now name a UTF-8 locale.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
POSIX.2024 sed ("Regular Expressions in sed"): the delimiter "shall not
terminate the RE when it appears within a bracket expression, and shall
have its normal meaning in the bracket expression". The scanner ended
the RE at the first unescaped delimiter wherever it was, so
`s/[/]/X/`, the address `/[/][/*]/` and autoconf's
`s/\.[^/.][^/.]*$//` all failed to compile ("Unmatched [").
Scan a bracket expression in an `s` RE or a context address as a unit:
a leading `]` (after any `^`) is a member, `[:class:]`, `[=equiv=]` and
`[.coll.]` may hold a `]`, and a backslash is copied as the ordinary
character it is there, so `[\/]` matches a backslash or a slash, as in
GNU sed. `\n` inside a bracket expression is still a <newline>, which
GNU-style `[^\n]` scripts rely on. The replacement and the `y` strings
have no bracket expressions and are scanned as before.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rrors restore_sigpipe reset SIGPIPE to the default unconditionally, so a utility started with it ignored (`trap '' PIPE`) was killed anyway and the shell saw 141 instead of a write error. The inherited disposition is now recorded by a constructor that runs before the Rust runtime overwrites it, and an inherited SIG_IGN is left in place. With SIGPIPE ignored, println! panics on EPIPE and exits 101; init_locale now installs a panic hook that turns a failed write to standard output into "UTILITY: write error: REASON" and exit 1, as GNU utilities report it. Other panics go to the previous hook. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…imiter
POSIX.2024 sed: "Any character other than <backslash> or <newline> can
be used instead of a <slash> to delimit the RE and the replacement".
The parser refused letters, digits, <space>, `;` and `{` with
"unterminated `s' command", so `s;a;X;`, autoconf's
`s;[[][0-9][]]\(.*\);\1;` and `y;abc;xyz;` all failed. Refuse only
<backslash> and <newline>. The command is scanned character by
character, so a `;` delimiter is never taken for a command separator.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…th an error grep read its input with read_line into a String, so any line that was not valid UTF-8 failed to read: it was reported as "error reading line N (stream did not contain valid UTF-8)", left out of the search, and the exit status became 2, in every locale, including the C locale and no locale at all (`printf 'x\xff foo\n' | env -i grep foo`). POSIX input files are byte strings; in the C locale every byte is a character. Read each line as bytes, match it with regexec on those bytes (and compare -F patterns as bytes, folding -i a locale character at a time), and write a selected line out byte for byte. In a UTF-8 locale a byte that is not part of a valid sequence is no character, so `.` does not match it, but the rest of the line is searched and the line is written as it is; nothing is reported. Commit 7b96e7f kept the "error reading line N" report for invalid UTF-8 while it fixed other read errors. That report was this bug, so the tests that pinned it (the *_03 and *_05 cases on tests/grep/invalid_line) now expect the valid lines, no message and status 0, which is what GNU grep gives for that file in the C locale; the *_no_messages_with_error_05 tests are renamed to say what they cover now. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
gsub(/b*/, "X") turned "abc" into "XaXXcX": after replacing "b" the
scan resumed at its end, found the empty match there, and replaced it
too. gawk, mawk and the one true awk all give "XaXcX"; an empty match
adjacent to the previous match is not another, non-overlapping match.
Skip it.
bugfix_gsub_multibyte.out pinned the doubled X ("XXéXbX" for
gsub(/a*/, "X") on "aéb"); it is now "XéXbX".
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
There was a problem hiding this comment.
🔵 Needs a closer look
It spans ~90 files and mixes broad mechanical argument-parsing changes with several substantive behavioral features (cp -l, rmdir, ls mode strings, SIGPIPE/byte-oriented rewrites), warranting final human review.
0 open findings
What changed in this PR
This PR bundles two kinds of changes across the posixutils-rs workspace. The dominant, low-risk theme is a mechanical POSIX-conformance fix: adding allow_hyphen_values = true to clap #[arg(...)] option-arguments so that an option-argument may begin with - (XBD 12.2, Guideline 7), with a matching option_argument_may_begin_with_hyphen regression test per utility (driven by the plib::testing::assert_hyphen_option_argument helper). The second theme adds/validates several substantive features and their tests — notably cp -l hard-linking, rmdir --ignore-fail-on-non-empty, and improved ls -l mode-string rendering of setuid/setgid/sticky bits and the ACL + flag for the link itself.
Changes:
- Add
allow_hyphen_values = trueto many utilities' option-argument definitions (compress, uudecode, uucp, uustat, logger, talkd, touch, mkdir, etc.) plus per-utility hyphen-argument regression tests. - Add behavior tests covering
cp -lhard links,rmdir --ignore-fail-on-non-empty(empty-chain removal, silent non-empty skip, permission-denied classification), andls -lspecial permission bits / ACL+on symlinks. - Introduce a shared
assert_hyphen_option_argumenttest helper and a newtree/tests/cp/link.rstest module.
| File | Description |
|---|---|
| xform/compress.rs | Adds allow_hyphen_values to -b and -m so negative-looking values are accepted. |
| xform/uudecode.rs | Adds allow_hyphen_values to -o/--outfile. |
| xform/tests/compress/mod.rs, xform/tests/uue/mod.rs | Hyphen-argument regression tests. |
| uucp/uucp.rs, uucp/uustat.rs | Adds allow_hyphen_values to value-bearing options. |
| uucp/tests/uucp/mod.rs, uucp/tests/uustat/mod.rs | Hyphen-argument regression tests. |
| users/logger.rs, users/talkd.rs | Adds allow_hyphen_values to -f/-p/-t and -s/--invite-timeout. |
| users/tests/logger/mod.rs, users/tests/talkd/mod.rs | Hyphen-argument regression tests. |
| tree/touch.rs | Adds allow_hyphen_values to -d/-t/-r. |
| tree/tests/touch/mod.rs, tree/tests/mkdir/mod.rs | Hyphen-argument regression tests. |
| tree/tests/rmdir/mod.rs | New tests for --ignore-fail-on-non-empty (parents, other errors, permission denied). |
| tree/tests/ls/mod.rs | New tests for -l special-bit mode strings, ACL + on the link itself, and EPIPE handling. |
| tree/tests/cp/mod.rs, tree/tests/cp/link.rs | Registers and adds a cp -l hard-link test module. |
🧠 Review effort: Balanced
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
gettext's po.m4 keeps a msgfmt only if `msgfmt --statistics /dev/null` exits 0. The counts use the singular for one, as GNU does; -v shares the same line. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
-s -r writes the source's path relative to the link's directory, both resolved through existing symbolic links without requiring either to exist, as GNU ln does; libselinux runs `ln -sf --relative`. -r without -s is an error. A single operand links into the current directory under its last component; perl's build runs `ln -s regen-configure/U`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Linux takes the type from the mount table, macOS from f_fstypename.
guile's build reads it with `df -T PATH | awk 'END{print $2}'`.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
UserInfo::from_username and from_uid, spool's User::current and the at and batch tests called getpwnam and getpwuid, whose result lives in a static buffer that a lookup on another thread overwrites; User::current even copied the passwd struct out and read its pointers later. They now use plib::user, which is reentrant and returns owned data. crond's UserInfo used to decode the name and home directory lossily and hand the result to initgroups and chdir. A user whose name or home is not UTF-8 now resolves to no UserInfo, so crond runs nothing for it rather than running it under some other name. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
fuser -u, renice -u and their tests called getpwuid and getpwnam, which return a pointer into a static buffer that another thread's lookup overwrites. They now use plib::user, which calls the _r functions. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
~user expansion called getpwnam, whose result lives in a static buffer that a lookup on another thread overwrites. It now uses plib::user, which calls getpwnam_r and returns the home directory's bytes exactly. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
util-linux's tools/poman-translate.sh runs `sort --check --version-sort`; GNU sort spells POSIX -c as --check. --check=quiet/silent/diagnose-first (GNU's long forms of -C and -c) are not accepted. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Keys compare as gnulib's filevercmp does: digit runs by value, '~' before everything, letters before other bytes, trailing file suffixes compared last, and '.', '..' and dot-names first. -n with -V is refused, as GNU does. -V no longer prints the version; --version still does. util-linux's tools/poman-translate.sh runs `sort --check --version-sort`. A -V key is stored as Text and dispatched on its KeySpec: a third KeyVal variant cost about 2.5% on a 2M-line -n sort. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
env read the inherited environment with std::env::vars(), which panics on
an entry that is not valid UTF-8 (perl's test suite sets
$ENV{k} = "eh zero \xA0" and runs env). Operands were clap Strings, so a
NAME=VALUE assignment, utility name or argument with such bytes was
rejected.
Read the environment with vars_os(), take operands as OsStrings, split
assignments on bytes, and write the dump as raw bytes through a buffered
writer, reporting a write error with exit status 1.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ps wrote its listing through Rust's line-buffered stdout, one write per line. A reader that stops early -- perl's dist/threads/t/join.t reads `ps -f |` up to its own line, closes the pipe and dies if close() reports a failure -- closed while ps was still writing, and ps died of SIGPIPE. procps' stdio fully buffers a pipe, so the whole listing is already in the pipe when the reader sees its first byte. Both the POSIX and the BSD-style listings now write through a 64 KiB BufWriter when stdout is not a terminal, flushed once at the end, with the existing write-error report. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The C field (-f, -l, -o c) was always "-". It is now the CPU time as a percentage of the time since the process started, capped at 99 and right-aligned, as procps computes it. A left-aligned last column (CMD, COMMAND) was padded to its width, so every line ended in blanks and the header was right-aligned far from its values. The last column is no longer padded on the right, and its header starts where its values do. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
kill and nohup read their arguments with std::env::args(), which panics on one that is not valid UTF-8. nice and timeout took the utility and its arguments as clap Strings and rejected such bytes. nohup ignored a HOME, and timeout a PATH, that was not valid UTF-8, as if unset. nohup, nice and timeout now pass the utility and its arguments through as OsStrings, nohup falls back to $HOME by var_os and timeout searches PATH by var_os and split_paths. kill reads args_os and reports a non-UTF-8 argument as an invalid signal or PID. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Deciding which makefile macros to export read the environment with std::env::vars(), which panics on an entry that is not valid UTF-8, so any such entry (perl's test suite sets one) made every make run abort. Only the inherited names are needed: read them with vars_os() and keep the ones that are valid UTF-8, as no other can name a macro. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
at read the environment with std::env::vars(), which panics on an entry that is not valid UTF-8, and wrote the working directory with to_string_lossy(), so a job submitted from a directory whose name is not valid UTF-8 could not cd back to it. The job script is now built as bytes from vars_os() and the directory's own bytes, single-quoted as before. An inherited name the shell cannot assign (not a letter or underscore followed by letters, digits and underscores) is left out instead of being written as a command line. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
dd and find read their arguments with std::env::args(), which panics on
one that is not valid UTF-8, so `dd if=NAME` or `find DIR` over such a
name aborted. find also replaced {} in -exec and -ok with a lossy copy
of the pathname, which names a different file.
dd now takes if= and of= as OsStrings and reports any other non-UTF-8
operand as invalid. find takes its starting points as bytes, hands -exec
and -ok the pathname itself, and reports an expression operand that is
not valid UTF-8 as an error.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each read its arguments with std::env::args(), which panics on one that is not valid UTF-8. who and pr take their file operands as pathnames, byte for byte (and plib's utmpx::load_from_file passes the path's bytes to utmpxname instead of refusing a non-UTF-8 one). tabs, and more with $MORE set, hand the words to clap, which refuses a non-UTF-8 operand with its own message. compress tells zcat and uncompress from a lossy copy of a program name that is not valid UTF-8. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…args sccs, admin, vi/ex, yacc, mailx and c17 parse their command lines by hand as Strings, read with std::env::args(), which panics on an argument that is not valid UTF-8. plib::optarg::args_utf8 reads args_os and reports such an argument as `UTILITY: ARG: argument is not valid UTF-8` with exit status 1, as clap does for the utilities it parses. vi also told ex from vi by its program name through env::args(); it now reads a lossy copy of args_os()'s first word. plib::testing::assert_non_utf8_argument_rejected checks the report. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The user and group stress tests expected a lookup by name to return the same record as a lookup by uid (gid). macOS answers the two from different sources and gives root /bin/bash by name and /bin/sh by uid, so the test failed there with no race involved. Each thread now compares a by-uid answer with a by-uid baseline and a by-name answer with a by-name baseline, both taken before the threads start; the baselines are checked to agree on name and uid (name for groups). With getpwuid/getpwnam/getgrgid/getgrnam swapped back in, the tests still failed 5 runs of 5 on Linux. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… pointer nssbuf::lookup dereferenced the result pointer that getpwuid_r and friends set on success. That pointer always names the caller's own entry, but CodeQL cannot see that and flagged the dereference as an access of a possibly invalid pointer. The result pointer is now only tested for null and (in debug builds) checked to equal the entry; a found entry is read from the MaybeUninit the call filled in. getlogin_r and ttyname_r return no pointers and needed no change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CodeQL flagged the group-id panic in group_name() as cleartext logging of sensitive information, since the id comes from the test user's passwd entry. get_group_id's panic printed a group name taken from the same lookups. Both messages now say only what failed. No other test touched by the plib lookup series formats a looked-up id or name into a panic or print. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ne, and survive a poisoned lock macOS (APFS) refuses a directory name that is not UTF-8 with EILSEQ, so test_at_non_utf8_environment_and_directory now falls back to a plain name there and still checks the non-UTF-8 environment. Its panic had poisoned TEST_MUTEX and failed test_multi_operand_timespec too; every at test now takes the lock with unwrap_or_else(into_inner), as some already did, so one failure no longer cascades. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Input was decoded as UTF-8 with invalid bytes replaced by U+FFFD, so `find . -print0 | xargs -0 rm` with a name holding \xff handed rm a different, usually nonexistent, file. Parse input as bytes in every mode (blank/newline/quote splitting, -0, -I, -L, -n, -s), compare -E's eofstr as bytes, take the utility, its arguments and -I's replstr as OsStrings, and pass each argument to the utility built from the raw bytes. -L now reads lines with read_until, which no longer fails on a non-UTF-8 line. -s and ARG_MAX accounting count bytes, and the queued argument size is kept as a running total instead of being re-summed on every check. -t and -p write the command line's bytes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
An argument too long for -s even alone made xargs run the utility with no arguments in an endless loop unless -x was given; it is now the "argument line too long" error, exit status 1, as in GNU xargs. And only one command was run for the arguments left at end of input, so an argument that overflowed that last batch was silently dropped (`printf 'aaa bbb ccc' | xargs -s 13 echo` lost ccc); the remainder now runs in as many commands as it needs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The default size was {ARG_MAX}-2048 with neither the environment nor the
argument pointers counted, so `seq 1 2000000 | xargs echo` packed one
command too large to exec and failed with E2BIG. POSIX bounds the
combined argument and environment lists by {ARG_MAX}-2048; subtract the
environment (strings, NULs and pointers) from that bound, and default to
GNU's 128 KiB within it. A larger -s is clamped to the bound. The
benchmark now batches exactly as GNU xargs does (114 commands, identical
output) in 0.33 s against GNU's 0.31 s.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…arch timeout ran the operand as given when a file of that name was in the current directory and otherwise joined it to each PATH directory itself. That search treated a name with a slash as a name to look up, so `timeout 5 ./x` with no ./x ran bin/./x from PATH, and it found nothing at all with PATH unset, where execvp uses the default search path. The cwd check never ran the shadowing file, since the bare name it kept was still searched by execvp. Hand the operand to Command, whose execvp does the POSIX lookup: a name with a slash is a pathname, a bare name is searched through PATH (an empty element meaning the current directory), and ENOENT is still exit status 127. nice, nohup, env, xargs, time and find -exec already pass the bare name to Command and were probed to behave the same. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
APFS refuses to create a file whose name is not valid UTF-8 (EILSEQ), so the test that has rm remove such files cannot set up on macOS. The other byte tests pass their bytes through arguments and stdin only and still run everywhere. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The primaries ran fnmatch(3) on a lossy copy of the name, each byte that is not valid UTF-8 replaced by U+FFFD, so in the C locale `-name 'a?b'` missed a file named a\xffb, and a pattern operand that is not valid UTF-8 was refused outright. The expression is now parsed from OsStr tokens; these four patterns are kept as bytes and matched against the raw name or pathname with fnmatch(3) in the current locale, with the same flags as before (neither FNM_PATHNAME nor FNM_PERIOD, per POSIX). Other operands are still read as text and refused when not UTF-8. Results agree with GNU find 4.9 in the C and C.UTF-8 locales. The existing test of the UTF-8 refusal now uses a -user operand. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A test that makes a file, directory or symlink whose name is not valid UTF-8 fails on macOS, where APFS refuses such a name with EILSEQ. create_non_utf8 makes the entry through a caller-supplied creator and returns None, with a note on stderr, when the filesystem refuses the name (EILSEQ, or EINVAL as some filesystems answer), so the test can skip just that part; any other error still panics. non_utf8_names_supported probes a directory for a test whose utility, not the test, makes the name. Unit tests fake the refusal with an injected creator. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ex_invoked_by_non_utf8_name and zcat_invoked_by_non_utf8_name made a symlink named \xffex / \xffzcat to choose the mode, which macOS APFS refuses with EILSEQ. vi and compress select their mode from argv[0] alone, so the tests now spawn the real binary with CommandExt::arg0 and run on every Unix filesystem. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every test that makes a file, directory or symlink with a name that is not valid UTF-8 now does so through plib::testing::create_non_utf8 (or non_utf8_names_supported, where the utility makes the name), and skips that part where the filesystem refuses the name, as macOS APFS does with EILSEQ. This replaces the ad-hoc fallbacks in the at, pax multivolume, patch and sh tests, and fixes tests that would have panicked on APFS: dd, find, nohup, pr, timeout and who. Tests gated to Linux only because of APFS now build everywhere and skip on refusal: find -name/-path byte matching, realpath -e, the four pax special name tests, pax -u -a, and xargs rm. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
No description provided.