tr: speed up -d - #14897
tr: speed up -d#14897charly-mp wants to merge 5 commits into
Conversation
|
@charly-mp Please share the hyperfine command line for your measurements. |
Merging this PR will improve performance by 14.52%
Performance Changes
Tip Curious why performance improved? Comment Comparing Footnotes
|
here it is: LC_ALL=C taskset -c 12 nice -n -20 hyperfine -N --warmup 5 --runs 40 \
--input text.txt --output=null \
-n gnu "gnutr -d ' '" \
-n trunk "./trunk/coreutils tr -d ' '" \
-n pr "./pr/coreutils tr -d ' '"text.txt is ~100 MB of lines of 3 to 15 random dictionary words. |
b6eeeaf to
d9d4e8a
Compare
|
GNU testsuite comparison: |
|
|
||
| // Check if this is a simple single-character translation | ||
| if let Some((source, target)) = | ||
| find_single_change(&self.translation_table, |i, &val| val != i as u8) |
There was a problem hiding this comment.
could translate also compute this once in new(), like you did for delete?
There was a problem hiding this comment.
good point, 7% fewer instructions on tr_delete_single_char
| output.extend_from_slice(&input[start..]); | ||
| } else if count < input.len() { | ||
| output.extend(input.iter().filter(|&&b| b != delete_char).copied()); | ||
| let mut keep = [true; 256]; |
There was a problem hiding this comment.
DeleteOperation already has a keep_table, could you pass it here instead of rebuilding one for each chunk?
Deleting filtered bytes with a branch per byte, which mispredicts on
real data. Store every byte in a small block and only advance past the
kept ones. After a block that kept almost nothing, go back to the
branch: it is then well predicted and stores less. When the single
byte to delete is rare, copy the runs between its occurrences instead.
Add a benchmark for deleting a single byte.
Mean of 40 runs in ms, 100 MB input unless noted (Ryzen 9 9950X,
GNU 9.12, pinned core, hyperfine -N):
GNU trunk opti
tr -d , (CSV) 90.1 87.0 33.7
tr -d ' ' 118.0 109.2 33.8
tr -d '\n' 47.8 57.8 37.0
tr -d '\r' (CRLF, 16 MB) 7.4 8.9 5.7
tr -d aeiou 272.4 261.0 32.9
tr -cd 'a-z\n' 130.5 137.9 32.9
tr -cd '[:print:]' (binary) 332.0 317.6 35.0
The single byte to delete or replace was looked up in the whole table, with a Vec allocation, on every read. Do it once when building the operation. About 1,300 fewer instructions per 32 KiB read: up to 3% faster when nothing is deleted, 7% fewer instructions on the single byte replace benchmark, noise otherwise.
When deleting or replacing a single byte, a chunk without it was still
copied to the output buffer before being written. ChunkProcessor now
returns the bytes to write, so such a chunk is written from the input.
Mean of 30 paired runs on a pinned core, 100 MB of text without CR:
before after
tr -d '\r' 8.54 ms 7.67 ms
tr '\r' '\n' 8.84 ms 7.97 ms
Stdout is line buffered: every chunk was searched for its last newline,
and the bytes after it cost an extra write. Write to the file
descriptor directly, as cat and tee do.
Keep write_output out of line: once inlined into the per-byte loop of
translate_input, it made the translator state go to memory on every
byte, doubling the time of tr -s.
Give test_broken_pipe_no_error more than a pipe buffer to write: it
only hit the broken pipe when the output came late, which was already
racy on main (1 failure in 60 runs) and became frequent here.
Mean of 30 paired runs on a pinned core, 100 MB input:
before after
tr -d '\n' (text) 36.23 ms 32.64 ms
tr -d '\r\n' (CSV) 35.66 ms 31.94 ms
6494452 to
281ebc3
Compare
Mean time in ms over 40 runs, 100 MB input unless noted (Ryzen 9 9950X, pinned core,
hyperfine -N).tr -d ,(CSV)tr -d ' 'tr -d '\n'tr -d '\n'(10 KB lines)tr -d '\r'(CRLF, 16 MB)tr -d '\r'(nothing to delete)tr -d aeioutr -cd 'a-z\n'tr -cd '[:print:]'(binary)