Skip to content

tr: speed up -d - #14897

Open
charly-mp wants to merge 5 commits into
uutils:mainfrom
charly-mp:tr-faster-delete
Open

charly-mp wants to merge 5 commits into
uutils:mainfrom
charly-mp:tr-faster-delete

Conversation

@charly-mp

Copy link
Copy Markdown

Mean time in ms over 40 runs, 100 MB input unless noted (Ryzen 9 9950X, pinned core, hyperfine -N).

Command GNU 9.12 trunk this PR vs trunk
tr -d , (CSV) 90.1 87.4 33.5 2.6x
tr -d ' ' 117.8 109.9 33.1 3.3x
tr -d '\n' 47.8 58.7 33.8 1.7x
tr -d '\n' (10 KB lines) 49.8 50.4 9.4 5.4x
tr -d '\r' (CRLF, 16 MB) 7.4 8.9 5.6 1.6x
tr -d '\r' (nothing to delete) 57.3 9.1 7.7 1.2x
tr -d aeiou 272.8 261.8 31.9 8.2x
tr -cd 'a-z\n' 130.3 136.8 32.0 4.3x
tr -cd '[:print:]' (binary) 332.0 316.6 32.6 9.7x

@xtqqczze

xtqqczze commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

@charly-mp Please share the hyperfine command line for your measurements.

@codspeed

codspeed Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Merging this PR will improve performance by 14.52%

⚡ 3 improved benchmarks
✅ 390 untouched benchmarks
🆕 2 new benchmarks
⏩ 54 skipped benchmarks1

Performance Changes

Mode Benchmark BASE HEAD Efficiency
⚡ Simulation three_39_bit_primes 533 ms 399.2 ms +33.51%
⚡ Simulation five_38_bit_primes 1.8 s 1.6 s +7.86%
⚡ Simulation tr_single_char_replace 14.5 ms 13.9 ms +4.3%
🆕 Memory tr_delete_single_char N/A 35.4 KB N/A
🆕 Simulation tr_delete_single_char N/A 103.9 ms N/A

Tip

Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.


Comparing charly-mp:tr-faster-delete (09734be) with main (2b85f6f)

Open in CodSpeed

Footnotes

  1. 54 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

@charly-mp

charly-mp commented Sep 27, 2026 •

Copy link
Copy Markdown
Author

@charly-mp Please share the hyperfine command line for your measurements.

here it is:

LC_ALL=C taskset -c 12 nice -n -20 hyperfine -N --warmup 5 --runs 40 \
    --input text.txt --output=null \
    -n gnu   "gnutr -d ' '" \
    -n trunk "./trunk/coreutils tr -d ' '" \
    -n pr    "./pr/coreutils tr -d ' '"

text.txt is ~100 MB of lines of 3 to 15 random dictionary words.
I also have csv.txt which is ~100 MB of 6-field CSV lines.

@github-actions

github-actions Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

GNU testsuite comparison:

Congrats! The gnu test tests/id/setgid is no longer failing!

Comment thread src/uu/tr/src/operation.rs Outdated

// Check if this is a simple single-character translation
if let Some((source, target)) =
find_single_change(&self.translation_table, |i, &val| val != i as u8)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

could translate also compute this once in new(), like you did for delete?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

good point, 7% fewer instructions on tr_delete_single_char

Comment thread src/uu/tr/src/simd.rs Outdated
output.extend_from_slice(&input[start..]);
} else if count < input.len() {
output.extend(input.iter().filter(|&&b| b != delete_char).copied());
let mut keep = [true; 256];

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

DeleteOperation already has a keep_table, could you pass it here instead of rebuilding one for each chunk?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done

Deleting filtered bytes with a branch per byte, which mispredicts on
real data. Store every byte in a small block and only advance past the
kept ones. After a block that kept almost nothing, go back to the
branch: it is then well predicted and stores less. When the single
byte to delete is rare, copy the runs between its occurrences instead.

Add a benchmark for deleting a single byte.

Mean of 40 runs in ms, 100 MB input unless noted (Ryzen 9 9950X,
GNU 9.12, pinned core, hyperfine -N):

                                GNU   trunk   opti
  tr -d ,  (CSV)               90.1    87.0   33.7
  tr -d ' '                   118.0   109.2   33.8
  tr -d '\n'                   47.8    57.8   37.0
  tr -d '\r'  (CRLF, 16 MB)     7.4     8.9    5.7
  tr -d aeiou                 272.4   261.0   32.9
  tr -cd 'a-z\n'              130.5   137.9   32.9
  tr -cd '[:print:]' (binary) 332.0   317.6   35.0
The single byte to delete or replace was looked up in the whole table,
with a Vec allocation, on every read. Do it once when building the
operation.

About 1,300 fewer instructions per 32 KiB read: up to 3% faster when
nothing is deleted, 7% fewer instructions on the single byte replace
benchmark, noise otherwise.
When deleting or replacing a single byte, a chunk without it was still
copied to the output buffer before being written. ChunkProcessor now
returns the bytes to write, so such a chunk is written from the input.

Mean of 30 paired runs on a pinned core, 100 MB of text without CR:

                   before   after
  tr -d '\r'      8.54 ms  7.67 ms
  tr '\r' '\n'    8.84 ms  7.97 ms
Stdout is line buffered: every chunk was searched for its last newline,
and the bytes after it cost an extra write. Write to the file
descriptor directly, as cat and tee do.

Keep write_output out of line: once inlined into the per-byte loop of
translate_input, it made the translator state go to memory on every
byte, doubling the time of tr -s.

Give test_broken_pipe_no_error more than a pipe buffer to write: it
only hit the broken pipe when the output came late, which was already
racy on main (1 failure in 60 runs) and became frequent here.

Mean of 30 paired runs on a pinned core, 100 MB input:

                          before     after
  tr -d '\n'   (text)   36.23 ms  32.64 ms
  tr -d '\r\n' (CSV)    35.66 ms  31.94 ms

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants