Skip to content

Repository files navigation

ARDFTSRC logo

ARDFTSRC

Crates.io Docs.rs Tests

A rust implementation of the Arbitrary Rate Discrete Fourier Transform Sample Rate Converter (ARDFTSRC) algorithm.

ardftsrc is a high-quality audio sample-rate converter, and is appropriate for both realtime and offline resampling.

Generally ardftsrc is preferred over other resamplers when quality is paramount. Although it is generic over both f32 and f64, it is highly recommended to use it with f64, even when processing an f32 audio stream.

It is ranked among the top scoring resampling libraries on HydrogenAudio Sample Rate Conversion Comparison, topping out with a balanced score of 99.84%. It is more compute intensive than other resamplers, so consider sinc rubato if you want more efficiency. See PERFORMANCE.md for a detailed speed and quality comparison vs rubato.

Quick Start

Use InterleavedResampler::process_all to resample a complete interleaved audio stream for a single track.

use ardftsrc::{InterleavedResampler, PRESET_HIGH};

fn resample_all(input: &[f32], in_rate: usize, out_rate: usize, channels: usize) -> Vec<f32> {
    // When using a preset other than "FAST", f64 processing is preferred.
    let input_f64: Vec<f64> = input.iter().map(|v| *v as f64).collect();

    let config = PRESET_HIGH
        .with_input_rate(in_rate)
        .with_output_rate(out_rate)
        .with_channels(channels);

    let mut resampler = InterleavedResampler::<f64>::new(config).unwrap();

    let output = resampler.process_all(&input_f64).unwrap();

    // Convert back to the original interleaved f32
    output.interleave().into_iter().map(|v| v as f32).collect()
}

Chunk Resampling

Use chunk resampling when you can control both read and write buffer sizes. Query input_buffer_size() and output_buffer_size() and size your input and output slices to the sizes required. The chunk API is more efficient than the streaming API and is preferred when you are not doing live resampling.

There are two chunked resamplers depending on the shape of your audio:

  1. InterleavedResampler - for interleaved audio
  2. PlanarResampler - for planar audio.

Internally ardftsrc uses planar representation, so PlanarResampler is more efficient, but if you're already working with interleaved audio, prefer InterleavedResampler since it has an optimized de-interleave / re-interleave path. Working with all chunked resamplers is the same:

  1. Create the resampler with let resampler = Resampler::new(config)
  2. Query the required input buffer size and output buffer size with resampler.input_buffer_size() and resampler.output_buffer_size()
  3. Call process_chunk(...) for each chunk, using the appropriate buffer sizes.
  4. Call process_chunk_final(...) for the final chunk, it can be undersized.
  5. Finally, call finalize(...) once per stream to emit delayed tail samples and reset stream state.
use ardftsrc::{InterleavedResampler, PRESET_GOOD};

fn resample_chunked(input: Vec<f32>, in_rate: usize, out_rate: usize, channels: usize) -> Vec<f32> {
    // When using a preset other than "FAST", f64 processing is preferred.
    let input_f64: Vec<f64> = input.into_iter().map(|v| v as f64).collect();

    let config = PRESET_GOOD
        .with_input_rate(in_rate)
        .with_output_rate(out_rate)
        .with_channels(channels);

    let mut resampler = InterleavedResampler::<f64>::new(config).unwrap();

    // Get the input and output chunk sizes
    // You must read and write in these buffer sizes
    let input_chunk_size = resampler.input_buffer_size();
    let output_chunk_size = resampler.output_buffer_size();
    let mut out_buf = vec![0.0_f64; output_chunk_size];
    let mut out_f64 = Vec::<f64>::new();
    let mut offset = 0;

    // Process whole chunks in the size of input_chunk_size
    while offset + input_chunk_size <= input_f64.len() {
        let chunk = &input_f64[offset..offset + input_chunk_size];

        // Process the chunk
        let written = resampler.process_chunk(chunk, &mut out_buf).unwrap();

        // Process output
        out_f64.extend_from_slice(&out_buf[..written]);
        offset += input_chunk_size;
    }

    // The final chunk can be undersized (or even zero sized)
    let final_chunk = &input_f64[offset..];

    // Process Output
    let written = resampler.process_chunk_final(final_chunk, &mut out_buf).unwrap();
    out_f64.extend_from_slice(&out_buf[..written]);

    // After processing the final chunk, you must call "finalize()" to get tail content.
    // finalize() also resets the resampler instance so it can be used again.
    let written = resampler.finalize(&mut out_buf).unwrap();
    out_f64.extend_from_slice(&out_buf[..written]);

    // Convert back into f32
    out_f64.into_iter().map(|v| v as f32).collect()
}

Gapless Context

For adjacent tracks, you can set edge context before processing:

  • pre(Vec<T>): tail frames from the previous track
  • post(Vec<T>): head frames from the next track

post(...) may be called any time while the current stream is still active, but it must be set before process_chunk_final(...).

This enables live gapless handoff: while track A is streaming, once track B is known you can call post(...) on A with B's head samples so A's stop-edge uses real next-track context.

Realtime Resampling

ardftsrc-rs provides both rodio integration via RodioResampler (rodio feature) and the ability to build your own custom realtime audio resampling pipeline via RealtimeResampler.

Rodio integration

Enable the rodio feature to use RodioResampler to wrap a rodio::Source and resample it in realtime in your rodio pipeline.

When playing from a buffered audio source such as a file or a buffered stream, it is recommended to use config.with_rodio_fast_start(true), which will avoid initial output delay by pulling samples from the upstream source to prime the resampler. For very-realtime sources such as microphones or similar, do not enable fast-start.

#[cfg(feature = "rodio")]
{
    let stream = rodio::DeviceSinkBuilder::open_default_sink()?;
    let mixer = stream.mixer();

    let tone = rodio::source::SignalGenerator::new(
        NonZero::new(44_100 as u32).unwrap(),
        400, // 400 Hz
        rodio::source::Function::Sine,
    )
    .take_duration(Duration::from_secs(3.0));

    let config = PRESET_FAST.with_channels(1).with_input_rate(44_100).with_output_rate(48_000);
    let resampled_tone = RodioResampler::new(tone, config)?;

    mixer.add(resampled_tone);
    std::thread::sleep(Duration::from_secs(4));
}

More examples can be found:

Batching

Use batching when you have multiple full tracks to convert with the same configuration.

Enable the rayon feature to parallelize work across tracks.

use ardftsrc::{InterleavedResampler, PRESET_GOOD, PlanarVecs};

fn resample_tracks(
    inputs: &[&[f64]],
    in_rate: usize,
    out_rate: usize,
    channels: usize,
) -> Vec<PlanarVecs<f64>> {
    let config = PRESET_GOOD
        .with_input_rate(in_rate)
        .with_output_rate(out_rate)
        .with_channels(channels);

    let driver = InterleavedResampler::<f64>::new(config).unwrap();

    // Independent tracks (podcasts, unrelated files, etc.).
    let _independent = driver.batch(inputs).unwrap();

    // Gapless sequence (album tracks played back-to-back).
    let gapless = driver.batch_gapless(inputs).unwrap();

    // Return one of the two results based on your use case.
    gapless
}

Quality Tuning and Presets

ARDFTSRC is built for quality over speed, and despite supporting both f32 and f64 should almost always be run as f64. To resample f32 audio, it is recommended to convert f32 samples to f64, resample them using InterleavedResampler<f64> or PlanarResampler<f64>, then convert back to f32.

If you want better performance than what this project offers, consider using a sinc resampler such as rubato.

Presets are pre-vetted Config for various quality levels.

let config = ardftsrc::PRESET_GOOD
  .with_input_rate(44_100)
  .with_output_rate(48_000)
  .with_channels(2);
Preset Quality Bandwidth Recommended use Quality metrics
PRESET_FAST 512 0.832 Fast preset for realtime workloads. f32, f64
PRESET_GOOD † 1878 0.911 Balanced preset for realtime quality. f64
PRESET_HIGH 73622 0.987 High quality for offline use. f64
PRESET_EXTREME 524514 0.995 Maximum quality, intended for offline use. f64

† You should probably use PRESET_GOOD. It's fast, has very high quality metrics, and has lower pre-ringing artefact as compared PRESET_HIGH and PRESET_EXTREME.

Extreme Downsampling Ratios

If you need to support very large downsampling ratios (e.g. 192kHz → 8kHz), set .with_decimate(true) to speed things up. It enables a conservative pre-decimator (only at ratios of 4:1 or higher) that applies progressive 2:1 decimations before the main resampling stage. It is designed so the FFT stage still performs at least a genuine 2:1 reduction of its own, and respects the configured bandwidth.

let config = ardftsrc::Config::new(192_000, 8_000, 1).with_decimate(true);

It's a speed/memory optimization, not a way to reduce buffering or latency — for that, lower quality instead. If you do lower quality for a large ratio, turn decimate on too: it keeps a low-quality conversion sounding good at ratios where it would otherwise struggle.

Transition-Band Aliasing

By default the low-pass transition ends at the lower Nyquist frequency and everything beyond it is suppressed. Setting an alias_floor lets the transition extend past Nyquist, which makes it wider and reduces ringing. Energy in the extended region is folded back (downsampling) or imaged (upsampling), but only down to the floor, and never into the passband set by bandwidth.

// Fold only down to where the filter response is -3 dB.
let config = ardftsrc::Config::new(48_000, 44_100, 2).with_alias_floor_db(-3.0);

This is not the same as removing the low-pass filter: the passband is unchanged, and content above the extended stopband is still suppressed. Alias rejection is intentionally reduced. Pre-decimation stages stay strict.

On the command line, use -a / --allow-aliasing (equivalent to --alias-floor-db -3), --alias-floor <fraction>, or --alias-floor-db <dB>.

Feature Flags

Flag Enables Default
rodio rodio integration via rodio::RodioResampler No
rayon Parallelized resampling (batch() and process_all() APIs) No
avx FFT AVX SIMD Yes
sse FFT SSE SIMD Yes
neon FFT NEON SIMD for ARM / Mac Yes
wasm_simd FFT WebAssembly SIMD Yes
audioadapter Experimental audioadapter support No
high_precision Double-double, f128, and f256 precision FFT backends for extreme quality No

Runtime feature detection is in place for all SIMD except webassembly.

Command Line

The workspace includes a small utility cli, ardftsrc-rs, for WAV/FLAC sample-rate conversion.

You can use this as a utility, or use it to benchmark this project.

Processing defaults to f64. Pass --f32 for 32-bit float processing (quality is capped at 8192, so --preset high and --preset extreme are rejected). --f32 decodes, resamples, and encodes in f32 without an f64 round-trip. A compatible Vulkan GPU is used automatically when one is available: f64 needs shaderFloat64 (not available on Apple GPUs), while --f32 can run on any real GPU. Pass --cpu to force CPU even when a GPU is available. --gpu-group-chunks and --gpu-ring-slots tune GPU batching (defaults are both 4). Configurations the GPU backend cannot run (--decimate, --high-precision) stay on CPU. Pass --high-precision double-double|f128|f256 to use a high-precision FFT backend (much slower; for extreme quality).

RUSTFLAGS="-C target-cpu=native" cargo build --release
./target/release/ardftsrc-rs --help
./target/release/ardftsrc-rs --input in.wav --output out.flac --output-rate 48000 --preset high
./target/release/ardftsrc-rs --input in.wav --output out.flac --output-rate 48000 --preset good --f32
./target/release/ardftsrc-rs --input in.wav --output out.flac --output-rate 48000 --preset extreme \
    --taper-type beta_cdf --alpha 10 --beta 10 --dd-fft --decimate --phase -0.5

Additional taper profiles

Select --taper-type tanh --alpha 3 for an endpoint-normalized hyperbolic tangent transition. With the bessel feature enabled, --taper-type kbd selects the descending half of a standard Kaiser–Bessel-derived window, and --taper-type half_kaiser (also half-kaiser) selects a descending half-Kaiser shifted and scaled to reach zero at the stopband. Both default to --alpha 6 and use Kaiser beta = pi × alpha, like the existing cumulative bessel taper. The library variants are TaperType::Tanh(alpha), TaperType::Kbd(alpha), and TaperType::HalfKaiser(alpha). All require finite, positive alpha.

Quality Reporting

Detailed quality reports are available for each preset:

You can generate reports using the ardftsrc-report command, a CLI that generates per-preset quality reports combining two independent measurements:

  • HydrogenAudio scores: runs ardftsrc through the HydrogenAudio SRC, test suite. Requires GNU Octave (with the signal and image packages) on PATH.

  • THD+N (total harmonic distortion + noise): Measure frequency, amplitude, sample-rate pair, and preset, fitting each resampled output against a steady-state sine to measure distortion.

cargo run -p ardftsrc-report --release -- all --out-dir reports

Contributing

Contributions are welcome!

Architectural Overview

At a high level there are two layers:

  • ArdftsrcCore<T> is the core DSP engine. It owns FFT and runs the core ARDFTSRC algorithm. It is private.
  • PlanarResampler<T> and InterleavedResampler<T> are fixed-size chunk resamplers. They own one ArdftsrcCore per channel and expose full-buffer, chunked, and batch processing APIs for planar or interleaved audio.
  • AdapterResampler<T> is optional behind the audioadapter feature and adapts generic audioadapter inputs and outputs onto the chunk resampling core. Right now there are performance issues with this.
  • RealtimeResampler<T> provides arbitrary-size sample buffering for live resampling.

Golden Hashes

The golden_hashes test validates resampler determinism against checked-in golden outputs in test_wavs/golden_hashes.<arch>.json. It is intended to catch unintended behavior changes.

Run it with:

cargo test -p ardftsrc --release --features=rayon golden_hashes -- --nocapture

To regenerate test_wavs/golden_hashes.<arch>.json:

rust-script scripts/generate_golden_hashes.rs

Updates to test_wavs/golden_hashes.<arch>.json are allowed, but only when accompanied by verifiable quality improvements demonstrated with the HydrogenAudio SRC test suite.

AI Usage Policy

AI use is allowed for the following:

  1. Code exploration and understanding
  2. Generating tests
  3. Creating normal code / function blocks as long as the code is then manually and carefully hand-edited by a human

Development TODOs:

  1. Add bindings to other languages, python, cpp, c#, ts (wasm) etc.
  2. Investigate why the optional audioadapter interface appears to be much slower than other paths.

About

ARDFTSRC - Arbitrary Rate Discrete Fourier Transform Sample Rate Converter

Resources

Stars

10 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages