This is the reference implementation accompanying the paper Pitch Smoothing Using Relative Interval Networks, submitted to ICASSP 2027. We propose a method for smoothing pitch estimates by combining absolute pitch measurements with multi-hop relative pitch differences, using a network-flow linear program for optimal fusion. Specifically, we solve the following objective function:
where
With pixi (recommended):
pixi installOr with pip:
pip install rin-pitchimport numpy as np
from rin import smooth_pitch
# x: mono waveform, sr: sample rate
# f0: (M,) absolute pitch in cents on a 20 ms grid, NaN = unvoiced
# strength: (M,) voicing confidence in [0, 1]
hop_length = int(0.02 * sr)
# the paper's pipeline in one call: VQT relative diffs -> dual LP fusion
# -> voicing, with the paper's fixed settings (hops 1,5; Pearson xcorr;
# arcsin x peak2mean weighting)
f0_smooth, voicing = smooth_pitch(x, f0, strength, sr, hop_length)The package never converts pitch units: absolute and relative estimates must share one pitch domain (the built-in estimator outputs cents), and converting to or from your own domain (Hz, MIDI, ...) is your responsibility.
If your tracker reports Hz, convert before calling; smooth_pitch expects cents and will not convert for you:
voiced = np.isfinite(f0_hz)
f0_cents = np.where(voiced, 1200 * np.log2(np.where(voiced, f0_hz, 1.0)), np.nan)
f0_smooth, voicing = smooth_pitch(x, f0_cents, strength, sr, hop_length)Arrays in, arrays out: the package does no file loading and no caching.
One high-level function plus three cores:
rin.smooth_pitch(x, f0, strength, sr, hop_length, hops=(1, 5), ...)runs the paper's pipeline in one call: VQT relative diffs, dual LP fusion, and voicing.f0in cents with NaN = unvoiced; returns(f0_smooth, voicing)in cents.f0defines the frame grid: its length sets the output length and is never checked against the audio, since trackers disagree on the exact frame count, and relative edges past its end are dropped. Each stage is swappable viadifference_estimator,solver, andvoicing_estimatorkeyword arguments (any callable obeying therin.interfacescontracts); configure a stage withfunctools.partial, e.g.difference_estimator=partial(vqt_diff_calculator, max_diff_cents=500.0).
Three core functions (for custom wiring):
rin.vqt_diff_calculator(x, sr, hop_length, hops=(1,), ...)returns multi-hop relative pitch differences (cents) with confidences. Uses the paper's fixed estimation path: Pearson (mean-subtracted) normalized cross-correlation of VQT magnitude slices with the arcsin x peak2mean confidence weighting.max_diff_cents(600),bins_per_octave(36),n_bins(252), and other VQT options are plain kwargs, so pass your own.max_diff_centsmust stay inside the VQT's range: it has to buy at least one bin of search and fewer thann_binsof it, or the correlation window runs off the spectrogram.rin.lp_smoother(abs_estimates, abs_confidences, rel_edges, rel_estimates, rel_confidences)does network-flow LP fusion (dual min-cost circulation, HiGHS); absolute and relative estimates must share one pitch domain (caller's choice, e.g. cents).rin.estimate_voicing(abs_confidences, rel_edges, rel_confidences)returns per-frame voicing probabilities from the two confidence streams.
Each core is a plain function obeying a contract in rin.interfaces (DifferenceEstimator, Solver, and VoicingEstimator are Callable type aliases).
Implement your own function with the same signature and pass it to smooth_pitch, or call it directly in your own wiring.
No classes or inheritance needed; any callable (function, lambda, functools.partial, callable object) works:
from functools import partial
from rin import smooth_pitch, vqt_diff_calculator
def my_estimator(x, sr, hop_length, hops):
# -> (edges (E,2) int, estimates (E,) cents, confidences (E,) in [0,1])
...
def my_solver(abs_estimates, abs_confidences, rel_edges, rel_estimates, rel_confidences):
# -> smooth_pitch (M,) cents
...
def my_voicing(abs_confidences, rel_edges, rel_confidences):
# -> voicing (M,) in [0,1]
...
# One call, custom stages (tune a stage via functools.partial):
f0_smooth, voicing = smooth_pitch(
x,
f0,
strength,
sr,
hop_length,
difference_estimator=partial(vqt_diff_calculator, max_diff_cents=500.0),
solver=my_solver,
voicing_estimator=my_voicing,
)
# ...or wire them by hand:
edges, estimates, confidences = my_estimator(x, sr, hop_length, hops=(1, 5))
smooth_cents = my_solver(f0_cents, abs_conf, edges, estimates, confidences)
voicing = my_voicing(abs_conf, edges, confidences)pixi run test # pytest with branch coverage
pixi run lint # ruff check
pixi run docstrings # numpydoc validation of the public API
pixi run format # ruff format
pixi run build # sdist + wheel
pixi run smoke # install the built wheel in a clean venv and check its versionCI runs lint and docstrings once, and test on Linux, macOS and Windows across Python 3.10-3.13 (pixi run -e py310 test reproduces one cell locally).
The version comes from the git tag (no file in the repo declares one), so a release is git tag v1.1.0 && git push --tags, which builds and publishes to PyPI.
MIT