Quantization: rules, sites, and open violations¶
A survey of every place doppler converts between fixed-width integers and floating point, written to settle the rounding question for every kind of conversion, not only the NCO's. It is a discussion page: it records what was decided, what the code does today, and what still disagrees with the rule. The decisions it proposes are not made yet.
Code links are permalinks to commit 51d34a29, so
their line numbers stay correct as the tree moves.
What was settled¶
doppler#1117, 2026-09-12, settled two rules, one for each side of a single question: can the converted value leave its range?
- An unclamped value going into wrapping (modular) arithmetic, like the
NCO: truncate. The rationale is
NCO design §7, "Why truncation, precisely".
Truncation has no tie to break, so it is bit-identical on every host.
doppler builds with
-ffast-math, and on FMA targets a rounding form like+ 0.5gets fused into the multiply, so the tie is decided by the compiler. Also,(uint32_t)of a value rounded up to 2³² is undefined behaviour: x86 wraps it to 0, arm64 saturates it. - A clamped sample quantiser: scale by 2^(N−1), clamp, then
lroundf. The clamp happens first, so the value can't wrap.lroundfrounds half away from zero regardless of the host's rounding mode (C99 §7.12.9.7), so it is deterministic. Truncating a bipolar signal instead costs 6.02 dB (6.0 dB measured on #1117) and leaves a ±1 LSB dead zone at zero. Full scale is 2^(N−1) in both directions: +1.0 maps one step past the top code and saturates, which is how the converter is defined.
The six kinds of conversion¶
| # | Kind | Rule | Implemented once in | Breaks the rule |
|---|---|---|---|---|
| A | float → phase word | truncate: clamp at the top, or reduce mod 2³² for a resampling step | nco_phase_units, nco_phase_units_mod, nco_norm_freq_to_inc; used by nco, lo, dll, resamp |
symsync |
| B | phase word → table index / float | truncate (phase truncation in the sine lookup); the conversion to float is exact | lo lookup, nco nmax scaling, nco_word_to_norm |
dll's own / 2³² twice (exact, only duplicated) |
| C | float sample → integer | 2^(N−1), clamp, lroundf |
the six f32_to_*_step, e.g. f32_to_i16, f32_to_uq15; called by wfm_sink and wfm_writer |
CIC encoder, ADC vector path |
| D | integer sample → float | multiply by 1/2^(N−1) | six *_to_f32, e.g. i16_to_f32; wfm_reader, stream |
none; int32 → float keeps only 24 bits by nature |
| E | fixed-point requantise (acc >> n) |
round half up, (x + 2^(n−1)) >> n, then saturate |
arith, e.g. mul_q15; hbdecim_q15 output |
CIC output shift |
| F | filter coefficients (design time) | clamp, lround |
hbdecim_q15 coefficients |
none |
Kind A already has a gate, check_phase_conversion_sites.py: any
2³² constant outside nco_core.h fails unless it's listed in its
allowlist, and that list may only shrink. Kinds C and E have no
equivalent gate.
The cvt vector loops (f32_to_*_steps) all call the scalar step, so a
vector path can't round differently from its scalar one. The ADC below is
the one family where it does.
The four violations¶
-
symsync: a private, truncating phase conversion with undefined behaviour at
sps = 1.nominal_incis(uint32_t)(4294967296.0 / s), andsymsync_initturnssps = 0into 1. At 1 that cast is undefined: x86 gives 0, arm64 gives0xFFFFFFFF. The gate's allowlist already marks it "VIOLATION", andnco_core.hcites it as one of the cases that motivated a single home. It was never fixed, and no issue tracks it. Proposed fix: refusesps < 2. A Gardner timing detector needs two samples per symbol anyway, and removing the cast shrinks the allowlist by one. -
The CIC encoder truncates.
cic_core.hclamps, then converts with a bare(int16_t)sr. It's a private copy off32_to_uq15_step(the same offset-binary encoding, plus headroom) that truncates where that function rounds: the 6 dB defect #1117 fixed everywhere else. QUANTIZATION.md §2.4 documents this truncation, while §3.1 says the Q15 encoder rounds. Proposed fix: callf32_to_uq15_stepwith scale32768 / CIC_PAPR_HEADROOM. -
The CIC output shift floors. The decoder computes
(uint16_t)(re >> shift)with no rounding bias. Because the value is offset-binary, that floor is a −½ LSB DC offset on the signed value. Proposed fix: add1 << (shift − 1)before the shift, the same form as kind E. Both CIC fixes change its numbers, so its validation evidence has to be regenerated. -
The ADC's vector and scalar paths disagree. With dither off,
adc_stepscomputes its vector body in float (llroundfof a float product), while the tail andadc_stepcompute in double. They also setclippedat different points: the vector path after rounding, the scalar path before. So a sample's code depends on where it falls in the block. Measured with two million uniform samples:bits dBFS samples that differ largest difference 8–24 0.0 0 0 12 −3.0 0.005% 1 LSB 16 −3.0 0.070% 1 LSB 20 −3.0 1.134% 1 LSB 24 −3.0 18.079% 1 LSB At 0 dBFS the scale is a power of two, so float and double agree exactly. Any other dBFS gives a scale float can't represent exactly. Proposed fix: do the vector multiply in double, and set
clippedfrom the same value in both paths.
Related, found in the same review: two of the twelve cvt cores abort on
out-of-memory where the other ten raise MemoryError
(doppler#1380), and the
case for writing each converter family's kernel once is on
just-makeit#1310.
Documentation today¶
The rule is spread across three pages, and none of them covers every kind:
- QUANTIZATION.md is from July, before #1117, and it documents the CIC truncation next to a formula that rounds (§2.4 vs §3.1).
- Fixed-point guide §8 covers only kind E.
- NCO design §7 covers only kind A.
Proposal¶
- One "Conversion rules" section in QUANTIZATION.md, built on the table above: each kind, its rule, why, and the one function that implements it. The fixed-point guide and the NCO page link to it instead of restating it.
- One fix per violation, as described above. The two CIC changes move its numbers and need its validation evidence regenerated.
- Gates for kinds C and E, so they have the same single-home check kind A has, with an allowlist that can only shrink.
Open decision: how ties break¶
The float quantisers (kind C) break ties away from zero (lroundf). The
integer requantisers (kind E) break them upward (+ 2^(n−1)).
- Recommended: keep both and document why. Each is the standard choice for
its domain. With float input an exact tie is rare, and
lroundfis symmetric about zero. With integer input a tie happens once every 2^n values, and half-up is the one-add form every DSP uses, at a cost of at most 2^−(n+1) LSB of bias. - Alternative: round half away from zero everywhere. One rule, no bias, at the cost of a sign-dependent branch or select in every integer requantiser.