<divclass="eli5"><b>ELI5.</b> The dealer keys the binomial coefficients of (center + eta) up to a chosen degree. After eta is public, a dot with those shares is the monomial or the polynomial, with no further tree walk.</div>
<divclass="eli5"><b>ELI5.</b> An n-bit limb is rewritten into another modulus, a field, or a P-256 scalar by an exact map on the opened residue. The value does not go through floating point, and the map does not expand another key.</div>
<divclass="eli5"><b>ELI5.</b> Powers of a public center are already shared. Shifting them by the opened eta, the binomial way, evaluates the polynomial at the secret. The degree here is fixed in the template.</div>
<divclass="eli5"><b>ELI5.</b> Same shift as offset Horner, but the degree is an argument, so the number of powered shares is chosen when the keys are built.</div>
<divclass="eli5"><b>ELI5.</b> A carry across a shift is a short list of comparisons and bit corrections, not a generic circuit. The request names the source width, the shift, and the width of what comes out. Truncate, arithmetic shift, and sign-extend are the same plan with different output widths.</div>
<divclass="eli5"><b>ELI5.</b> One walk of an existing key stops at the public endpoints and folds XOR or addition along that prefix. The fold is O(number of endpoints), not a new key per prefix.</div>
<divclass="eli5"><b>ELI5.</b> Stack the breakpoints of every table into one sorted list. One comparison and one prefix walk label the pieces of that list. Each table then sums the labels that fall inside its own intervals.</div>
\endhtmlonly
`make_lut_union_plan(luts, eta)` shifts every piecewise LUT by the public
`eta`, inserts the same domain-minimum and carry cuts as
[offset polynomial](@ref offset_poly), and sorts the union. A piece of one
LUT is a span of those union knots: `[begin, end)`, or
`[begin, end-of-union) ∪ [0, end)` when the piece wraps. The span stores
that piece's binomial shift by its public `kappa`. Spans of one LUT
partition the union.
The interactive plan is one comparison, whatever the number of LUTs and
whatever the number of union knots:
-`plan.comparisons` and `plan.prefix_walks` are 1.
-`plan.depth` and `plan.geneval_rounds()` are the bitlength of the input.
-`plan.degree` is the widest polynomial. The payload is
`1, center, …, center^degree`.
-`schedule_lut_union` records one `fss_cmp` of that depth. The slot is
`lut_union_slot_bytes`: one AES block, or `lanes * 8` when the power
vector is wider. Prefix parity of the union is local after that
comparison.
`lut_union_eval<Party>` reads one `make_offset_poly_keys` key of degree
at least `plan.degree` and returns a share per LUT. `geneval_lut_union`
opens that comparison from XOR shares of the center, as in
`geneval_offset_horner`. Pass `dpf::arith_input` when the shares add to
the center in the input group.
`piecewise_from_easy` and `piecewise_from_constant` adapt the cleartext
tables. An `easy_lut` denominator other than 1 is a rounding division, so
`piecewise_from_easy` rejects it. Powers that are zero on every piece are
dropped, and the shared payload stays only as wide as the widest remaining
<divclass="eli5"><b>ELI5.</b> The product lives in a ring wide enough for both fixed-point operands. One Beaver triple in that ring is the product; the binary point is placed by a public shift afterward.</div>
<divclass="eli5"><b>ELI5.</b> A cleartext approximation is replaced by a table addressed with the secret. Constant, easy, range, window, and principal tables differ in how many bits of the input they consume and how the correction is added.</div>
auto sign = grotto::make_exact_constant_lut<std::int32_t>(
grotto::exact_constant::signum, 0);
auto s = sign(std::int32_t{-3});
auto relu = grotto::make_relu_lut<std::int32_t>(8);
auto ln = grotto::eval_reduced(grotto::reduced::ln, 16, raw);
auto gelu = grotto::eval_window(grotto::window::gelu, 16, raw);
auto sine = grotto::eval_principal(grotto::principal::sin, 16, raw);
\endcode
The degree-0 exact tables follow Storrier, Vadapalli, Lyons, and Henry,
[ePrint 2023/108](@ref bib_grotto), Appendix D.
Sign predicates are a constant number of cuts, `Θ(1)`. `clz` and
`ilogb` cut once per bit of the raw width, `Θ(w)`. `ilog10` binary-searches
the raw domain once per decimal exponent, `Θ(w²)` probes. `make_msb_lut`
emits `Θ(2^index)` cuts and rejects `index` at or above 8.
`make_clipped_quotient_lut` is linear in `(high-low)/modulus`, capped at
`2^16` pieces. Calling a constant or easy table binary-searches its `P`
pieces, `O(log P)`. `eval_principal` and `eval_window` binary-search the
static knots and then run one cubic. `eval_reduced` adds a short series
on the small interval (`expm1` 24 terms, `log1p` 80). No DPF and no
communication.
## Appendix D maps that were still cleartext {#appendix_d_gaps}
Appendix D of [ePrint 2023/108](@ref bib_grotto) lists HardELiSH, LeCun tanh, and leaky
ReLU with slope `1/100`. Those three now have fixed-point evaluators.
`one_minus_sigmoid` is the sigmoid table complemented, which rounds out
the logistic pair.
The cubics were built on mocha2. Sollya chose the longest pieces whose
absolute error stays within half an ulp. Mathematica (Remez), Maple
(`numapprox[minimax]`), and MATLAB/Chebfun (`minimax`) fitted each
piece, and the shipped polynomial is the one with the lowest error
after the coefficients are rounded to `k+16` fraction bits. Of the 443
cubics, Sollya won 229, Maple 83, MATLAB 74, and Mathematica 57.
Half an ulp at 16 fraction bits is `2^{-17} ≈ 7.63e-6`. Appendix D's
own columns are tighter (`4.2e-8`) and therefore use more pieces
(HardELiSH 38, LeCun tanh 89). The counts below are this library's
half-ulp partitions.
| map | degree | pieces at k = 8, 12, 16, 20, 24, 28, 32 | evaluation |
| --- | --- | --- | --- |
| `window::hardelish` | 3 on `(-1, 0)`; exact quadratic on `[0, 1]` | 1, 2, 3, 6, 12, 23, 45 | `Θ(log P)` knot search and one cubic on `(-1, 0)`. Elsewhere `Θ(1)`: `0`, `round(x(x+1)/2)`, or `x` |
| `window::lecun_tanh` | 3 | 3, 6, 11, 22, 44, 88, 177 | `Θ(log P)` on the positive knots of the absolute value, then a sign. Past the last knot the value is the constant `±round(1.7159 · 2^k)` |
| `window::one_minus_sigmoid` | 3 | same as `sigmoid`: 8, 16, 64, 128, 256, 1024, 2048 | one sigmoid evaluation and one subtraction. `Θ(log P)` |
| `make_leaky_relu_hundredth_lut` | 1 | 2 | `Θ(1)`. Identity on the right, `round(x/100)` on the left. Error at most half a unit in the last place |
`make_leaky_relu_lut(shift)` is still the dyadic slope `1/2^shift`.
The hundredth factory is the Appendix D slope and does not depend on
the fractional width.
`grotto::polynomials::eval_horner` evaluates a `poly_constant`,
`poly_linear`, `poly_quadratic`, or `poly_cubic` (a `std::array` of
coefficients, constant term first). `piecewise_eval(polys, bounds, x)`
<divclass="eli5"><b>ELI5.</b> A wavelet step is a fixed linear combination. The LUT stores that combination so the signal stays in shares and never enters a floating-point routine. Haar and biorthogonal 5/3 are the two filters.</div>
\endhtmlonly
`make_haar_dwt_lut` and `make_bior53_dwt_lut` compress a real signal of
length \f$2^n\f$ and evaluate it as a fixed-point word. The construction
is the cleartext Haar and bior(5,3) lookup of Reis, Ugurbil, Wagh, Henry,
and de Vega, [ePrint 2025/013](@ref bib_wave), Equations (7) and (8).
`sample_dwt_signal(domain_bits, fractional_bits, f)` writes the grid
\f$i \cdot 2^{-f}\f$ for \f$i \in [0, 2^n)\f$.
Both builders run the depth-\f$j\f$ low-pass with the smooth edge
extension used for that paper's accuracy tables. PyWavelets calls these
filters `haar` and `bior2.2`; bior(5,3) is the same pair, named there by
vanishing moments. Building either table is \f$\Theta(N)\f$ arithmetic
and extra memory, \f$N = 2^n\f$.
Haar then multiplies the approximation coefficients by \f$2^{-j/2}\f$
and rounds down to \f$f\f$ fraction bits. On this grid that coefficient
is the mean of each block of \f$2^j\f$ samples. Evaluation reads
`coeff[raw >> j]`, one indexing step, \f$\Theta(1)\f$.
bior(5,3) multiplies by \f$2^{j/2}\f$ and rounds down the same way.
Smooth extension prepends two coefficients, so the bin `msb = raw >> j`
lives at index `msb + 2`, and the next tap at `msb + 3`, wrapping in the
stored vector. With `lsb = raw mod 2^j`,
\f[
y = \bigl\lfloor\bigl(c_{\mathrm{msb}+2}\,(2^j - \mathrm{lsb})
<divclass="tldr"><b>TL;DR.</b> Open eta = x − r once. Jets and Horner turn that public distance into polynomial powers. Ring switch and carry move the integer. Prefix parity folds a key you already hold. Lookup tables, including the wavelet tables, replace cleartext math. A fixed-point product is one Beaver triple.</div>