Publish a scannable manual and a bibliography that links each paper back to the pages that use it.

Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
Ryan Henry 2026-09-26 23:51:06 -06:00
parent cf8054a0b3
commit 695f8e84f7
45 changed files with 4848 additions and 338 deletions

425
doc/pages/jet_and_ring.md Normal file
View file

@ -0,0 +1,425 @@
# Jet and exact ring switch {#jet_and_ring}
One opened offset `eta = x - r` drives two cheap corrections. The binomial
jet returns shares of \f$\binom{x}{0},\ldots,\binom{x}{d}\f$ after a public
Chu–Vandermonde shift. The ring switch returns shares of `x` in any residue
group whose comparison payload is the destination modulus.
The same offset also drives [offset Horner](@ref offset_horner),
[offset polynomials](@ref offset_poly), and [carry](@ref carry).
[Prefix parity](@ref prefix_parity) reads a key's path.
[Cleartext maps](@ref grotto_luts) evaluate fixed-point functions with no tree.
## Binomial jet {#offset_jet}
`make_offset_jet_keys(center, degree)` keys one incremental `gt` whose
payload is the vector of \f$\binom{\mathrm{center}}{k}\f$ in
\f$\mathbb{Z}/2^{64}\f$. After `eta` opens,
the same knot shift and carry cut as offset poly refine the pieces. On the
piece with carry `kappa`,
\f[
\binom{c+\kappa}{k}
=\sum_j\binom{c}{j}\binom{\kappa}{k-j}.
\f]
`make_offset_jet_keys` writes one incremental comparison for degree `d`
(`d ≤ 16`). The seed spine is `Θ(n λ)` bits, with `n` the center's bit
length and `λ` the seed width. Value words grow with the `d+1` binomial
lanes. After `η` is public, `offset_jet_shares` evaluates that one key
on the `K` knots, the same order as one `eval_sequence` on the knots.
The Chu–Vandermonde
update after those walks is `Θ(P · d²)` arithmetic, where `P` is the
number of refined pieces (the knots, plus the domain minimum, plus the
carry cut when the input width is at most 62). `offset_jet_dot` is
`Θ(d)`. No further round when the coefficients are public.
`offset_jet_shares` returns that shifted jet. Public dots are free:
- value of \f$\sum a_k\binom{x}{k}\f$ via `offset_jet_dot`;
- forward difference via `offset_jet_difference_coeff` (Pascal);
- hockey-stick prefix via `offset_jet_prefix_coeff`.
The prefix needs \f$\binom{x}{k+1}\f$, so the key degree must be one larger
than the polynomial degree. Degree 16 therefore prefix-sums polynomials
through degree 15.
Binomials modulo \f$2^{64}\f$ use a falling factorial modulo
\f$2^{64+v_2(k!)}\f$ (\f$v_2(16!)=15\f$), then multiply by the inverse of the
odd part of \f$k!\f$. Dividing by \f$k!\f$ inside \f$\mathbb{Z}/2^{64}\f$ alone
is not exact.
A Padé pair or one Newton correction is two dots against the same jet and
one reciprocal after the shares are opened. Those are not separate APIs.
**Code samples**\n
<div class="tabbed">
- <b class="tab-title">jet_and_ring.cpp</b> \include{cpp} grotto/jet_and_ring.cpp
</div>
## Exact ring switch {#ring_switch}
For an unsigned \f$n\f$-bit limb (\f$n\le 64\f$) with representatives in
\f$[0,2^n)\f$,
\f[
\eta + r = x + w\cdot 2^n,\qquad
w=\mathbf{1}[r+\eta\ge 2^n].
\f]
In any modulus \f$M\f$,
\f[
x \equiv \eta + (r\bmod M) - w\cdot(2^n\bmod M)\pmod M.
\f]
A `uint64` comparison share is not a share mod \f$M\f$. The payload of the
wrap comparison is the destination element \f$2^n\bmod M\f$. The dealer keys
`lt(2^n \bmod M)` at the secret `r` and stores an additive split of `r` in
the residue group. After `eta` opens, each party evaluates at the public
query \f$2^n-1-\eta\f$. That indicator is hot exactly on wrap, including the
`eta = 0` case. Party 0 adds public `eta`.
Destination groups:
- `grotto::zn64<Mod>` and `grotto::zn128<Lo,Hi>` ([residue.hpp](@ref grotto/residue.hpp));
- `dpf::field128`;
- `dpf::p256_scalar` (NIST P-256 order, not the point group).
`ring_switch_factor<Factor>` reduces a share when `Factor` divides the
modulus. One switch into an lcm yields every factor by local reduction.
The dealer material is one `lt` key on that limb, `Θ(n λ)` bits for
limb width `n ≤ 64` and seed width `λ`, plus two residue shares of `r`. After `η` is
public, each party does one point evaluation (`Θ(n)` expands) and a
constant amount of arithmetic in the destination group.
`ring_switch_factor` is local.
This is the exact neighbour of truncated Barrett `nmod`.
`grotto::nmod(x_raw, x_bits, recip_raw, recip_bits, residue_bits)` splits
`x / M` when `recip_raw / 2^recip_bits` is a positive approximation of `1/M`.
The result is an `nmod_result`: `quotient` is `floor(x/M)`, and `residue`
is the fractional part truncated onto `residue_bits`.
`nmod_pow2(x_raw, x_bits, exp, residue_bits)` is the same split when the
modulus is a power of two. Both are one product by a reciprocal of at
most 128 bits, so time and extra memory are constant in the word size.
\code{cpp}
auto split = grotto::nmod_pow2(raw, 16, 0, 16);
\endcode
**Defined in**\n
@ref grotto/nmod.hpp
See also [representation shift and twisted jets](@ref repr_and_twist).
## Offset Horner {#offset_horner}
`make_offset_horner_keys<Input, Degree>(center)` keys one `gt` whose
payload is `center^m` for `m = 0 .. Degree`. `Degree` is at most 3
(`offset_horner_max_degree`). Pass `dpf::verifiable{}` for proof tokens.
After `eta` opens, `offset_horner_eval<Party, Degree>` returns that party's
share of the cubic at the wrapped point. Coefficients are one
`std::array<uint64_t, Degree + 1>` per knot, low degree first.
\code{cpp}
const std::uint8_t center = 12;
auto mat = grotto::make_offset_horner_keys<std::uint8_t, 2>(center);
std::vector<std::uint8_t> knots{0};
std::vector<std::array<std::uint64_t, 3>> coeff{{4, 2, 1}};
auto s0 = grotto::offset_horner_eval<0, 2>(mat, knots, coeff, eta);
\endcode
`geneval_offset_horner` runs the same cubic from Jack Doerner and abhi shelat shares of
`x` and of the center, on a `dpf::ds_randomness` tape.
Degree is at most 3. The seed spine is one comparison, `Θ(n λ)` bits,
and the value words hold the four powers. Evaluation after `η` opens is
one sequence-shaped walk on the knots plus `O(1)` arithmetic. The
geneval form generates that same comparison once, opens one correction
word per level, as in [geneval](@ref tour_ds), and does not store a
reusable key.
**Defined in**\n
@ref grotto/offset_horner.hpp
## Offset polynomial {#offset_poly}
`make_offset_poly_keys(center, degree)` is offset Horner at a runtime
degree, at most 16 (`offset_poly_max_degree`). One incremental `gt`
whose payload is the vector of powers. `offset_poly_eval<Party>` dots the shifted powers.
`offset_poly_clear` is the same polynomial in the clear.
`offset_poly_kappas` is the public carry of each piece.
Shared coefficients use `offset_poly_shift_share` (the binomial map is
linear) and `offset_poly_beaver_share` for the dot.
Degree `d` is at most 16: one key, `Θ(n λ)` bits of seed spine plus
value words that grow with `d`. The clear and public-coefficient evals
are one sequence-shaped walk on the `K` knots, then `O(d^2)` arithmetic. A shared-coefficient dot is one Beaver
inner product: one opening round of the masked vectors, communication
linear in the flattened length (pieces times `d+1` coefficients), and
one product share per coefficient in preprocessing. The shift of each
party's coefficient share is local.
\code{cpp}
auto mat = grotto::make_offset_poly_keys(std::uint8_t{12}, 4);
std::vector<std::uint8_t> knots{0};
std::vector<std::vector<std::uint64_t>> coeff{{4, 2, 1, 0, 0}};
auto s0 = grotto::offset_poly_eval<0>(mat, knots, coeff, eta);
auto opened = s0 + grotto::offset_poly_eval<1>(mat, knots, coeff, eta);
\endcode
**Defined in**\n
@ref grotto/offset_poly.hpp
## Carry {#carry}
A `carry_request` names the source width `n`, the shift `s`, the output
width `out_n`, a `carry_mode` (`truncate_reduce`, `same_ring`, `extend`,
`window`), and a `sign_knowledge` (`unknown`, `nonnegative`, `negative`).
`plan_carry` returns a `carry_recipe` whose flags are the steps that are
still live. `plan_carry_in(n, s)`, `plan_carry_out(n, s, sign)`, and
`plan_carry_fused(n, s, out_n, sign)` fill the common requests.
`make_carry_keys(recipe)` (and `make_carry_in_keys`, `make_carry_out_keys`,
`make_carry_fused_keys`) builds the dealer keys. `finalize_carry_in_blinds`
adjusts a truncate-reduce split. Online, `eval_carry_in(keys, party, opened)`
returns a `carry_eval_share` whose `value` is that party's share.
`opened` is `(x0 + x1 + rin) mod 2^n`. The other online entry points are
`eval_carry_out_known`, `eval_carry_out_unknown`, `eval_carry_extend`,
`eval_carry_window`, and `eval_carry_fused`.
Cleartext twins, for tests and for a public limb, are `eval_carry_clear`,
`carry_in_clear`, `carry_out_clear`, `carry_asr`, and `carry_mask`.
`plan_carry` is a constant-time inspection of the request. Each live
comparison flag becomes one DPF key whose domain is the limb width `w`
of that comparison, `Θ(w λ)` bits, and the online step is one point
walk of that key. A share-MSB AND adds one Beaver bit triple in
preprocessing and one opening round of a bit.
\code{cpp}
auto keys = grotto::make_carry_in_keys(32, 8);
auto share = grotto::eval_carry_in(keys, /*party*/ 0, opened);
auto clear = grotto::eval_carry_clear(keys.recipe, x0, x1);
\endcode
**Defined in**\n
@ref grotto/carry_plan.hpp, @ref grotto/carry.hpp
## Prefix parity {#prefix_parity}
`prefix_parities(key, endpoints)` walks a key to the sorted endpoints and
returns XOR shares of the prefix parities, plus the index of the first
endpoint on the wrap. `segment_parities` turns those into one share per
segment. `all_segment_parities_from_prefix_parities` is the same conversion
when you already hold the prefix array.
`signed_prefix_parities(key, endpoints)` needs a comparison channel
(assigned, if the payload was a wildcard). It returns one additive
`uint64_t` share per endpoint: for `dpf::gt(1)` that share opens to 1 when
the secret point is below the endpoint. `signed_prefix_parities_into`
writes a runtime-length buffer.
The prefix walk follows Storrier, Vadapalli, Lyons, and Henry, ePrint
2023/108: one key's prefix parity in place of a comparison per piece.
On `m` endpoints the walk resumes one path memoizer (`Θ(n)` nodes, `n`
the key depth). Expands are the nodes on those paths, `O(m n)` in the
worst case, and less when endpoints share a prefix or the zero-suffix
stop hits. `signed_prefix_parities` adds an `O(n)` sum of
value-correction words on each endpoint. Both calls are local.
`segment_parities` is the prefix walk plus an `O(m)` XOR of those bits.
\code{cpp}
std::array<std::uint8_t, 2> ends{10, 40};
auto [bits, first] = grotto::prefix_parities(k0, ends);
auto segs = grotto::segment_parities(k0, ends);
auto signs = grotto::signed_prefix_parities(cmp0, ends);
\endcode
**Defined in**\n
@ref grotto/prefix_parity.hpp
## Cleartext maps {#grotto_luts}
These functions take a raw fixed-point word (`n << fractional_bits`) and
return a raw word. They do not build a DPF. The type
`grotto::fixedpoint` itself is a domain and an output; see
[Input types](@ref input_types) and [Output types](@ref output_types).
## Fixed-point product {#fixedpoint_mul}
`fixed_mul<IntegerBits, FractionalBits>(lhs, rhs)` multiplies two
`fixedpoint` values and keeps that many integer bits (including the sign)
and fraction bits. Bits below the fraction are floored. The product type
is the `result_type` of `fixed_mul_plan`. The plan uses at most 8
limbs and refuses a wider window, so the product is a constant amount
of 64-bit arithmetic and `O(1)` extra memory.
\code{cpp}
using q16 = grotto::fixedpoint<16, std::int32_t>;
auto prod = grotto::fixed_mul<16, 16>(q16{1.5}, q16{2.0});
\endcode
**Defined in**\n
@ref grotto/fixedpoint_mul.hpp
## Lookup tables {#lookup_tables}
Constant, easy, principal, range, and window tables are included from
`grotto.hpp`. The dyadic table comes in through `exact_steps.hpp`, which
`grotto.hpp` also includes.
- **Constant.** `make_exact_constant_lut<Raw>(exact_constant::signum, fractional_bits)`
and the other `exact_constant` names (`positive`, `negative`, `nonneg`,
`nonpos`, `zero`, `nonzero`, `ilogb`, `ceil_ilogb`, `ilog10`, `clz`,
`clrsb`). `make_threshold_lut`, `make_interval_lut`, and
`make_clipped_quotient_lut` build a `constant_lut<Raw>` you call as
`table(raw)`.
- **Easy.** Few-piece polynomials with integer knots:
`make_abs_lut`, `make_relu_lut`, `make_clip_lut`, `make_hardsigmoid_lut`,
`make_hardswish_lut`, `make_leaky_relu_hundredth_lut`, and the other
`make_*_lut` factories in [easy_lut.hpp](@ref grotto/easy_lut.hpp).
The result is an `easy_lut<Raw>`. `make_leaky_relu_lut(shift)` is the
dyadic slope `1/2^shift`. The Appendix D leaky ReLU is slope `1/100`.
- **Dyadic.** Exact steps on powers of two: `make_signum_lut`,
`make_msb_lut(index)`, `make_ilogb_lut`, `make_ilog10_lut`, `make_clz_lut`,
`make_clrsb_lut`, and the sign predicates `make_positive_lut` through
`make_nonzero_lut`. `ilog_of_zero` is the sentinel for a zero argument.
`msb_bit_limit` is 8.
- **Range.** `eval_reduced(reduced::ln, fractional_bits, raw)` and the
other `reduced` names (`lg`, `log10`, `exp`, `exp2`, `exp10`, `sin`,
`cos`, `tan`, `cot`, `sec`, `csc`, the hyperbolics, `sqrt`, `inv`,
`rsqrt`, `invsq`, `expm1`, `log1p`). `split_positive` is the dyadic
mantissa split those reductions use.
- **Window.** `eval_window(window::gelu, fractional_bits, raw)`. The
`window` names cover `smoothstep`, `sigmoid`, `tanh`, `erf`, `erfc`,
`softplus`, `gelu`, `silu`, `asin`, `acos`, `probit`, `hardelish`,
`lecun_tanh`, `one_minus_sigmoid`, and the rest of the enum in
[window_lut.hpp](@ref grotto/window_lut.hpp).
- **Principal.** `eval_principal(principal::sin, fractional_bits, raw)` on
the closed principal interval. Precisions are 8, 12, …, 32
(`principal_precision`). Names: `ln`, `exp`, `sin`, `tanf`, `tang`,
`sinh`, `cosh`, `sqrt`, `coth`, `sec`, `gsec`, `csch`, `inv`, `rsqrt`,
`invsq`.
\code{cpp}
auto sign = grotto::make_exact_constant_lut<std::int32_t>(
grotto::exact_constant::signum, 0);
auto s = sign(std::int32_t{-3});
auto relu = grotto::make_relu_lut<std::int32_t>(8);
auto ln = grotto::eval_reduced(grotto::reduced::ln, 16, raw);
auto gelu = grotto::eval_window(grotto::window::gelu, 16, raw);
auto sine = grotto::eval_principal(grotto::principal::sin, 16, raw);
\endcode
The degree-0 exact tables follow Storrier, Vadapalli, Lyons, and Henry,
[ePrint 2023/108](@ref bib_grotto), Appendix D.
Sign predicates are a constant number of cuts, `Θ(1)`. `clz` and
`ilogb` cut once per bit of the raw width, `Θ(w)`. `ilog10` binary-searches
the raw domain once per decimal exponent, `Θ(w²)` probes. `make_msb_lut`
emits `Θ(2^index)` cuts and rejects `index` at or above 8.
`make_clipped_quotient_lut` is linear in `(high-low)/modulus`, capped at
`2^16` pieces. Calling a constant or easy table binary-searches its `P`
pieces, `O(log P)`. `eval_principal` and `eval_window` binary-search the
static knots and then run one cubic. `eval_reduced` adds a short series
on the small interval (`expm1` 24 terms, `log1p` 80). No DPF and no
communication.
## Appendix D maps that were still cleartext {#appendix_d_gaps}
Appendix D of [ePrint 2023/108](@ref bib_grotto) lists HardELiSH, LeCun tanh, and leaky
ReLU with slope `1/100`. Those three now have fixed-point evaluators.
`one_minus_sigmoid` is the sigmoid table complemented, which rounds out
the logistic pair.
The cubics were built on mocha2. Sollya chose the longest pieces whose
absolute error stays within half an ulp. Mathematica (Remez), Maple
(`numapprox[minimax]`), and MATLAB/Chebfun (`minimax`) fitted each
piece, and the shipped polynomial is the one with the lowest error
after the coefficients are rounded to `k+16` fraction bits. Of the 443
cubics, Sollya won 229, Maple 83, MATLAB 74, and Mathematica 57.
Half an ulp at 16 fraction bits is `2^{-17} ≈ 7.63e-6`. Appendix D's
own columns are tighter (`4.2e-8`) and therefore use more pieces
(HardELiSH 38, LeCun tanh 89). The counts below are this library's
half-ulp partitions.
| map | degree | pieces at k = 8, 12, 16, 20, 24, 28, 32 | evaluation |
| --- | --- | --- | --- |
| `window::hardelish` | 3 on `(-1, 0)`; exact quadratic on `[0, 1]` | 1, 2, 3, 6, 12, 23, 45 | `Θ(log P)` knot search and one cubic on `(-1, 0)`. Elsewhere `Θ(1)`: `0`, `round(x(x+1)/2)`, or `x` |
| `window::lecun_tanh` | 3 | 3, 6, 11, 22, 44, 88, 177 | `Θ(log P)` on the positive knots of the absolute value, then a sign. Past the last knot the value is the constant `±round(1.7159 · 2^k)` |
| `window::one_minus_sigmoid` | 3 | same as `sigmoid`: 8, 16, 64, 128, 256, 1024, 2048 | one sigmoid evaluation and one subtraction. `Θ(log P)` |
| `make_leaky_relu_hundredth_lut` | 1 | 2 | `Θ(1)`. Identity on the right, `round(x/100)` on the left. Error at most half a unit in the last place |
`make_leaky_relu_lut(shift)` is still the dyadic slope `1/2^shift`.
The hundredth factory is the Appendix D slope and does not depend on
the fractional width.
`grotto::polynomials::eval_horner` evaluates a `poly_constant`,
`poly_linear`, `poly_quadratic`, or `poly_cubic` (a `std::array` of
coefficients, constant term first). `piecewise_eval(polys, bounds, x)`
picks the piece and calls that Horner step.
**Defined in**\n
@ref grotto/constant_lut.hpp, @ref grotto/easy_lut.hpp,
@ref grotto/dyadic_lut.hpp, @ref grotto/range_lut.hpp,
@ref grotto/window_lut.hpp, @ref grotto/principal_lut.hpp,
@ref grotto/piecewise.hpp
## Closed form {#closed_form}
`eval_closed(closed::atanh, fractional_bits, raw)` composes
`eval_reduced` and `eval_window`. The `closed` names are the inverse
hyperbolics and inverse trig functions, `selu`, `elu`, `celu`,
`softsign`, `tanhshrink`, the `logistic` / `exponential` / `laplace` /
`cauchy` quantiles, `sinc`, and the extra powers `cbrt`, `qtrt`,
`icbrt`, `iqtrt`, `pow_m01`, `pow_p15`, `pow_m3`. Precision is one of
8, 12, …, 32. Each call runs a constant number of `eval_reduced` or
`eval_window` evaluations.
\code{cpp}
auto y = grotto::eval_closed(grotto::closed::atan, 16, raw);
\endcode
**Defined in**\n
@ref grotto/closed_form.hpp
## Exact steps {#exact_steps}
Counts and booleans come back as fixed-point integers,
`n << fractional_bits`.
\code{cpp}
auto floor_x = grotto::eval_dec_floor(raw, 16);
auto digits = grotto::eval_dec_width(raw, 16);
auto bits = grotto::eval_bit_width<std::int32_t>(raw, 16);
\endcode
Decimal digit counts divide in a loop, so the time follows the number
of digits. `eval_bit_width`, `eval_bit_floor`, `eval_bit_ceil`,
`eval_countl_one`, and `eval_has_single_bit` scan the raw width,
`Θ(width)` bit operations and at most 64 shifts. `eval_logstar` is five
magnitude comparisons, `Θ(1)`. `eval_deg2rad` and `eval_rad2deg` are one
scale each.
Also `eval_dec_ceil`, `eval_oct_width`, `eval_b64_width`,
`eval_value_length` (base 8, 10, or 64), `eval_has_single_digit`,
`eval_bit_floor`, `eval_bit_ceil`, `eval_countl_one`, `eval_has_single_bit`,
`eval_deg2rad`, and `eval_rad2deg`. `make_exact_step_lut` builds the
matching `easy_lut`.
**Defined in**\n
@ref grotto/exact_steps.hpp
## Gadget functors {#grotto_gadgets}
`grotto/gadgets.hpp` still includes the decimal and exponential reference
headers. The functors that those headers used to provide are deprecated:
call `eval_reduced`, `eval_window`, `make_*_lut`, or `exact_constant`
instead. `gadget_hints<T>` holds the old domain, degree, and pole notes
for a functor type.
**Defined in**\n
@ref grotto/gadgets.hpp, @ref grotto/gadget_hints.hpp