libdpf/doc/pages/jet_and_ring.md
2026-09-26 23:51:06 -06:00

19 KiB
Raw Blame History

Jet and exact ring switch

One opened offset eta = x - r drives two cheap corrections. The binomial jet returns shares of \f$\binom{x}{0},\ldots,\binom{x}{d}\f$ after a public Chu–Vandermonde shift. The ring switch returns shares of x in any residue group whose comparison payload is the destination modulus.

The same offset also drives [offset Horner](@ref offset_horner), [offset polynomials](@ref offset_poly), and [carry](@ref carry). [Prefix parity](@ref prefix_parity) reads a key's path. [Cleartext maps](@ref grotto_luts) evaluate fixed-point functions with no tree.

Binomial jet

make_offset_jet_keys(center, degree) keys one incremental gt whose payload is the vector of \f$\binom{\mathrm{center}}{k}\f$ in \f$\mathbb{Z}/2^{64}\f$. After eta opens, the same knot shift and carry cut as offset poly refine the pieces. On the piece with carry kappa,

\f[ \binom{c+\kappa}{k} =\sum_j\binom{c}{j}\binom{\kappa}{k-j}. \f]

make_offset_jet_keys writes one incremental comparison for degree d (d ≤ 16). The seed spine is Θ(n λ) bits, with n the center's bit length and λ the seed width. Value words grow with the d+1 binomial lanes. After η is public, offset_jet_shares evaluates that one key on the K knots, the same order as one eval_sequence on the knots. The Chu–Vandermonde update after those walks is Θ(P · d²) arithmetic, where P is the number of refined pieces (the knots, plus the domain minimum, plus the carry cut when the input width is at most 62). offset_jet_dot is Θ(d). No further round when the coefficients are public.

offset_jet_shares returns that shifted jet. Public dots are free:

  • value of \f$\sum a_k\binom{x}{k}\f$ via offset_jet_dot;
  • forward difference via offset_jet_difference_coeff (Pascal);
  • hockey-stick prefix via offset_jet_prefix_coeff.

The prefix needs \f$\binom{x}{k+1}\f$, so the key degree must be one larger than the polynomial degree. Degree 16 therefore prefix-sums polynomials through degree 15.

Binomials modulo \f$2^{64}\f$ use a falling factorial modulo \f$2^{64+v_2(k!)}\f$ (\f$v_2(16!)=15\f$), then multiply by the inverse of the odd part of \f$k!\f$. Dividing by \f$k!\f$ inside \f$\mathbb{Z}/2^{64}\f$ alone is not exact.

A Padé pair or one Newton correction is two dots against the same jet and one reciprocal after the shares are opened. Those are not separate APIs.

Code samples\n

  • jet_and_ring.cpp \include{cpp} grotto/jet_and_ring.cpp

Exact ring switch

For an unsigned \f$n\f$-bit limb (\f$n\le 64\f$) with representatives in \f$[0,2^n)\f$,

\f[ \eta + r = x + w\cdot 2^n,\qquad w=\mathbf{1}[r+\eta\ge 2^n]. \f]

In any modulus \f$M\f$,

\f[ x \equiv \eta + (r\bmod M) - w\cdot(2^n\bmod M)\pmod M. \f]

A uint64 comparison share is not a share mod \f$M\f$. The payload of the wrap comparison is the destination element \f$2^n\bmod M\f$. The dealer keys lt(2^n \bmod M) at the secret r and stores an additive split of r in the residue group. After eta opens, each party evaluates at the public query \f$2^n-1-\eta\f$. That indicator is hot exactly on wrap, including the eta = 0 case. Party 0 adds public eta.

Destination groups:

  • grotto::zn64<Mod> and grotto::zn128<Lo,Hi> ([residue.hpp](@ref grotto/residue.hpp));
  • dpf::field128;
  • dpf::p256_scalar (NIST P-256 order, not the point group).

ring_switch_factor<Factor> reduces a share when Factor divides the modulus. One switch into an lcm yields every factor by local reduction.

The dealer material is one lt key on that limb, Θ(n λ) bits for limb width n ≤ 64 and seed width λ, plus two residue shares of r. After η is public, each party does one point evaluation (Θ(n) expands) and a constant amount of arithmetic in the destination group. ring_switch_factor is local.

This is the exact neighbour of truncated Barrett nmod. grotto::nmod(x_raw, x_bits, recip_raw, recip_bits, residue_bits) splits x / M when recip_raw / 2^recip_bits is a positive approximation of 1/M. The result is an nmod_result: quotient is floor(x/M), and residue is the fractional part truncated onto residue_bits. nmod_pow2(x_raw, x_bits, exp, residue_bits) is the same split when the modulus is a power of two. Both are one product by a reciprocal of at most 128 bits, so time and extra memory are constant in the word size.

\code{cpp} auto split = grotto::nmod_pow2(raw, 16, 0, 16); \endcode

Defined in\n @ref grotto/nmod.hpp

See also [representation shift and twisted jets](@ref repr_and_twist).

Offset Horner

make_offset_horner_keys<Input, Degree>(center) keys one gt whose payload is center^m for m = 0 .. Degree. Degree is at most 3 (offset_horner_max_degree). Pass dpf::verifiable{} for proof tokens. After eta opens, offset_horner_eval<Party, Degree> returns that party's share of the cubic at the wrapped point. Coefficients are one std::array<uint64_t, Degree + 1> per knot, low degree first.

\code{cpp} const std::uint8_t center = 12; auto mat = grotto::make_offset_horner_keys<std::uint8_t, 2>(center); std::vectorstd::uint8_t knots{0}; std::vector<std::array<std::uint64_t, 3>> coeff{{4, 2, 1}}; auto s0 = grotto::offset_horner_eval<0, 2>(mat, knots, coeff, eta); \endcode

geneval_offset_horner runs the same cubic from Jack Doerner and abhi shelat shares of x and of the center, on a dpf::ds_randomness tape.

Degree is at most 3. The seed spine is one comparison, Θ(n λ) bits, and the value words hold the four powers. Evaluation after η opens is one sequence-shaped walk on the knots plus O(1) arithmetic. The geneval form generates that same comparison once, opens one correction word per level, as in [geneval](@ref tour_ds), and does not store a reusable key.

Defined in\n @ref grotto/offset_horner.hpp

Offset polynomial

make_offset_poly_keys(center, degree) is offset Horner at a runtime degree, at most 16 (offset_poly_max_degree). One incremental gt whose payload is the vector of powers. offset_poly_eval<Party> dots the shifted powers. offset_poly_clear is the same polynomial in the clear. offset_poly_kappas is the public carry of each piece. Shared coefficients use offset_poly_shift_share (the binomial map is linear) and offset_poly_beaver_share for the dot.

Degree d is at most 16: one key, Θ(n λ) bits of seed spine plus value words that grow with d. The clear and public-coefficient evals are one sequence-shaped walk on the K knots, then O(d^2) arithmetic. A shared-coefficient dot is one Beaver inner product: one opening round of the masked vectors, communication linear in the flattened length (pieces times d+1 coefficients), and one product share per coefficient in preprocessing. The shift of each party's coefficient share is local.

\code{cpp} auto mat = grotto::make_offset_poly_keys(std::uint8_t{12}, 4); std::vectorstd::uint8_t knots{0}; std::vector<std::vectorstd::uint64_t> coeff{{4, 2, 1, 0, 0}}; auto s0 = grotto::offset_poly_eval<0>(mat, knots, coeff, eta); auto opened = s0 + grotto::offset_poly_eval<1>(mat, knots, coeff, eta); \endcode

Defined in\n @ref grotto/offset_poly.hpp

Carry

A carry_request names the source width n, the shift s, the output width out_n, a carry_mode (truncate_reduce, same_ring, extend, window), and a sign_knowledge (unknown, nonnegative, negative). plan_carry returns a carry_recipe whose flags are the steps that are still live. plan_carry_in(n, s), plan_carry_out(n, s, sign), and plan_carry_fused(n, s, out_n, sign) fill the common requests.

make_carry_keys(recipe) (and make_carry_in_keys, make_carry_out_keys, make_carry_fused_keys) builds the dealer keys. finalize_carry_in_blinds adjusts a truncate-reduce split. Online, eval_carry_in(keys, party, opened) returns a carry_eval_share whose value is that party's share. opened is (x0 + x1 + rin) mod 2^n. The other online entry points are eval_carry_out_known, eval_carry_out_unknown, eval_carry_extend, eval_carry_window, and eval_carry_fused.

Cleartext twins, for tests and for a public limb, are eval_carry_clear, carry_in_clear, carry_out_clear, carry_asr, and carry_mask. plan_carry is a constant-time inspection of the request. Each live comparison flag becomes one DPF key whose domain is the limb width w of that comparison, Θ(w λ) bits, and the online step is one point walk of that key. A share-MSB AND adds one Beaver bit triple in preprocessing and one opening round of a bit.

\code{cpp} auto keys = grotto::make_carry_in_keys(32, 8); auto share = grotto::eval_carry_in(keys, /party/ 0, opened); auto clear = grotto::eval_carry_clear(keys.recipe, x0, x1); \endcode

Defined in\n @ref grotto/carry_plan.hpp, @ref grotto/carry.hpp

Prefix parity

prefix_parities(key, endpoints) walks a key to the sorted endpoints and returns XOR shares of the prefix parities, plus the index of the first endpoint on the wrap. segment_parities turns those into one share per segment. all_segment_parities_from_prefix_parities is the same conversion when you already hold the prefix array.

signed_prefix_parities(key, endpoints) needs a comparison channel (assigned, if the payload was a wildcard). It returns one additive uint64_t share per endpoint: for dpf::gt(1) that share opens to 1 when the secret point is below the endpoint. signed_prefix_parities_into writes a runtime-length buffer.

The prefix walk follows Storrier, Vadapalli, Lyons, and Henry, ePrint 2023/108: one key's prefix parity in place of a comparison per piece. On m endpoints the walk resumes one path memoizer (Θ(n) nodes, n the key depth). Expands are the nodes on those paths, O(m n) in the worst case, and less when endpoints share a prefix or the zero-suffix stop hits. signed_prefix_parities adds an O(n) sum of value-correction words on each endpoint. Both calls are local. segment_parities is the prefix walk plus an O(m) XOR of those bits.

\code{cpp} std::array<std::uint8_t, 2> ends{10, 40}; auto [bits, first] = grotto::prefix_parities(k0, ends); auto segs = grotto::segment_parities(k0, ends); auto signs = grotto::signed_prefix_parities(cmp0, ends); \endcode

Defined in\n @ref grotto/prefix_parity.hpp

Cleartext maps

These functions take a raw fixed-point word (n << fractional_bits) and return a raw word. They do not build a DPF. The type grotto::fixedpoint itself is a domain and an output; see [Input types](@ref input_types) and [Output types](@ref output_types).

Fixed-point product

fixed_mul<IntegerBits, FractionalBits>(lhs, rhs) multiplies two fixedpoint values and keeps that many integer bits (including the sign) and fraction bits. Bits below the fraction are floored. The product type is the result_type of fixed_mul_plan. The plan uses at most 8 limbs and refuses a wider window, so the product is a constant amount of 64-bit arithmetic and O(1) extra memory.

\code{cpp} using q16 = grotto::fixedpoint<16, std::int32_t>; auto prod = grotto::fixed_mul<16, 16>(q16{1.5}, q16{2.0}); \endcode

Defined in\n @ref grotto/fixedpoint_mul.hpp

Lookup tables

Constant, easy, principal, range, and window tables are included from grotto.hpp. The dyadic table comes in through exact_steps.hpp, which grotto.hpp also includes.

  • Constant. make_exact_constant_lut<Raw>(exact_constant::signum, fractional_bits) and the other exact_constant names (positive, negative, nonneg, nonpos, zero, nonzero, ilogb, ceil_ilogb, ilog10, clz, clrsb). make_threshold_lut, make_interval_lut, and make_clipped_quotient_lut build a constant_lut<Raw> you call as table(raw).
  • Easy. Few-piece polynomials with integer knots: make_abs_lut, make_relu_lut, make_clip_lut, make_hardsigmoid_lut, make_hardswish_lut, make_leaky_relu_hundredth_lut, and the other make_*_lut factories in [easy_lut.hpp](@ref grotto/easy_lut.hpp). The result is an easy_lut<Raw>. make_leaky_relu_lut(shift) is the dyadic slope 1/2^shift. The Appendix D leaky ReLU is slope 1/100.
  • Dyadic. Exact steps on powers of two: make_signum_lut, make_msb_lut(index), make_ilogb_lut, make_ilog10_lut, make_clz_lut, make_clrsb_lut, and the sign predicates make_positive_lut through make_nonzero_lut. ilog_of_zero is the sentinel for a zero argument. msb_bit_limit is 8.
  • Range. eval_reduced(reduced::ln, fractional_bits, raw) and the other reduced names (lg, log10, exp, exp2, exp10, sin, cos, tan, cot, sec, csc, the hyperbolics, sqrt, inv, rsqrt, invsq, expm1, log1p). split_positive is the dyadic mantissa split those reductions use.
  • Window. eval_window(window::gelu, fractional_bits, raw). The window names cover smoothstep, sigmoid, tanh, erf, erfc, softplus, gelu, silu, asin, acos, probit, hardelish, lecun_tanh, one_minus_sigmoid, and the rest of the enum in [window_lut.hpp](@ref grotto/window_lut.hpp).
  • Principal. eval_principal(principal::sin, fractional_bits, raw) on the closed principal interval. Precisions are 8, 12, …, 32 (principal_precision). Names: ln, exp, sin, tanf, tang, sinh, cosh, sqrt, coth, sec, gsec, csch, inv, rsqrt, invsq.

\code{cpp} auto sign = grotto::make_exact_constant_lutstd::int32_t( grotto::exact_constant::signum, 0); auto s = sign(std::int32_t{-3}); auto relu = grotto::make_relu_lutstd::int32_t(8); auto ln = grotto::eval_reduced(grotto::reduced::ln, 16, raw); auto gelu = grotto::eval_window(grotto:🪟:gelu, 16, raw); auto sine = grotto::eval_principal(grotto::principal::sin, 16, raw); \endcode

The degree-0 exact tables follow Storrier, Vadapalli, Lyons, and Henry, [ePrint 2023/108](@ref bib_grotto), Appendix D. Sign predicates are a constant number of cuts, Θ(1). clz and ilogb cut once per bit of the raw width, Θ(w). ilog10 binary-searches the raw domain once per decimal exponent, Θ(w²) probes. make_msb_lut emits Θ(2^index) cuts and rejects index at or above 8. make_clipped_quotient_lut is linear in (high-low)/modulus, capped at 2^16 pieces. Calling a constant or easy table binary-searches its P pieces, O(log P). eval_principal and eval_window binary-search the static knots and then run one cubic. eval_reduced adds a short series on the small interval (expm1 24 terms, log1p 80). No DPF and no communication.

Appendix D maps that were still cleartext

Appendix D of [ePrint 2023/108](@ref bib_grotto) lists HardELiSH, LeCun tanh, and leaky ReLU with slope 1/100. Those three now have fixed-point evaluators. one_minus_sigmoid is the sigmoid table complemented, which rounds out the logistic pair.

The cubics were built on mocha2. Sollya chose the longest pieces whose absolute error stays within half an ulp. Mathematica (Remez), Maple (numapprox[minimax]), and MATLAB/Chebfun (minimax) fitted each piece, and the shipped polynomial is the one with the lowest error after the coefficients are rounded to k+16 fraction bits. Of the 443 cubics, Sollya won 229, Maple 83, MATLAB 74, and Mathematica 57.

Half an ulp at 16 fraction bits is 2^{-17} ≈ 7.63e-6. Appendix D's own columns are tighter (4.2e-8) and therefore use more pieces (HardELiSH 38, LeCun tanh 89). The counts below are this library's half-ulp partitions.

map degree pieces at k = 8, 12, 16, 20, 24, 28, 32 evaluation
window::hardelish 3 on (-1, 0); exact quadratic on [0, 1] 1, 2, 3, 6, 12, 23, 45 Θ(log P) knot search and one cubic on (-1, 0). Elsewhere Θ(1): 0, round(x(x+1)/2), or x
window::lecun_tanh 3 3, 6, 11, 22, 44, 88, 177 Θ(log P) on the positive knots of the absolute value, then a sign. Past the last knot the value is the constant ±round(1.7159 · 2^k)
window::one_minus_sigmoid 3 same as sigmoid: 8, 16, 64, 128, 256, 1024, 2048 one sigmoid evaluation and one subtraction. Θ(log P)
make_leaky_relu_hundredth_lut 1 2 Θ(1). Identity on the right, round(x/100) on the left. Error at most half a unit in the last place

make_leaky_relu_lut(shift) is still the dyadic slope 1/2^shift. The hundredth factory is the Appendix D slope and does not depend on the fractional width.

grotto::polynomials::eval_horner evaluates a poly_constant, poly_linear, poly_quadratic, or poly_cubic (a std::array of coefficients, constant term first). piecewise_eval(polys, bounds, x) picks the piece and calls that Horner step.

Defined in\n @ref grotto/constant_lut.hpp, @ref grotto/easy_lut.hpp, @ref grotto/dyadic_lut.hpp, @ref grotto/range_lut.hpp, @ref grotto/window_lut.hpp, @ref grotto/principal_lut.hpp, @ref grotto/piecewise.hpp

Closed form

eval_closed(closed::atanh, fractional_bits, raw) composes eval_reduced and eval_window. The closed names are the inverse hyperbolics and inverse trig functions, selu, elu, celu, softsign, tanhshrink, the logistic / exponential / laplace / cauchy quantiles, sinc, and the extra powers cbrt, qtrt, icbrt, iqtrt, pow_m01, pow_p15, pow_m3. Precision is one of 8, 12, …, 32. Each call runs a constant number of eval_reduced or eval_window evaluations.

\code{cpp} auto y = grotto::eval_closed(grotto::closed::atan, 16, raw); \endcode

Defined in\n @ref grotto/closed_form.hpp

Exact steps

Counts and booleans come back as fixed-point integers, n << fractional_bits.

\code{cpp} auto floor_x = grotto::eval_dec_floor(raw, 16); auto digits = grotto::eval_dec_width(raw, 16); auto bits = grotto::eval_bit_widthstd::int32_t(raw, 16); \endcode

Decimal digit counts divide in a loop, so the time follows the number of digits. eval_bit_width, eval_bit_floor, eval_bit_ceil, eval_countl_one, and eval_has_single_bit scan the raw width, Θ(width) bit operations and at most 64 shifts. eval_logstar is five magnitude comparisons, Θ(1). eval_deg2rad and eval_rad2deg are one scale each.

Also eval_dec_ceil, eval_oct_width, eval_b64_width, eval_value_length (base 8, 10, or 64), eval_has_single_digit, eval_bit_floor, eval_bit_ceil, eval_countl_one, eval_has_single_bit, eval_deg2rad, and eval_rad2deg. make_exact_step_lut builds the matching easy_lut.

Defined in\n @ref grotto/exact_steps.hpp

Gadget functors

grotto/gadgets.hpp still includes the decimal and exponential reference headers. The functors that those headers used to provide are deprecated: call eval_reduced, eval_window, make_*_lut, or exact_constant instead. gadget_hints<T> holds the old domain, degree, and pole notes for a functor type.

Defined in\n @ref grotto/gadgets.hpp, @ref grotto/gadget_hints.hpp