Checkpoint the party/runtime stack before share-program and malicious-mode work.

Ship the TLS mesh, composer, Beaver/Yao/leaf MPC, prep/online paths, apps, and docs so the tree is pushable before elevating share_expr, security_mode, and prep resume.

Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
Ryan Henry 2026-09-28 05:59:19 -06:00
parent 695f8e84f7
commit 0d22946a0e
1835 changed files with 170291 additions and 2849 deletions

View file

@ -30,7 +30,10 @@ you want.
| [dpf::eval_interval](@ref dpf/eval_interval.hpp) | An inclusive range, one share per input. |
| [dpf::eval_sequence](@ref dpf/eval_sequence.hpp) | A sorted list of inputs. |
| [dpf::eval_full](@ref dpf/eval_full.hpp) | Every input in the domain. |
| [dpf::reconstruct](@ref dpf/secret_share.hpp) | Both shares. Leaf shares subtract. Comparison shares add. |
| [dpf::reconstruct](@ref dpf/secret_share.hpp) | Both shares. Leaf shares subtract. Comparison shares add. Shamir shares use Lagrange. |
| [dpf::shamir::deal](@ref dpf/shamir.hpp) / `make_shamir_shares` | (K,N) Shamir shares over `fp61` or `gf2n`. For `gf2n`, `N < 2^k`. `(2,3)` is `make_shamir_shares(secret, slope)`. |
| [dpf::shamir::share_secret](@ref dpf/random.hpp) | The same split with uniform higher coefficients. |
| [dpf::shamir::reconstruct](@ref dpf/shamir.hpp) | Any K shares. Further shares are a consistency check, not a correction. |
| [dpf::eval_point](@ref dpf/interval.hpp) `(dpf::ic, ...)` | One public input on an interval key. |
## Comparisons and tags
@ -50,19 +53,55 @@ you want.
The catalogs, with the types that are inputs and the types that are outputs:
- [Input types](@ref input_types): integers, `modint`, `xint`, `bitstring`, `keyword`, `keyword2`, fixed-point.
- [Output types](@ref output_types): `bit`, `twobit`, `nyble`, fields, curve points, shares, `vec`.
- [Output types](@ref output_types): `bit`, `twobit`, `nyble`, `gf2` through `gf264`, prime fields, curve points, shares, `vec`.
`bit`, `twobit`, and `nyble` are packed output lanes. `keyword2` is a domain, not a leaf.
`bit`, `twobit`, `nyble`, and `gf2` / `gf22` / `gf24` are packed output lanes. `keyword2` is a domain, not a leaf.
## After the offset is public
[Grotto](@ref guided_tour) evaluates a function of x once the public offset is open.
Several piecewise LUTs share one comparison via [make_lut_union_plan](@ref grotto/lut_union.hpp).
The pages are [offset Horner, jets, and ring switch](@ref jet_and_ring) and
[representation shift and twisted jets](@ref repr_and_twist).
Haar and bior(5,3) tables are [make_haar_dwt_lut](@ref grotto/dwt_lut.hpp)
and [make_bior53_dwt_lut](@ref grotto/dwt_lut.hpp).
## Multiplication and sessions
| Call | What it does |
| --- | --- |
| [dpf::beavers::session](@ref dpf/beaver.hpp) | ABY2.0 blinds, products, dots, and polynomial schedules. |
| [dpf::yao::b2y](@ref dpf/yao_share.hpp) / `a2y` / `fss2y` / `rss2y` | A leaf share to LSB-first XOR bits. Point leaves use `b2y`. Comparisons use `a2y`. |
| [dpf::yao::y2b](@ref dpf/yao_share.hpp) / `y2a` / `y2fss` / `y2rss` | Those bits back to the leaf's share type. |
| [dpf::yao::netlist](@ref dpf/yao.hpp) / `session::eval` | The boolean circuit on those bits. Party 0 garbles. |
| [dpf::arith_garble::circuit](@ref dpf/arith_garble.hpp) | Free add, public scale, projection. Ball–Malkin–Rosulek. |
| [dpf::flute::eval_pair](@ref dpf/flute.hpp) / `eval_trio` | Public LUT on masked bits. Two or three online bits per output. A DPF point stays a key. |
| [dpf::yao::eval_if](@ref dpf/yao_stack.hpp) / `eval_one_hot` | Stacked branch and k-way switch. Rows follow the heaviest branch. |
| [dpf::yao::aes128](@ref dpf/yao_aes.hpp) / `aes_mmo` | Packaged AES-128 and the zero-key MMO block on that session. |
| [dpf::beavers::schedule_objective](@ref dpf/beaver.hpp) | `prep` peels for Appendix E; `rounds` keeps one online round. |
| [dpf::protocol::composer](@ref dpf/compose.hpp) | Domain-tagged FSS / ABY / RSS strands on one RoundSink plan. |
| [dpf::shuffle::shuffle_hidden_pass](@ref dpf/shuffle.hpp) | Hidden reorder of an RSS column. A secret index stays a DPF. |
| [dpf::protocol::composer::shuffle_hidden](@ref dpf/compose.hpp) | Three `shuffle_send` waves for that reorder. |
| [dpf::protocol::composer::client_servers](@ref dpf/compose.hpp) | PIR: one upload round, one answer round, no server-server open. |
| [dpf::protocol::plan_to_schedule](@ref dpf/compose.hpp) / `drive_via_schedule` | Lower a plan onto `schedule_session` (edge, receive rule, branch/next). |
| [dpf::protocol::schedule_session](@ref dpf/protocol.hpp) | Ready instance runs on this thread; flush sends the largest prefix per edge. |
| [dpf::net::edge_mesh](@ref dpf/net/edge_mesh.hpp) / `make_memory_star` | N duplex RoundSinks (star / clique / dealer). |
| [dpf::protocol::session_host](@ref dpf/session_host.hpp) | Queue micro-plans on a durable mesh. |
| [dpf::protocol::iknp_setup_graph](@ref dpf/iknp_graphs.hpp) / `du_atallah_mul_graph` | IKNP / Du-Atallah / star upload-answer as schedule rounds. |
| [dpf::protocol::pirsona_bitmore_fetch](@ref dpf/mesh_apps.hpp) / `hushmap_add_schedule` | PIRsona BitMore fetch and hushmap ADD skeletons. |
| [dpf::protocol::drive_star](@ref dpf/app_plans.hpp) / named `*_plan` helpers | Star drive + application micro-plans (PIR, mailbox, SUBLEQ, Pika, …). |
| [dpf::app::run](@ref dpf/app_flow.hpp) / `run_plan` | Drive both parties and print `rounds` and `bytes`. |
| [dpf::log](@ref dpf/log.hpp) / [app::start_logging](@ref dpf/run_log.hpp) | Leveled run log and provenance banner. |
| [dpf::experiment](@ref dpf/experiment.hpp) / [Logging & statistics](@ref experiment_costs) | Replayable master seed; CSV cost breakdown. |
| [dpf::app::measure_plan](@ref dpf/app_flow.hpp) / `run_measured` | Drive a plan under an experiment; optional `DPF_EXPERIMENT_DIR` CSV dump. |
| [dpf::app::run_fleet](@ref dpf/app_flow.hpp) | Many instances. Parked receives yield to the side that is behind. |
## Where to read next
- [Which DPF?](@ref which_dpf) if you are still choosing the object.
- [Evaluating DPFs](@ref evaluation) for point, interval, sequence, and full-domain cost.
- [Network, parties, and MPC](@ref network_and_mpc) for the runtime around the keys.
- [Logging, statistics, and experiments](@ref experiment_costs) for the run log and cost CSVs.
- [Protocol composition](@ref protocol_compose) for fused walks, early-stop, and RSS refresh.
- [Bibliography](@ref bibliography) for the papers behind the keys.
- [Application mockups](@ref applications) for the DPF step inside a larger protocol.

View file

@ -5,9 +5,18 @@ Each one is a single process.
A dealer stands in where the paper generates keys from shares.
They compile from the repository root:
c++ -std=c++17 -march=native -I include -I thirdparty examples/applications/duoram3.cpp
c++ -std=c++17 -march=native -pthread -I include -I thirdparty examples/applications/duoram3.cpp
The online compose sketch at the end of each file uses
[`dpf::app::run_measured`](@ref dpf/app_flow.hpp): it prints rounds, live
bytes, wall/CPU, PRG evals, random bytes, and the master seed hex. Set
`DPF_EXPERIMENT_DIR=/tmp/run` to also write the CSV tables described in
[Experiments](@ref experiment_costs).
The same compile line with the other filenames builds the rest. Timing for
the party protocols, including word garbling, stacked branches, FLUTE, and
the hidden column reorder, is `party_bench` / `profile_party` (see
[The battery](@ref experiment_costs)).
The same line with the other filenames builds the rest.
Optional Python bindings configure with `-DLIBDPF_PYTHON=ON` in the test
build directory, then `make pydpf` (and `make pydpf_pytest`).
`pydpf` exposes point / interval / full / sequence / recipe eval on
@ -16,7 +25,6 @@ and `it_dpf3`. Not in this module yet: geneval, Doerner–Shelat, VDPF
`prove`/`sketch`, DCF, or grotto.
What those programs had to do by hand is [the library surface underneath](@ref application_gaps).
| Sketch | What the DPF step is |
| --- | --- |
| [3-party Duoram](@ref app_duoram) | Unit key, rotate, inner product |
@ -38,17 +46,25 @@ What those programs had to do by hand is [the library surface underneath](@ref a
| [Range count](@ref app_range_count) | Interval payload |
| [Floram](@ref app_floram) | ORAM read |
| [Three-server PIR](@ref app_pir3) | Information-theoretic DPF |
| [Protocol composition](@ref protocol_compose) | Fused walks, early-stop, RSS refresh, ABY scale |
| PIRsona fetch | BitMore star upload+answer (`pirsona_fetch.cpp`) |
| Hushmap ADD | Dealer tape + two opens (`hushmap_add.cpp`) |
| [What the walk folds in](@ref application_gaps) | Library surface those programs used to do by hand |
## 3-party Duoram {#app_duoram}
\htmlonly
<div class="eli5"><b>ELI5.</b> Preprocessing plants a unit DPF at a random index r. Online the parties open i* − r, rotate the expanded unit vector by that public shift, and dot with the memory. The update rotates a payload vector the same way and adds it in.</div>
\endhtmlonly
Vadapalli, Henry, and Goldberg ([USENIX Security 2023](@ref bib_duoram)) keep a memory in
shares and read or add at a secret index.
Preprocessing builds unit DPFs at a random index `r`.
Online, the parties open `i* - r` and cyclic-shift the expanded vector.
The read is the dot product of that vector with the memory.
The update adds a payload vector, shifted the same way.
Opening the whole memory, rather than one index, is a hidden shuffle of
that column ([an array of shares](@ref share_shuffle)), not another DPF.
The program uses one dealer unit key for the read and one payload key
for the update.
@ -71,6 +87,10 @@ three evaluators.
## MPC SUBLEQ {#app_subleq}
\htmlonly
<div class="eli5"><b>ELI5.</b> The address is not known when the keys are built, so the unit vectors are expanded early into a deferred buffer. Online, opening the address rotates that buffer. The read is a dot with memory; the write adds the scaled unit vector back.</div>
\endhtmlonly
Jiang and Henry ([MSc thesis, University of Calgary](@ref bib_subleq)) emulate the
subtract-and-branch-if-less-than-or-equal-to-zero (SUBLEQ) OISC for
private function evaluation. One instruction is
@ -109,6 +129,10 @@ and out-of-bounds prefix-parity checks from the thesis.
## BitMore, `2^L` servers {#app_bitmore}
\htmlonly
<div class="eli5"><b>ELI5.</b> The label is L bits. Each bit is its own 1-bit DPF, expanded over the whole domain. Stacking the L bit-vectors and reading a column produces the 2^L-server answer for that label.</div>
\endhtmlonly
Hafiz and Henry ([PoPETs 2019](@ref bib_bitmore), §5.2) query `ell = 2^L` servers with `L`
independent 1-bit DPFs, all at the same row.
Server `j` receives key number `j_e` from DPF `e`.
@ -131,6 +155,10 @@ The two answers XOR to the record.
## Keyword PIR {#app_keyword}
\htmlonly
<div class="eli5"><b>ELI5.</b> The keyword is hashed into cuckoo buckets. Each bucket is a point key, and the record is the inner product of the probed buckets with the dictionary. The S&amp;P 2025 seed-packing of those buckets is not what this program does.</div>
\endhtmlonly
Gilboa and Ishai ([EUROCRYPT 2014](@ref bib_dpf2014)) retrieve one record by a keyword.
[dpf::keyword](@ref dpf/keyword.hpp) is the domain, so the DPF point is
the keyword itself.
@ -158,6 +186,10 @@ stack.
## Prio and the heavy-hitter prefix walk {#app_prio}
\htmlonly
<div class="eli5"><b>ELI5.</b> A one-hot vote is a unit DPF in the Prio field. Each server adds the expanded vector into a running histogram. The heavy-hitter pass is the same key read at successive prefixes: the servers add the opened prefix shares and keep the heavy nodes.</div>
\endhtmlonly
Corrigan-Gibbs and Boneh ([NSDI 2017](@ref bib_prio)) aggregate client encodings.
A frequency count is a one-hot vector.
A unit DPF is that vector, compressed.
@ -182,6 +214,10 @@ and add the opened values.
## I-DPF max and k-th {#app_idpf_agg}
\htmlonly
<div class="eli5"><b>ELI5.</b> Each secret integer is one incremental DPF with a unit payload on every prefix. At each bit the servers open the two children. Max keeps the child that holds mass. The k-th keeps the 1-child when its count covers k, and otherwise descends the 0-child with k reduced.</div>
\endhtmlonly
Cheng, Mitrokotsa, Zhang, and Hartmann ([ePrint 2024/1190](@ref bib_idpfagg)) aggregate
secret values with an incremental DPF.
Communication tracks the bit length of the domain, not how many secret
@ -200,6 +236,10 @@ share that walk with the gtest.
## LLAMA {#app_llama}
\htmlonly
<div class="eli5"><b>ELI5.</b> The comparison is the gate. eval_point on a gt or lt key opens to the payload when the public query is on the true side of the secret, including across the sign bit, because signed inputs flip the high bit before the walk.</div>
\endhtmlonly
Gupta, Kumaraswamy, Chandran, and Gupta ([ePrint 2022/793](@ref bib_llama)) evaluate a
nonlinear gate from a dealer key and one opened masked input
`x_hat = x + r`.
@ -214,6 +254,10 @@ to 1.
## Pika {#app_pika}
\htmlonly
<div class="eli5"><b>ELI5.</b> The dealer keys a unit DPF at a fresh r and the parties open x = r − a. Rotating the public table by x and dotting with the DPF reads the entry at a. The sign of an early-stop bit leaf is recorded at keygen, so the evaluators never open r.</div>
\endhtmlonly
Wagh ([PoPETs 2022](@ref bib_pika), Fig. 1) looks up `Func(a)` in a table of a bounded
domain.
The dealer keys a unit DPF at a fresh index `r` and the parties open
@ -224,11 +268,17 @@ A word payload of `1` opens to `+1`.
The paper's early-stop bit leaf opens to `+1` or `-1`; the dealer records
that sign at keygen with `dpf::unit_sign` (the final control bit `Gen`
sees), so the evaluators never open `r`.
A networked early-stop *walk* (BGI Remark 3.4) drops ν CW rounds with
`fss_point_early_stop` on a [composer](@ref protocol_compose).
\include{cpp} applications/pika.cpp
## Express {#app_express}
\htmlonly
<div class="eli5"><b>ELI5.</b> A mailbox write is a full-domain add of one DPF into the shared array. Every box is touched by the expand; only the programmed box survives when the shares are opened.</div>
\endhtmlonly
Eskandarian, Corrigan-Gibbs, Zaharia, and Boneh ([USENIX Security 2021](@ref bib_express),
§3.1) write one mailbox.
Two servers hold subtractive shares of the mailboxes.
@ -246,10 +296,19 @@ The extractable full-domain leaf now matches point evaluation on every
lane, so this one call replaces the earlier `eval_point`-per-address
loop and the separate `sketch_fold` pass.
On a networked walk the audit share rides in the last correction-word
flush — `fss_point_fused` / `level_walk_fused` on a
[composer](@ref protocol_compose) — so the sketch does not add a round.
The same shape is Sabre's proof token.
\include{cpp} applications/express.cpp
## PRAC {#app_prac}
\htmlonly
<div class="eli5"><b>ELI5.</b> Binary search needs a unit vector on a stride that grows by one bit per comparison. One incremental key holds all of those prefixes, and a prefix inner product dots a stride without a separate point key per slot. A heap update is a three-lane vector at the parent and its two children.</div>
\endhtmlonly
Sasy, Vadapalli, and Goldberg ([ePrint 2023/1897](@ref bib_prac)) run dynamic data
structures on a 3-party Duoram.
The new DPF shapes are an incremental key and a wide leaf.
@ -280,6 +339,10 @@ The protocol appends each comparison bit after the key exists.
## Splinter {#app_splinter}
\htmlonly
<div class="eli5"><b>ELI5.</b> The secret WHERE value is a unit DPF. The server dots it with a column that was already summed by attribute, which is the grouped SUM, and with an all-ones column, which is the COUNT. The queried attribute is not revealed.</div>
\endhtmlonly
Wang, Yun, Goldwasser, Vaikuntanathan, and Zaharia ([NSDI 2017](@ref bib_splinter)) answer
private queries on public data with two-server FSS.
The client's private `WHERE` value is a unit DPF at that attribute.
@ -296,6 +359,10 @@ one selector DPF; Splinter composes several FSS instances for those.
## Mastic {#app_mastic}
\htmlonly
<div class="eli5"><b>ELI5.</b> This is the Poplar prefix walk with a weight instead of 1 on every prefix. Servers sum the prefix shares across clients and drop prefixes under the threshold. A path sketch can check that each client programmed a single path.</div>
\endhtmlonly
Mastic (private weighted heavy-hitters and attribute-based metrics) is
Poplar's prefix walk with a weight payload.
Each client keys an [idpf](@ref dpf/placement.hpp) whose β on every
@ -311,6 +378,10 @@ heavy with total weight 8.
## Waldo {#app_waldo}
\htmlonly
<div class="eli5"><b>ELI5.</b> Each append is a fresh unit DPF added into the value shares; old events are not rewritten. A threshold query is a comparison inner product of those shares with a public magnitude column, so the sum past a secret threshold does not reveal the threshold or the matches.</div>
\endhtmlonly
Dauterman, Rathee, Popa, and Stoica ([S&P 2022](@ref bib_waldo)) build a private time-series
database from FSS.
The store is append-only: each event is a fresh unit DPF folded into the
@ -330,6 +401,10 @@ comparison key.
## Sabre {#app_sabre}
\htmlonly
<div class="eli5"><b>ELI5.</b> The write is the same full-domain add as Express. The audit folds a constant-size proof token along that key. verify accepts one honest point and rejects a key that was hot in more than one place.</div>
\endhtmlonly
Vadapalli, Storrier, and Henry ([S&P 2022](@ref bib_sabre)) send anonymous messages with a
fast audit.
The write is Express's full-domain add (`eval_full_add_into`).
@ -337,6 +412,8 @@ The audit is a *verifiable* DPF proof rather than Express's `fp61`
sketch: `prove_full(key, dpf::prove(π))` folds a constant-size token per
party, and `dpf::verify(π0, π1)` accepts an honest single-point write and
rejects the mismatched fold a multi-point key produces.
Networked audits pack the proof share into the last CW with
`fss_point_fused` ([protocol composition](@ref protocol_compose)).
Still by hand: Sabre's blame / accountability phase that identifies a
cheating client is protocol logic above the DPF proof.
@ -345,6 +422,10 @@ cheating client is protocol logic above the DPF proof.
## A (2,3) ledger {#app_ledger23}
\htmlonly
<div class="eli5"><b>ELI5.</b> An append is one verifiable three-party point at the slot. The three proof tokens are checked before the point is added into the slot shares. Any two servers reconstruct a balance.</div>
\endhtmlonly
A replicated ledger held as (2-of-3) shares by three servers, on this
group's `dpf3` VDPF+ construction.
Each append is one `make_dpf3(slot, amount, dpf::verifiable{})`; a
@ -362,6 +443,10 @@ Still by hand: the transaction / consensus layer around the append
## Private set intersection {#app_psi}
\htmlonly
<div class="eli5"><b>ELI5.</b> Each element on one side is a unit DPF. Dotting it with the other side's table is the membership test. The cuckoo layout and the OPRF that would sit around that test are not in the DPF call.</div>
\endhtmlonly
Kolesnikov, Kumaresan, Rosulek, and Trieu ([CCS 2016](@ref bib_kkrt)) test membership
with an oblivious PRF.
On this domain the PRF is a table both servers hold.
@ -380,6 +465,10 @@ DMPF seed packing needs a PCG this library does not provide.
## Range count {#app_range_count}
\htmlonly
<div class="eli5"><b>ELI5.</b> A value falls in [lo, hi) when the greater-than bit at hi and the greater-than bit at lo differ. Each secret value is one comparison key. The count is the sum of those opened bits.</div>
\endhtmlonly
Each secret value is one comparison.
The interval `[lo, hi)` is public.
`eval_point(dpf::cmp, key, q)` opens to 1 when `q` is strictly above the
@ -396,6 +485,10 @@ This program is the other direction, secret values and a public range.
## Floram {#app_floram}
\htmlonly
<div class="eli5"><b>ELI5.</b> The address is shared, not known to a dealer. The Doerner–Shelat opening builds the unit key level by level, and the read or write is the inner product of that key with the array.</div>
\endhtmlonly
Doerner and shelat ([CCS 2017](@ref bib_ds)) read and write an array at a secret
address.
Both parties see the memory.
@ -413,6 +506,10 @@ is the FSS access.
## Three-server PIR {#app_pir3}
\htmlonly
<div class="eli5"><b>ELI5.</b> All three servers hold the database. A Shamir DPF3 key makes each server return an inner product; any two of those field elements open the record. The information-theoretic key instead adds all three inner products.</div>
\endhtmlonly
The database is public and replicated on three servers.
The computational path is one (2,3) point key from
@ -432,8 +529,21 @@ three dots sum to the record. Distinct from `make_dpf3`.
## What the walk now folds in {#application_gaps}
\htmlonly
<div class="eli5"><b>ELI5.</b> Rotate-then-dot, the sign of a bit leaf, bit columns, prefix dots, and path sketches used to be loops around eval. They are parameters of one walk now.</div>
\endhtmlonly
The calls the eight programs used to build by hand are now the library
surface. See [dpf/eval_walk.hpp](@ref dpf/eval_walk.hpp).
Multi-protocol *schedules* (FSS + ABY + RSS on one RoundSink) are
[protocol composition](@ref protocol_compose). Each listing under
`examples/applications/` records that paper's online flow on a
`composer` and calls `dpf::app::run`, which drives both parties on an
in-process sink and prints `name rounds= bytes=`. That line is the
experiment: compare it with the round and bandwidth column of the
paper. PIR listings are a client and two or three servers (one upload
round, one answer round). The servers do not open shares with each
other.
## Shift, then add {#gap_shift}
@ -562,3 +672,7 @@ full-domain `H` publishes nothing. The walk is O(|H| · n) and never
materializes the domain. An audit opening of a replica-seed pool is that
copath with `program_hidden = false`, so the live seeds stay out. The
one-point layout stays for PSI; `{α}` with programming agrees with it.
\htmlonly
<div class="tldr"><b>TL;DR.</b> Each program is only the DPF step, in one process, with a dealer standing in for shared keygen. Memory reads are a unit vector, a public shift or a prefix, and a dot. PIR and grouped sums are inner products. Heavy hitters are prefix walks. Mailbox writes and the ledger are full-domain adds, plus a proof when the write must be a single point.</div>
\endhtmlonly

151
doc/pages/arith.md Normal file
View file

@ -0,0 +1,151 @@
# Arithmetic share runtime {#arith_runtime}
\htmlonly
<div class="eli5"><b>ELI5.</b> Beside FSS keys you get ordinary secret shares: add them locally, multiply with Beaver or RSS, flip between bits and numbers with edaBits, and truncate fixed-point products. The composer schedules those opens next to FSS walks.</div>
\endhtmlonly
Semi-honest 2PC and honest-majority 3PC. Openings may carry Shark/SPDZ IT-MACs
from the beaver session. A leaf that must enter a boolean netlist uses
[b2y / a2y](@ref yao_leaf) and comes back with `y2b` / `y2a`. There is no
Yao domain on the composer.
## Domains
| Domain | Meaning |
| --- | --- |
| `a` / `b` / `fss` | Existing additive, subtractive, FSS leaf |
| `rss` / `y` | Existing replicated / product factor |
| `bin` | (2,2) XOR bit shares (packed) |
| `bin_rss` | (2,3) replicated bits |
`composer::as` never casts to or from `bin` / `bin_rss`. Use `bin_a2b` /
`bin_b2a` / `bin_inject` (see [compose.hpp](@ref dpf/compose.hpp) opcodes
400–405).
## Building blocks
| Header | Role |
| --- | --- |
| [rss_seed.hpp](@ref dpf/rss_seed.hpp) | Pairwise PRG seeds, zero-sharing, RSS mul local |
| [ot_pack.hpp](@ref dpf/ot_pack.hpp) | Consumable B2A / bit pads or dealer dabits |
| [edabit.hpp](@ref dpf/edabit.hpp) | daBits and edaBits; A2B |
| [bit_inject.hpp](@ref dpf/bit_inject.hpp) | Bit × arithmetic (2PC / RSS) |
| [trunc.hpp](@ref dpf/trunc.hpp) | Probabilistic and exact truncate; mul_trunc |
| [share_cmp.hpp](@ref dpf/share_cmp.hpp) | Share–share compare, ReLU, max, range, div |
| [share_vec.hpp](@ref dpf/share_vec.hpp) | Lane-block share vectors |
| [fixed_share.hpp](@ref dpf/fixed_share.hpp) | `fixed<Int,Frac>` |
| [share_expr.hpp](@ref dpf/share_expr.hpp) | Imperative recorder |
| [gilboa.hpp](@ref dpf/gilboa.hpp) | One-off product / tape fill |
| [matmul.hpp](@ref dpf/matmul.hpp) | Matrix triples and matmul |
| [shuffle.hpp](@ref dpf/shuffle.hpp) | Hidden reorder of an RSS column. A secret index stays a DPF |
| [arith_garble.hpp](@ref dpf/arith_garble.hpp) | Free add and a unary projection, after a word has left the key |
| [flute.hpp](@ref dpf/flute.hpp) | Public table on short masked bits. A DPF point stays a key |
| [cost_pass.hpp](@ref dpf/cost_pass.hpp) | DCF vs edaBit vs A2B strategy |
| [yao.hpp](@ref dpf/yao.hpp) | Half-gates netlist on a leaf's XOR bits. Party 0 garbles |
| [yao_share.hpp](@ref dpf/yao_share.hpp) | Leaf share to those bits and back (`b2y`, `a2y`, `fss2y`, `rss2y`) |
| [yao_aes.hpp](@ref dpf/yao_aes.hpp) | The PRG's zero-key MMO block, and AES-128 under a shared key |
## When to use which comparison
- **One input public or keyed:** DCF / `geneval_*` / interval keys.
- **Both inputs arithmetic shares:** `share_cmp` (mask, open, MSB / DCF at the public difference).
- **The leaf must enter a deep bit circuit:** [Yao](@ref yao_leaf). Not a comparison, a mux, or a public-offset LUT.
## A small word after the key {#arith_garble_word}
Comparisons, intervals, and public-offset polynomials stay on the key.
[Grotto](@ref jet_and_ring) covers a public offset. This section is the
word you already hold as shares, when the next step is addition, a public
scale, or a unary map, and you do not want an opening between the gates.
[arith_garble.hpp](@ref dpf/arith_garble.hpp) is that circuit
([Ball, Malkin, and Rosulek, CCS 2016](@ref bib_garble_gadgets)). Addition
is free. Scaling by a public constant coprime to the modulus is free. A
unary map of a mod-`m` wire sends `m − 1` ciphertexts. A threshold of `b`
bits that already live in `Z_{b+1}` is one such map, so the row count is
`b`. A product in a small prime field is the discrete-log reduction: project
to the exponent, add, project back, and drop the zero cases.
The beaver [session](@ref beaver_triples) is the other tool. It opens one
masked wire and multiplies interactively. It does not garble these gadgets.
```cpp
dpf::arith_garble::circuit c;
auto x = c.input(7);
auto y = c.input(7);
c.out(c.mul(c.add(x, y), x));
auto opened = dpf::arith_garble::eval_pair(c, in);
```
`opened.mask` is the garbler's share and `opened.color` is the evaluator's.
`open_shares` subtracts them mod the output modulus.
## A public table on short shares {#flute_lut}
A secret point in a public table is a DPF dotted with that table. FLUTE is
the case where the index is already a handful of masked bits sitting in an
ABY2.0 or three-party XOR sharing, and the result has to stay in that
sharing.
[flute.hpp](@ref dpf/flute.hpp) follows Brüggemann, Hundt, Schneider, Suresh,
and Yalame, [IEEE S&P 2023](@ref bib_flute). The table is an inner product of
those bits. The online exchange is two bits per output bit for two parties,
and three bits per output bit for `eval_trio`, independent of how wide the
index is. The index is at most 8 bits. Bit 0 of the index is the least
significant bit of the row.
```cpp
auto pair = dpf::flute::eval_pair(delta, n_out, columns, bits);
auto trio = dpf::flute::eval_trio(delta, n_out, columns, bits);
```
`columns[w * 2^δ + row]` is output bit `w` on that row. `opened` is the
clear bit. `masked` XOR the party's mask shares is the same bit.
## An array of shares, not a secret index {#share_shuffle}
A secret index is a DPF. One read of a shared memory is a unit key, a public
rotate, and a dot product, as in [3-party Duoram](@ref app_duoram). The
shuffle is the other job: the parties already hold a share of every row, and
the next step opens values. The opened order must not be the stored order.
`shuffle_party` derives one permutation from `k01` and applies it to every
component. A caller who passes the whole seed bundle can recompute that
order. Use it when the order is allowed to be known.
`shuffle_hidden_pass` is the order no single party should learn. Three
passes use the pairwise seeds `k01`, `k12`, and `k20`, leaving out the party
who does not hold that seed (party 2, then 0, then 1). Each party is given
`rss::party_seeds` only. On a pass, the two parties who share the seed
permute a two-party split of the column. The left-out party sends nothing
and receives one fresh component. Both messages go to the sender's RSS
neighbor. `permute` means `out[i] = in[pi[i]]`. Replaying the clear column
is `π01`, then `π12`, then `π20`.
```cpp
auto opened = dpf::shuffle::shuffle_hidden_triple(column, bundle, /*index=*/0);
```
`composer::shuffle_hidden(n, value_bytes)` records those three exchanges as
`shuffle_send`. `aux` on each node is the left-out party.
This is for a column you already share: a Duoram memory before a bulk open,
a histogram, a PSI payload. It does not replace a key. One hidden cell is
still a DPF. A comparison or a sort by a secret key is a DCF or
`share_cmp`. A gather to data-dependent indices is a DPF per access. The
seed permutation is chosen before the data.
## Truncation
- `trunc_prob` — local right shift, error in `{0,1}`, no round.
- `trunc_exact` — edaBit / carry correction.
- `mul_trunc` — product then shift (fixed-point multiply).
## Networking
`schedule_round` carries `phase::{setup,online}`, `round_dir::{duplex,send_next,recv_prev}`,
and `receive_rule::ring_next`. `drive_options` adds `pipeline_credit`, `cleartext`,
and `check_open`. Pairwise seeds replace `dealer_zero` via `rss_zero_mask`.
**Go deeper:** [Beaver](@ref beaver_triples), [compose](@ref protocol_compose),
[Grotto carry](@ref jet_and_ring).

View file

@ -62,6 +62,10 @@ Width literals (`100_u12`, `7_x12`, `1.5_fixed16`, `1_bit`, `2_twobit`,
## DPF Trees {#dpf_trees}
\htmlonly
<div class="eli5"><b>ELI5.</b> A key stores a root seed and one correction word per level. Expanding the seed walks the tree. On the secret path the correction forces the leaf to the programmed value; off that path the two parties' corrections cancel, so the opened value is zero.</div>
\endhtmlonly
Keys store a seed and a list of *correction words*.
Evaluation walks a binary tree from the root toward `x`.
At each level a correction word mixes the two children so only the secret
@ -79,3 +83,7 @@ full-domain evaluation versus `2N`. See [tree_traits.hpp](@ref dpf/tree_traits.h
For a slow, friendly walk through every feature, start at the
[guided tour](@ref guided_tour).
\htmlonly
<div class="tldr"><b>TL;DR.</b> A (2,2) DPF gives each party a short key for one secret point. make_dpf builds it. Point leaves open by subtraction and comparisons by addition. The key is a seed plus one correction word per level.</div>
\endhtmlonly

View file

@ -1,16 +1,39 @@
# Beaver triples {#beaver_triples}
\htmlonly
<div class="eli5"><b>ELI5.</b> A product opens as d = x − a and e = y − b, with a and b the preprocessing blinds. The product share is de plus the blinded cross terms, all local once d and e are public. The session keeps blinds that were already opened and only samples monomials it has not seen.</div>
\endhtmlonly
ABY2.0-style sessions open masked wires once (Patra, Schneider, Suresh,
and Yalame, USENIX Security 2021 / [ePrint 2020/1225](@ref bib_aby2)).
A fresh triple follows Beaver, [CRYPTO 1991](@ref bib_beaver): both masked factors are
reconstructed, and the product share is a local correction.
Optional MAC tags are the Shark/SPDZ check.
Constant-round word arithmetic is a separate gadget. [Ball, Malkin, and
Rosulek, CCS 2016](@ref bib_garble_gadgets) give free addition, free scaling
by a public constant, and a unary projection of `m − 1` ciphertexts.
[arith_garble.hpp](@ref dpf/arith_garble.hpp) is that circuit. A session does
not become one: it still opens δ once per wire. A public table on masked
bits is [FLUTE](@ref bib_flute) in [flute.hpp](@ref dpf/flute.hpp): the table
is a multi-fan-in inner product, and the online exchange is two bits per
output bit. `eval_trio` is the same product on three XOR shares of each mask.
One call that samples a list of formulae opens the new wires in one
round. Communication is one masked value per newly opened wire, plus a
tag share of the same width when MACs are on. Preprocessing is one blind
per wire and one product share per monomial.
**Go deeper:** [beaver.hpp](@ref dpf/beaver.hpp),
`schedule_objective::prep` (default) peels shared factors for Appendix-E
prep savings and may add interactive rounds.
`schedule_objective::rounds` emits the polynomial in one online round
(Pika / online Grotto). Composer-owned sessions use `rounds`.
Compose FSS walks, ABY products, and RSS refreshes on one sink with
[protocol composition](@ref protocol_compose).
**Go deeper:** [a small word](@ref arith_garble_word),
[a public table](@ref flute_lut), [beaver.hpp](@ref dpf/beaver.hpp),
[compose.hpp](@ref dpf/compose.hpp),
[F_Beaver](@ref beaver.hpp), [F_BeaverAuth](@ref beaver.hpp),
and the cost notes in the [guided tour](@ref tour_beaver).

View file

@ -169,7 +169,121 @@ USENIX Security 2021. Full version:
One public reconstruction per newly opened wire.
**Used in** [Beaver triples](@ref beaver_triples) · [Guided tour](@ref tour_beaver)
**Used in** [Beaver triples](@ref beaver_triples) · [Arithmetic share runtime](@ref arith_runtime) · [Guided tour](@ref tour_beaver)
### ABY {#bib_aby}
Daniel Demmler, Thomas Schneider, and Michael Zohner.
*ABY — A Framework for Efficient Mixed-Protocol Secure Two-Party Computation.*
NDSS 2015.
[ePrint 2014/386](https://eprint.iacr.org/2014/386)
Arithmetic / boolean / Yao sharing and conversions.
**Used in** [Arithmetic share runtime](@ref arith_runtime) · [A boolean function of a leaf](@ref yao_leaf)
### Half-gates {#bib_halfgates}
Samee Zahur, Mike Rosulek, and David Evans.
*Two Halves Make a Whole: Reducing Data Transfer in Garbled Circuits using Half Gates.*
EUROCRYPT 2015.
[ePrint 2014/756](https://eprint.iacr.org/2014/756)
Free-XOR AND rows. `yao::session` sends two blocks per AND.
**Used in** [A boolean function of a leaf](@ref yao_leaf) · [Arithmetic share runtime](@ref arith_runtime)
### Garbling gadgets {#bib_garble_gadgets}
Marshall Ball, Tal Malkin, and Mike Rosulek.
*Garbling Gadgets for Boolean and Arithmetic Circuits.*
CCS 2016.
[ePrint 2016/969](https://eprint.iacr.org/2016/969)
Free addition and public scaling in `(Z_m)^k`, and a unary projection of
`m − 1` ciphertexts. A fan-in-`b` symmetric gate is a projection of the sum.
`arith_garble.hpp` is this gadget. The session in `beaver.hpp` stays the
interactive one-open product.
**Used in** [Beaver triples](@ref beaver_triples) · [Arithmetic share runtime](@ref arith_runtime)
### Stacked garbling {#bib_stacked}
David Heath and Vladimir Kolesnikov.
*Stacked Garbling: Garbled Circuit Proportional to Longest Execution Path.*
CRYPTO 2020.
[ePrint 2020/973](https://eprint.iacr.org/2020/973)
One XOR-stack of the branch materials. Inactive branches are rebuilt from seeds.
**Used in** [A boolean function of a leaf](@ref yao_leaf)
### One-hot garbling {#bib_onehot}
David Heath and Vladimir Kolesnikov.
*One Hot Garbling.*
CCS 2021.
The same stack over `k` branches. The demux carries the inactive seeds.
**Used in** [A boolean function of a leaf](@ref yao_leaf)
### FLUTE {#bib_flute}
Andreas Brüggemann, Robin Hundt, Thomas Schneider, Ajith Suresh, and Hossein Yalame.
*FLUTE: Fast and Secure Lookup Table Evaluations.*
IEEE S&P 2023.
[ePrint 2023/499](https://eprint.iacr.org/2023/499)
A public LUT is a multi-fan-in inner product on ABY2.0 masked bits.
Online cost is two bits per output bit.
**Used in** [Beaver triples](@ref beaver_triples) · [Arithmetic share runtime](@ref arith_runtime)
### ABY3 {#bib_aby3}
Payman Mohassel and Peter Rindal.
*ABY3: A Mixed Protocol Framework for Machine Learning.*
CCS 2018.
[ePrint 2018/403](https://eprint.iacr.org/2018/403)
Honest-majority 3PC, RSS, probabilistic truncation, matrix triples.
**Used in** [Arithmetic share runtime](@ref arith_runtime)
### edaBits / CrypTFlow2 {#bib_edabits}
Daniel Escudero, Satrajit Ghosh, Marcel Keller, Rahul Rachuri, and Peter Scholl.
*Improved Primitives for MPC over Mixed Arithmetic-Binary Circuits.*
CRYPTO 2020.
[ePrint 2020/338](https://eprint.iacr.org/2020/338)
**Used in** [Arithmetic share runtime](@ref arith_runtime) · [A boolean function of a leaf](@ref yao_leaf)
### EzPC {#bib_ezpc}
Nishanth Chandran, Divya Gupta, Aseem Rastogi, Rahul Sharma, and Shardul Tripathi.
*EzPC: Programmable and Efficient Secure Two-Party Computation for Machine Learning.*
EuroS&P 2019.
[ePrint 2017/1109](https://eprint.iacr.org/2017/1109)
**Used in** [Arithmetic share runtime](@ref arith_runtime)
### Gilboa multiplication {#bib_gilboa}
Niv Gilboa.
*Two Party RSA Key Generation.*
CRYPTO 1999, LNCS 1666, pp. 116–129.
**Used in** [Arithmetic share runtime](@ref arith_runtime)
### Sharemind {#bib_sharemind}
Dan Bogdanov, Sven Laur, and Jan Willemson.
*Sharemind: A Framework for Fast Privacy-Preserving Computations.*
ESORICS 2008, LNCS 5283, pp. 192–206.
**Used in** [Arithmetic share runtime](@ref arith_runtime)
### Beaver triples {#bib_beaver}
@ -221,6 +335,22 @@ Offset Horner is not that piecewise-polynomial construction.
**Used in** [Grotto](@ref jet_and_ring) · [Guided tour](@ref tour_grotto)
### Wave Hello {#bib_wave}
José Reis, Mehmet Ugurbil, Sameer Wagh, Ryan Henry, and Miguel de Vega.
*Wave Hello to Privacy: Efficient Mixed-Mode MPC using Wavelet Transforms.*
PoPETs 2025(2), pp. 697–718.
<a href="https://doi.org/10.56553/popets-2025-0083">Publisher</a> ·
[ePrint 2025/013](https://eprint.iacr.org/2025/013) ·
<a href="reis-ugurbil-wagh-henry-de-vega-wave-hello-eprint-2025-013.pdf">PDF</a>
Equations (7) and (8) are the cleartext Haar and bior(5,3) lookup.
`make_haar_dwt_lut` and `make_bior53_dwt_lut` evaluate those tables.
The online phase pairs Haar with a deterministic Pika truncation and
bior(5,3) with segment parity.
**Used in** [Wavelet lookup tables](@ref dwt_luts) · [Guided tour](@ref tour_grotto)
## Protocols {#bib_sec_protocols}
### Duoram {#bib_duoram}
@ -379,3 +509,7 @@ The DPF work is a prepaid wildcard unit vector, rotated once the
address is opened.
**Used in** [MPC SUBLEQ](@ref app_subleq)
\htmlonly
<div class="tldr"><b>TL;DR.</b> The point key is ePrint 2018/707, not the longer EUROCRYPT 2015 key. Proofs and cuckoo multipoint are 2021/580. Comparisons are 2020/1392. The dealer-free opening is 2017/827, and the pads are IKNP plus the 2015/267 base OT. Three-party spines are 2024/1658. The information-theoretic table is 2023/028. Grotto is 2023/108.</div>
\endhtmlonly

View file

@ -3,13 +3,25 @@
Everything past a plain two-party point key.
The home page is the short version. Open a page for the snippet and the header.
## Keys and evaluation
- [Verifiability & authenticity](@ref verifiability)
- [Programmability](@ref programmability)
- [Comparisons & ranges](@ref comparisons)
- [Multipoint keys](@ref multipoint_keys)
- [Multiparty & 3-server](@ref multiparty)
- [Dealer-free keygen](@ref dealer_free)
- [Beaver triples](@ref beaver_triples)
- [Grotto](@ref jet_and_ring)
- [Point-programmable vector commitments](@ref ppvc_manual)
- [Application sketches](@ref applications)
## Network and MPC around the keys
These are first-class in the tree — see the overview, then dig in:
- [Network, parties, and MPC](@ref network_and_mpc) — map of the runtime stack
- [Protocol composition](@ref protocol_compose)
- [Beaver triples](@ref beaver_triples)
- [Arithmetic share runtime](@ref arith_runtime)
- [A boolean function of a leaf](@ref yao_leaf)
- [Logging, statistics, and experiments](@ref experiment_costs) — run log, CSVs, replayable seeds

View file

@ -1,5 +1,9 @@
# Comparisons & ranges {#comparisons}
\htmlonly
<div class="eli5"><b>ELI5.</b> The predicate is part of the key, not a test you run after expanding a point key. Greater-than, less-than, and equality each return one payload on the true side and another on the false side. Interval containment is one key for both endpoints. Shares of a comparison add.</div>
\endhtmlonly
A distributed comparison function returns a payload when a predicate holds
on the secret point. Shares are additive: `reconstruct` adds them.

204
doc/pages/compose.md Normal file
View file

@ -0,0 +1,204 @@
# Protocol composition {#protocol_compose}
\htmlonly
<div class="eli5"><b>ELI5.</b> Record every expand, multiply, and open as a node with a share domain. Identical work is interned. Independent opens share one RoundSink round; a dependency chain becomes successive waves. The schedule is what a hand-tuned Express or Duoram party would have written by hand.</div>
\endhtmlonly
`dpf::protocol::composer` ([compose.hpp](@ref dpf/compose.hpp)) records
multi-protocol strands on one sink. Values are tagged with a share
domain matching [secret_share.hpp](@ref dpf/secret_share.hpp):
| Domain | Meaning |
| --- | --- |
| `fss` | FSS / DPF leaf share |
| `a` | (2,2) additive |
| `b` | (2,2) subtractive |
| `rss` | (2,3) replicated |
| `y` | (3,3) additive (RSS product factor) |
Party-count changes are never implied: use `reshare` (or `rss_from_y`
for `y`→`rss`). Local casts use `as`.
## Building a schedule
```cpp
dpf::protocol::composer c(/*party=*/0);
auto seed = c.input(dpf::protocol::domain::fss, 16);
auto leaf = c.fss_point(seed, /*depth=*/8, /*slot_bytes=*/16);
auto p = c.default_plan(); // == schedule(); RoundSink uses this
// p.rounds() == p.exchange_waves() == 8 (no empty compute-only sink rounds)
```
Drive with [drive](@ref dpf::protocol::drive) or party
`util::drive_composed` / `util::drive_composed_trio`, optionally with
`drive_options` (`from_exchange_wave`, `compact_sink`, `beavers`).
Express/Sabre audits use `util::schedule_fused_audit` (or `fss_point_fused`).
The same plan lowers onto the RoundSink batch loop with
`plan_to_schedule` → [`schedule_session`](@ref dpf::protocol::schedule_session)
→ `finish_schedule`, or the one-shot `drive_via_schedule`. Each exchange wave
becomes a [`schedule_round`](@ref dpf::protocol::schedule_round) whose
`produce` runs on this thread when that instance's peer slot is ready (the
`schedule_session::drive` scan). Rounds carry an
[`edge_id`](@ref dpf/net/edge_mesh.hpp) on an [`edge_mesh`](@ref dpf/net/edge_mesh.hpp)
(star / clique / dealer), a [`receive_rule`](@ref dpf::protocol::receive_rule)
(`domain_open`, `copy_peer`, `field_sum`, `any_two`, `verify_*`, `eq_check`),
optional `branch` (skip send) and `next` (jump after an open).
[`session_host`](@ref dpf/session_host.hpp) queues micro-plans on a durable mesh.
Pad graphs live in [pad_graphs.hpp](@ref dpf/pad_graphs.hpp)
(`du_atallah_mul_graph`, `star_upload_answer_graph`); IKNP setup in
[iknp_graphs.hpp](@ref dpf/iknp_graphs.hpp); PIRsona / hushmap builders in
[mesh_apps.hpp](@ref dpf/mesh_apps.hpp). Named application plans and
`drive_star` / `submit_and_drive` live in [app_plans.hpp](@ref dpf/app_plans.hpp).
Replayable seeds and paper cost CSVs: see [Experiments](@ref experiment_costs).
In short, `dpf::experiment` / `app::measure_plan` / `app::run_measured` attach a
per-thread master seed (default random, or `replay`) and emit CSV breakdowns.
`plan::rounds()` equals `exchange_waves()` and `slot_bytes_all().size()` —
the sink and the scheduler share one count. Trailing compute-only DAG waves
still appear in `waves()` / `wave(i)` but do not allocate RoundSink rounds.
Composer-owned ABY sessions default to
`beavers::schedule_objective::rounds` so online sign×polynomial stays
one round. Dealer benches that want Appendix-E peels keep `prep`.
Independent circuits use `aby_lane(i)`. Live δ openings use
`drive_options::beavers` (`util::u64_beaver_host` wraps
`party_batch_stepper`).
## Hand-schedule shapes the API covers
| Shape | Call | What it saves |
| --- | --- | --- |
| L‖R PRG stretch | `expand_pair` / default `level_walk` | Half the AES vs separate child expands |
| Shared-seed fan (Grotto LUT) | `fan` of `fss_cmp` on one seed | One expand per level, not N×depth |
| Several piecewise LUTs | `schedule_lut_union` | One `fss_cmp`; the union's prefix walk is local |
| Express / Sabre audit | `fss_point_fused` / `schedule_fused_audit` | Sketch in last CW — no +1 exchange wave |
| BGI Remark 3.4 early-stop | `fss_point_early_stop` | Drop ν interactive CW rounds |
| Poplar / idpf prefix checkpoints | `level_walk_prefixes` | Prefix share after each CW, no extra rounds |
| Adaptive idpf_agg | `step` → drive tail → `retain` → `step`… | One packed L‖R open per depth; child bit never on the wire |
| DCF `block_width` | `level_walk_sized` | Per-level CW slot bytes |
| Doerner–Shelat keygen | `level_walk_ds` / `level_walk_ds_sized` | blind‖CW‖advice‖AND + OH AND-layers |
| Keyword PIR / PSI buckets | `multipoint_fan` / `exchange_pack` | Bucket CWs pack; answers one open |
| RSS mul + neighbor refresh | `rss_product_replicated` / `rss_from_y` | One y-exchange, not a reconstructing open |
| FSS leaf → ABY scale | `aby_product` after `fss_point` | Leaf stays on the beaver barrier critical path |
| Multi-lane ABY | `aby_lane(i)` / `aby_product(..., lane)` | Independent barriers, shared wave |
| Round-aware Beaver | `composer::aby<Ring>()` | `schedule_objective::rounds` |
| Prepaid expand / rotate | `defer_expand` / `rotate_share` | Zero online FSS rounds (SUBLEQ offline) |
| Duoram leaf_later | `leaf_later_walk` / `apply_leaf_correction` | Path CWs only; leaf apply is local |
| RSS column, hidden order | `shuffle_hidden` | Three `shuffle_send` waves. `aux` is the left-out party: 2, then 0, then 1 |
Adaptive prefixes: `step_adaptive_prefix` schedules only the next packed
open; drive with `from_exchange_wave = exchanges_flushed` and a sink sized
by `plan::slot_bytes_from` (`compact_sink = true`); then
`retain_adaptive_prefix` adds one local tip. Never unroll a full-depth
adaptive walk into one static schedule.
3PC RSS: `rss_from_y` schedules a `domain::y` exchange (neighbor receive).
`util::drive_composed_trio` splits peer maps per exchange — `y` on the
neighbor ring, `dealer_deliver` from p2, everything else on p0↔p1.
Cross-party `reshare` refuses a reconstructing open; use
`reshare_with_mask`. Opens use domain algebra (`a`/`fss` sum, `b`
subtractive, `y` copy, optional `field_open::fp61`). Authenticated Beaver
sessions size barriers as `auth_opening` and drive through
`u64_auth_beaver_host`. `exchange_fuse` packs extra payloads into one open.
`schedule_cuckoo_probes` turns occupied bucket ids into `multipoint_fan`.
## Client and servers {#compose_client}
Keyword PIR and both three-server PIRs are a star, not a correction-word
walk. `client_servers(servers, query_bytes, answer_bytes)` records two
waves and nothing between the servers:
1. Upload. The slot is `servers * query_bytes` (every key in one round).
2. Answer. The slot is `servers * answer_bytes`. It depends on the upload,
so it cannot share that wave.
`drive` copies the peer's message into the exchange node. It does not add
the shares. A two-party sink stands in for one client and one server; the
slot width is still the full parallel payload, which is what
`dpf::app::exercise` reports as `bytes`.
```cpp
auto q = c.client_servers(/*servers=*/2, /*query_bytes=*/256, /*answer_bytes=*/4);
auto p = c.default_plan();
// p.rounds() == 2
// p.slot_bytes(0) == 512, p.slot_bytes(1) == 8
```
## Parking and the fleet {#compose_fleet}
`drive` used to spin until the peer flushed, then throw. A step that
simply takes a long time looks like that failure, and a pool that always
resumes the side already waiting on a receive never runs the peer who
could unblock it.
`drive_options::park_if_waiting` returns instead. `drive_cursor` remembers
the wave, the exchange index, and whether the submit already happened.
The next `drive` with that cursor receives if the peer has caught up, or
parks again. `one_exchange` stops after a single completed round so a
scheduler can interleave other instances.
`dpf::app::run_fleet(composer, instances, chaos_seed)` is that scheduler.
Each instance is a pair of parties on an in-process sink. Workers prefer
a side that still has a submit to do over a side parked on a receive, and
among those they prefer the one further behind. Delays are per side and
per step. They do not line up, which is the case that makes round-robin
and "always resume the waiter" stall.
`dpf::app::run(name, composer, expect_rounds)` is the one-shot experiment
used by `examples/applications/`. It drives both parties and prints
```
name rounds=R bytes=B
```
`bytes` is the sum of one lane's exchange slots.
Opcodes from `beaver_delta` (300) through `beaver_delta + 1023` are δ
barriers. `user_base` (1000) sits inside that window, so a hand-rolled
kernel opcode in that range is skipped as a beaver. Use 5000 and up for
kernels `drive` must run or reject.
## What eval_full reuses
`eval_full(key, buffer)` keeps a thread-local workspace for that key type.
The same key evaluated again does not rebuild the interior. A different
key still does. Leaves of one block are stretched eight at a time
(`eval_x8`). Pass your own memoizer when two evaluations on one thread
must not share that cache.
DS OH: `level_walk_ds(..., oh=true)` emits
`net::ds_oh_exchanges_per_level` (80) AND-layer opens per level. Size sinks
with `net::compose_ds_slot_bytes` (schedule-exact) or the conservative
`net::ds_walk_slot_bytes` for hand `dist_ds` paths.
\include{cpp} protocol/compose_schedule.cpp
## Gap analysis {#compose_gaps}
| Item | Status |
| --- | --- |
| Point-key stretch | Builtin `fss_expand_pair` runs `prg::aes128`; `fss_step+1` XORs the opened correction onto the children |
| Setup plus walk | `iknp_setup_graph` / `plan_with_pad_setup` — pad frames are `schedule_round`s |
| Edge mesh | `edge_mesh` + `make_memory_star` / `make_memory_clique`; `drive_via_schedule(plan, mesh, …)` |
| Receive rules | `field_sum`, `any_two`, `verify_sketch` / `verify_proof`, `eq_check` in `apply_peer_slot` |
| Jump / host | `schedule_round::next` + `session_host` for SUBLEQ / idpf_agg / hushmap ops |
| PIRsona / hushmap | `pirsona_bitmore_fetch`, `pirsona_gd_update_graph`, `hushmap_add_schedule` |
| 2PC → RSS | `reshare` casts to a `y` share (p2 holds 0) and `rss_from_y`. `reshare_fresh` adds a dealer-sampled zero mask |
| Auth openings | `auth_beaver_host<Ring>` / `u64_auth_beaver_host`. `party_auth_batch_stepper::apply_peer` rejects a bad tag |
| One DS sink | `ds_walk_slot_bytes` is the envelope of the hand walk and `compose_ds_slot_bytes` (5 peer opens per level + OH + mux) |
| Wide opens | Every multiple of 8 bytes is a lane add/sub (`field_open::fp61` for the Mersenne field) |
| Fuse offset | `exchange_fuse` stores the segment offset in `aux` (`uint32_t`) |
Still outside this scheduler: the cuckoo PRP, OPRF evaluation, the full IKNP wire body inside each pad `produce` (frames and count are in the schedule; `iknp::sample` still owns the crypto), and key-specific control-bit / verifiable / Doerner–Shelat advice logic. Those last ones override the builtin kernels. Trio multi-edge `drive_via_schedule` still needs `edge_sinks` wired from the party helper.
### Closed earlier
Incremental drive, domain-correct `a`/`b`/`y` opens, split y/2PC peer maps, OH AND-layers, `level_walk_ds_sized`, `default_plan`, staged adaptive retain, multipoint fan, multi-lane ABY, `defer_expand` / `leaf_later_walk`, Express trailer fuse.
**Go deeper:** [compose.hpp](@ref dpf/compose.hpp),
[beaver.hpp](@ref dpf/beaver.hpp),
[sink_exchange.hpp](@ref dpf/net/sink_exchange.hpp),
[Application sketches](@ref applications),
[guided tour](@ref tour_party).

View file

@ -1,5 +1,9 @@
# Dealer-free keygen {#dealer_free}
\htmlonly
<div class="eli5"><b>ELI5.</b> The index is already split. Doerner–Shelat turns those shares into a reusable key by opening one masked correction per level. geneval stops after a single public query and returns answer shares, not a key. IKNP is how the pads are sampled when no dealer supplies them.</div>
\endhtmlonly
The two parties already hold shares of the secret index.
Nobody sends `alpha` to a dealer.

View file

@ -75,6 +75,10 @@ mockup [I-DPF max and k-th](@ref app_idpf_agg).
## Assigning a wildcard leaf {#wildcard_assign}
\htmlonly
<div class="eli5"><b>ELI5.</b> The blank leaf is a Beaver slot from keygen. Filling it in rewrites the correction word: the parties exchange one blinded share of the new payload and both apply the same patch. assign_cmp rewrites the n comparison words locally and sends nothing. An updatable leaf is the same patch later, O(λ) and independent of the depth.</div>
\endhtmlonly
An output wildcard is a placeholder for a payload filled after keygen.
The type and the `dpf::wildcards` names are on
[Output types](@ref output_types). Evaluation of an unassigned slot throws
@ -149,6 +153,10 @@ either party's type accepts both parties. Name that type with
## Memoizers {#memoizers}
\htmlonly
<div class="eli5"><b>ELI5.</b> The first walk stores interior nodes. The next query starts from the deepest stored node that still lies on its path, instead of from the root. An interval memoizer stores the nodes that cover a range; a sequence memoizer stores the nodes along a sorted list.</div>
\endhtmlonly
## Path memoizers {#path_memoizers}
`eval_point` walks one root-to-leaf path. `make_basic_path_memoizer<Key>()`
@ -331,6 +339,10 @@ size both for that domain.
## Deferred input evaluation {#defer_eval}
\htmlonly
<div class="eli5"><b>ELI5.</b> The PRG expand runs once, into a full-domain buffer, before the index is known. When the offset opens, get() rotates that buffer. The tree is not expanded again.</div>
\endhtmlonly
Additive input blinding evaluates in tree coordinates `x ↦ x + δ`, where
`δ` is reconstructed by `assign_wildcard_input` into `offset_x`. Eager
`eval_interval` folds that map into the traversed range and throws if the
@ -364,6 +376,10 @@ Buffers must outlive both.
## dpf::eval_inner_product {#eval_inner_product}
\htmlonly
<div class="eli5"><b>ELI5.</b> The walk is the same as an interval or full-domain eval, but each leaf is multiplied by a public weight and added into one accumulator. The expanded vector is not stored. A sequence inner product does that only at the listed points.</div>
\endhtmlonly
`eval_inner_product` multiply-accumulates DPF shares against another vector
during the walk. It does not write the output vector.
@ -469,6 +485,10 @@ on that list: at most `O(n m)` expands and `m` output slots.
## Buffered PRG {#buffered_prg}
\htmlonly
<div class="eli5"><b>ELI5.</b> One AES expand produces more blocks than a single tree node needs. The buffered PRG keeps the leftover blocks and serves the next nodes from them, so a wide walk makes fewer expands. The keys do not change.</div>
\endhtmlonly
`dpf::randomness::buffered_prg<PRG, Ts...>` (alias
`dpf::randomness::aes_buffered_prg<Ts...>`) is a forward cursor with one
PRG stream per value type. `get<I>()` and `fill<I>(out, n)` consume the
@ -496,6 +516,10 @@ element. `fill_values` / `fill_masks` of `q` elements are `Θ(q)`.
## Three-party (2,3) DPF {#dpf3}
\htmlonly
<div class="eli5"><b>ELI5.</b> Each evaluator key is two VDPF+ spines. eval_point walks both, Θ(n) expands, then scales into fp61. Opening two or three as_share values is a constant amount of field arithmetic. An updatable rewrite patches four leaves and refreshes an offset, independent of n.</div>
\endhtmlonly
`make_dpf3(α, β)` builds three evaluator keys after Guy Zyskind, Avishay Yanai, and Alex "Sandy" Pentland, [ePrint 2024/1658](@ref bib_dpf3), Figure 3:
two VDPF+ spines plus Shamir embedding in `fp61`. `eval_point` returns a
field share; open with `dpf::reconstruct` on any two (or all three)
@ -525,6 +549,10 @@ two-party comparison, interval, or cuckoo packing, on top of those spines.
## Information-theoretic 3-server DPF {#it_dpf3}
\htmlonly
<div class="eli5"><b>ELI5.</b> There is no PRG tree. For this domain each key is an additive share of a 256-word table. The three shares sum to beta at alpha and to zero elsewhere. A PIR answer is three inner products with the database; those three dots sum to the record.</div>
\endhtmlonly
`make_it_dpf3(α, β)` ([ePrint 2023/028](@ref bib_itdpf)) is a different object from
`make_dpf3`. Each of three parties holds an additive share of the
characteristic vector on `{0..255}`; the **sum** of all three
@ -573,7 +601,13 @@ claimed speedup over dealer keygen or over Half-Tree §5.2.
<div class="tabbed">
- <b class="tab-title">eval_dpf3_point.cpp</b> \include{cpp} evaluation/eval_dpf3_point.cpp
- <b class="tab-title">eval_dpf3_doerner_shelat.cpp</b> \include{cpp} evaluation/eval_dpf3_doerner_shelat.cpp
- <b class="tab-title">eval_dpf3_cmp_ic.cpp</b> \include{cpp} evaluation/eval_dpf3_cmp_ic.cpp
</div>
\htmlonly
<div class="tldr"><b>TL;DR.</b> eval_point is one path. An interval costs the path plus the length of the range. A full domain costs the size of the domain; an inner product does that walk and keeps only the accumulator. Wildcard assign and an updatable rewrite patch the leaf and do not depend on the depth. Three-party eval is two walks and a short field open. The information-theoretic key is a 256-word share, not a walk.</div>
\endhtmlonly

192
doc/pages/experiment.md Normal file
View file

@ -0,0 +1,192 @@
# Logging, statistics, and experiments {#experiment_costs}
Paper and production runs need three things the bare key API does not:
**leveled run logs**, **granular cost statistics**, and **replayable coins**.
All three live in the same measurement stack —
[`dpf/log.hpp`](@ref dpf/log.hpp), [`dpf/run_log.hpp`](@ref dpf/run_log.hpp),
[`dpf::experiment`](@ref dpf/experiment.hpp), and the
[`prg::count`](@ref dpf/prg_count.hpp) / [`thread_work`](@ref dpf/thread_work.hpp)
counters.
Uninstrumented code stays quiet and keeps reading `/dev/urandom`. Logging and
experiments are opt-in. `run_measured`, the battery, and `party_node` call
[`app::start_logging`](@ref dpf/run_log.hpp) and install an experiment covering
the thread that constructs it. `run_parties`, `measure_plan`, and the battery
give party `i` its own stream, `ex.derive_party(i)`, whose master is SHA-256 of
the experiment's master and `i`, and they use it for every trial, warmup
included. A kernel handed to a compute pool draws from its party's stream.
Replaying the master replays every party, as long as each party makes the same
draws in the same order. The party masters are noted as `p0/master`,
`p1/master`, and so on.
## Run log {#run_log}
[`dpf/log.hpp`](@ref dpf/log.hpp) writes one `key=value` line per event to
stderr, syslog, or a file. Nothing is written until `log::configure` runs
(normally through `app::start_logging`), so library code can log freely and a
test that never configures the log stays quiet. Every line starts with
`ts`, `lvl`, `inv` (invocation id), `pid`, `tid`, then `role=<party>` when the
thread has one, then `ev=<event>` and the event's fields. Seed bytes go through
`record::seed` as hex, as a SHA-256 fingerprint, or not at all.
[`app::start_logging`](@ref dpf/run_log.hpp) turns it on and writes the
provenance banner first, at `info`:
| Event | Contents |
| --- | --- |
| `ev=start` | program, cwd, UTC/local start, invocation id shared with `runs.csv` |
| `ev=build` | git rev (`LIBDPF_GIT_REV`), compiler, `NDEBUG`, ISA, sanitizers |
| `ev=host` | hostname, kernel, CPU model, affinity, governor, turbo, memory, load; `ev=if` per interface |
| `ev=argv` | shell-quoted command line |
| `ev=env` | `DPF_*` / `LIBDPF_*` and perf-relevant vars (`LD_PRELOAD`, …) |
| `ev=config` | every `run_config` setting |
After that come seeds with their source, listeners and links (endpoints, peer
authentication, encryption, applied socket options, kernel RTT), each party's
plan and costs, trial statistics, and failures. `ev=end` at exit gives elapsed
time. The settings are `run_config` keys:
```
DPF_LOG_LEVEL=debug DPF_LOG=file:/tmp/run.log,stderr DPF_LOG_SEEDS=hash
--log_level=debug --log=syslog --log_seeds=off
```
Levels are `silent`, `error`, `warning`, `info` (the default), `debug`, and
`trace`. With `log_seeds=full` the log holds every master, so anyone who has
the file can regenerate that party's randomness. `hash` prints a SHA-256
fingerprint instead.
## Replayable seeds
```cpp
dpf::experiment ex("keyword_pir"); // fresh 32-byte master
auto master = ex.seed(); // record for the paper
auto x = dpf::uniform_sample<std::uint64_t>();
auto again = dpf::experiment::replay("keyword_pir", master);
assert(dpf::uniform_sample<std::uint64_t>() == x);
```
While the experiment is installed, every `uniform_fill` / `uniform_sample`
draw on that thread (DPF roots, Beaver default sampler, pads, Shamir, IKNP,
field rejection) comes from AES-CTR of the master. The master itself is always
drawn from OS entropy, even when nested under another experiment.
Constructors that already own a seed note it automatically:
| Type | Seed name |
| --- | --- |
| (master) | `master` |
| `beavers::oracle` | `beavers::oracle` (+ `lane_table`) |
| `buffered_prg` | `buffered_prg` |
| `prg_pad_rng` | `prg_pad_rng` |
| `pseudorandom_root_sampler` | `pseudorandom_root_sampler` |
Call `ex.note_seed("label", value)` for anything else. Replay is
**order-sensitive**: the same party must make the same `uniform_sample` calls.
Indexed `beavers::oracle(seed)` stays the seekable Beaver path; its seed still
appears in the report when constructed under the experiment.
## Measuring a compose plan {#statistics}
```cpp
auto plan = dpf::protocol::fss_point_plan(0);
auto ex = dpf::app::measure_plan("fss_point", plan);
ex.write_csv("/tmp/run1");
```
Or from an application sketch:
```cpp
// Prints rounds, live bytes, wall/CPU, PRG evals, random bytes, seed hex.
// Set DPF_EXPERIMENT_DIR=/tmp/run1 to also emit CSVs.
dpf::app::run_measured("keyword_pir",
dpf::protocol::keyword_pir_compose_plan(0, depth), 2);
```
What is recorded:
| Metric | Source |
| --- | --- |
| Interactive rounds / DAG depth | `plan::exchange_waves()` / `plan::waves()` |
| Critical path | back-walk of max-`wave_of` inputs |
| Schedule bytes per edge | `slot_bytes` + `wave_channel` |
| Bytes out/in per round | the round's send and receive slots, on its own channel |
| Wall / CPU | `steady_clock` / party 0's thread CPU plus its pool kernels' CPU |
| Symmetric-key blocks | `prg::count(purpose, primitive)` per thread |
| Random bytes | `uniform_fill` TLS counter |
| Seeds | master + every `note_experiment_seed` |
Symmetric-key blocks are counted by purpose and primitive. The purposes are
`expand` (tree expansion, leaf conversion, label and column PRGs), `hash`
(IKNP row hashes, garbled-gate hashes), and `harness` (the experiment's own
seed stream). The primitives are AES-128, AES-256, ChaCha, and LowMC.
`prg_evals` is `expand` plus `hash` over every primitive. Harness blocks are
never part of it. The counters are per thread; a kernel a party hands to a
compute pool (`compute_threads > 0`) is charged to that party when it returns
([`dpf/thread_work.hpp`](@ref dpf/thread_work.hpp)).
## CSV layout
`write_csv(dir)` appends. Every row ends with the id of the process that
wrote it (`invocation`), so rows from different runs into one directory stay
apart. A file whose header differs from this layout is renamed to
`<name>.before-<UTC>.csv` before new rows go in.
- `summary.csv` — one row per run, including `master_seed` hex. `wall_ns`,
`cpu_ns`, and the counters come from the last (instrumented) trial;
`median_ns` and `slowest_median_ns` are medians over the timed trials for
party 0 and for the slowest party.
- `rounds.csv` — per-round deltas and edge name
- `edges.csv` — protocol totals per peer / rss_next / dealer
- `seeds.csv` — every noted seed as hex
- `critical_path.csv` — node id, wave, opcode, effect
- `config.csv` — every `run_config` setting
- `trials.csv` — every party's wall time for each timed trial
- `wire.csv` — party 0's link counters, headers included
- `sym.csv` — symmetric-key blocks by purpose and primitive
- `runs.csv` — the invocation id that ties rows to the run log, and whether
the master was fresh, provided, or derived
Before any party starts its clock, all parties wait at a start gate until
every one has finished setup.
## The battery
The harness is `party_bench` and `profile_party`. Both spawn `p0`, `p1`, and
`p2` and dial [`net::trio`](@ref dpf/net/trio.hpp), the library's localhost
mesh (`p0-p1`, `p0-p2`, `p1-p2`). A flow's bytes and rounds are the frames on
that mesh. The repeat barrier is not included.
```
party_bench --tag bench --repeat 3 --warmup 1
party_bench --case arith_mul_p11 --repeat 5
profile_party --suite gadget --repeat 3 --warmup 1
```
`--tag bench` is the default list. The gadget rows are tagged `bench` as well,
so they sit on that list. `profile_party --suite gadget` runs only those rows.
`profile_party --suite all` appends them after the core and extreme suites.
A secret index stays a DPF. The gadget rows time what you do after a key, or
on shares the parties already hold:
- `arith_proj_*`, `arith_mul_*`, `arith_thresh_*`, `arith_chain_mul4` —
Ball–Malkin–Rosulek word garbling. Party 0 sends the evaluator view on the
p0–p1 link. Party 1 evaluates that view and opens.
- `yao_if_*`, `yao_onehot_*` — stacked and one-hot garbling on the garbler.
The active branch's tables and the evaluator's labels go across the p0–p1
channel, and party 1 evaluates those bytes. Base OT stays on the `iknp` tag.
- `flute_d*` — FLUTE. Party 2 deals the mask shares. Parties 0 and 1 exchange
the online bits and open.
- `shuffle_n*` — three hidden-shuffle passes. Each pass's array is sent on
the ring and consumed as the next party's inbound. Party 0 opens the sum.
Sizes are the modulus, the branch width, the table width, and the column
length. The manual pages are [a small word](@ref arith_garble_word),
[a public table](@ref flute_lut), [a secret branch](@ref yao_stack), and
[an array of shares](@ref share_shuffle).
See also [Network, parties, and MPC](@ref network_and_mpc),
[Protocol composition](@ref protocol_compose), and the
[application mockups](@ref applications).

View file

@ -19,7 +19,12 @@ A *distributed point function* (DPF) is a way to share it with short keys.
`libdpf++` builds those keys and evaluates them quickly in C++17.
People use DPFs for private lookup (PIR), multi-party computation (MPC),
and other privacy tools. See also the [ideal functionalities](@ref ideal_functionalities)
and other privacy tools. This tree also ships the **network and MPC stack**
those protocols need: TLS party sessions, RoundSink rounds, Beaver / Yao /
arithmetic shares on leaf values, composition, leveled run logs, and
paper-cost statistics. That map is [Network, parties, and MPC](@ref network_and_mpc);
logging and CSVs are [Logging, statistics, and experiments](@ref experiment_costs).
See also the [ideal functionalities](@ref ideal_functionalities)
for what each protocol is allowed to learn.
## Prior work {#tour_prior}
@ -87,8 +92,8 @@ auto [k0, k1] = dpf::make_dpf(X{1000}, std::uint64_t{1});
Width literals sit next to those types: `100_u12` (`modint`), `7_x12`
(`xint`), `dpf::literals::operator""_bitstring`, and `1.5_fixed16` through
`_fixed64`. `dpf::bit`, `dpf::twobit`, and `dpf::nyble` are outputs, not
domains. See [Input types](@ref input_types).
`_fixed64`. `dpf::bit`, `dpf::twobit`, `dpf::nyble`, and `dpf::gf2` through
`dpf::gf264` are outputs, not domains. See [Input types](@ref input_types).
## Outputs: what sits at that point {#tour_outputs}
@ -102,6 +107,7 @@ Many outputs can share one leaf when they fit.
| `dpf::bit` | One XOR bit, packed | [bit.hpp](@ref dpf/bit.hpp) |
| `dpf::twobit` | Z/4Z, packed 2-bit lanes | [twobit.hpp](@ref dpf/twobit.hpp) |
| `dpf::nyble` | Z/16Z, packed nibbles | [nyble.hpp](@ref dpf/nyble.hpp) |
| `gf2` … `gf264` | GF(2^k), XOR add, field multiply | [gf2.hpp](@ref dpf/gf2.hpp) |
| `dpf::bitstring` | XOR string | [bitstring.hpp](@ref dpf/bitstring.hpp) |
| `grotto::fixedpoint` | Fixed-point raw word | [fixedpoint.hpp](@ref grotto/fixedpoint.hpp) |
| `dpf::vec<T, N>` | `N` lanes, no carry between them | [vec.hpp](@ref dpf/vec.hpp) |
@ -109,7 +115,7 @@ Many outputs can share one leaf when they fit.
| `field64` / `field128` | Prime fields | [field64.hpp](@ref dpf/field64.hpp) |
| `fp61` | Field for 3-party DPFs | [fp61.hpp](@ref dpf/fp61.hpp) |
| `p256` / `p256_scalar` | Curve and order | [p256.hpp](@ref dpf/p256.hpp) |
| Typed shares | (2,2) additive and subtractive, (3,3) additive, (2,3) replicated | [secret_share.hpp](@ref dpf/secret_share.hpp) |
| Typed shares | (2,2) additive and subtractive, (3,3) additive, (2,3) replicated, (K,N) Shamir | [secret_share.hpp](@ref dpf/secret_share.hpp), [shamir.hpp](@ref dpf/shamir.hpp) |
```cpp
auto [k0, k1] = dpf::make_dpf(
@ -118,6 +124,12 @@ auto [k0, k1] = dpf::make_dpf(
// assign the payload later; eval before assign throws
```
Shamir shares are `shamir::share<T, Party, K, N>`. Any `K` of `N` open the
constant term. `(2,3)` is `shamir_share`. The fields are `fp61` and `gf2n`.
For `gf2n`, `N` must be less than `2^k`. `make_dpf3` uses the `(2,3)` case
on `fp61`. `examples/mwe/shamir.cpp` deals a `(3,5)` secret in `fp61` and a
`(2,3)` secret in `gf28`.
Packed output literals are `1_bit`, `2_twobit`, and `10_nyble`.
`dpf::vec<T, N>` is `N` lanes of an ordinary output with no carry between
lanes; the construction is on [Output types](@ref output_types).
@ -216,6 +228,10 @@ or zip two parties' buffers.
## Comparisons and ranges {#tour_dcf}
\htmlonly
<div class="eli5"><b>ELI5.</b> A comparison key is not a single spike. It returns the true payload on one side of the secret point and the false payload on the other, and the shares add instead of subtract. An interval key packs the two endpoint comparisons into one key. idcf repeats a correction at every depth; cmp_prefix stops after L bits.</div>
\endhtmlonly
A *distributed comparison function* (DCF) returns a payload when a predicate
holds on the secret point. The four predicates are `dpf::lt`, `dpf::leq`,
`dpf::gt`, and `dpf::geq`. Each takes the true payload and an optional false
@ -284,6 +300,10 @@ ideal figures [F_DCF](@ref dcf.hpp), [F_BDCF](@ref blocked_dcf.hpp), [F_IC](@ref
## Verifiable and extractable keys {#tour_vdpf}
\htmlonly
<div class="eli5"><b>ELI5.</b> verifiable carries an extra seed on each correction word and folds it into one proof token. Equal tokens across parties mean the seeds were the honest ones. extractable is a separate weight-1 sketch in fp61: any second hot point fails it. output_mac authenticates the opened leaf, not the path.</div>
\endhtmlonly
Pass `dpf::verifiable{}` or `dpf::extractable{}` as an extra `make_dpf`
argument. `verifiable` follows de Castro and Polychroniadou, EUROCRYPT
2022 ([ePrint 2021/580](@ref bib_vdpf)): one extra correction seed per level (their hash
@ -308,6 +328,10 @@ auto [e0, e1] = dpf::make_dpf(std::uint8_t{42}, std::uint64_t{7},
## Many points at once {#tour_multipoint}
\htmlonly
<div class="eli5"><b>ELI5.</b> The m secret points are placed in cuckoo buckets with three hashes, one ordinary point key per bucket, so the key grows with m and not with the domain. Evaluation probes the three buckets that could hold the query. One verifiable tag is a single proof for the whole set.</div>
\endhtmlonly
`make_multipoint(alphas, betas)` packs many points into cuckoo buckets,
following de Castro and Polychroniadou, EUROCRYPT 2022, §4 (ePrint
2021/580): `κ = 3` hashes, one point key per bucket.
@ -326,6 +350,10 @@ auto [k0, k1] = dpf::make_multipoint(alphas, betas);
## A vector with one programmable coordinate {#tour_ppvc}
\htmlonly
<div class="eli5"><b>ELI5.</b> Commit binds both roots of each aligned 1-bit DPF pair under a Naor string, before anyone chooses the coordinate. Open reveals one side of each pair, which writes that coordinate or the sum of the vector onto a public index. verify checks the opened roots against the committed strings.</div>
\endhtmlonly
`dpf::ppvc` commits to a vector in `(Z/2^s Z)^n` before the hidden
coordinate is chosen. The commitment binds both roots of `s` aligned
1-bit DPF pairs. Opening one side of each pair writes that coordinate,
@ -343,6 +371,10 @@ const auto op = scheme::open(st, 0, 0x5a, std::uint8_t{40});
## Tree shapes: classic and Half-Tree {#tour_trees}
\htmlonly
<div class="eli5"><b>ELI5.</b> The default key is the CCS 2016 layout: one correction word per level, with the last few levels packed into the leaf when the output group is small. Half-Tree keeps that key shape and changes the expand to H(s) and H(s) XOR s, which is about half as many permutation calls. Section 5.2 of that paper is a different thing: two-party keygen in the COT/OLE hybrid, which this generator does not use.</div>
\endhtmlonly
The default interior PRG walks a Boyle–Gilboa–Ishai tree (CCS 2016,
full version [ePrint 2018/707](@ref bib_fss2018)).
Select `prg::aes128_ccr` as the *interior* PRG to use Half-Tree expands
@ -358,6 +390,10 @@ about `4n`, and `1.5N` calls for a full-domain evaluation versus `2N`.
## Two-party keygen without a dealer: Doerner–Shelat {#tour_ds}
\htmlonly
<div class="eli5"><b>ELI5.</b> Each party holds a share of the index and neither sends the index. Level by level they open a masked correction word. The key that comes out is the same object a dealer would have built with make_dpf. geneval uses the same opening on one public query and does not return a reusable key.</div>
\endhtmlonly
Two parties hold XOR (or additive) shares of `alpha`.
A third party deals pads and learns nothing.
The result matches what an honest `make_dpf` would emit.
@ -410,6 +446,10 @@ a round of its own beyond that split.
### Two-party socket walk without p2 (IKNP) {#tour_iknp}
\htmlonly
<div class="eli5"><b>ELI5.</b> IKNP turns a few slow base oblivious transfers into a long tape of correlated pads. Those pads stand in for the dealer in the Doerner–Shelat walk. The base OTs are Chou–Orlandi on P-256. When a level must stay hidden, the hash is the Boyar–Peralta AES S-box: 32 ANDs per block, which is why an oblivious tape is much longer than a reveal tape.</div>
\endhtmlonly
When there is no pad dealer, p0 and p1 sample the same Doerner–Shelat
pad correlations with semi-honest OT extension after Yuval Ishai, Joe
Kilian, Kobbi Nissim, and Erez Petrank, CRYPTO 2003 (`dpf::iknp::sample`),
@ -482,6 +522,10 @@ party mesh [trio.hpp](@ref dpf/net/trio.hpp).
## Multiplication and circuits: Beaver {#tour_beaver}
\htmlonly
<div class="eli5"><b>ELI5.</b> Preprocessing gives every wire a blind. To multiply, the parties open the two inputs masked by those blinds, then fix the product with a local correction. A later gate reuses blinds it already holds and only samples new monomials. A MAC is a second share of the same width, checked in batch.</div>
\endhtmlonly
ABY2.0-style sessions open masked wires once, following Patra, Schneider,
Suresh, and Yalame, USENIX Security 2021 (full version [ePrint 2020/1225](@ref bib_aby2)).
A fresh triple follows Donald Beaver, [CRYPTO 1991](@ref bib_beaver), which reconstructs
@ -497,12 +541,75 @@ A later round reuses blinds it already holds and samples only the new
monomials. The dealer keeps each full blind; the parties receive the
additive splits.
A word that has already left the key can be added and projected in one shot
([a small word](@ref arith_garble_word)). A public table on a short masked
index is [FLUTE](@ref flute_lut). A secret point in a public table stays a
DPF. The session above is still the interactive product.
**Go deeper:** [beaver.hpp](@ref dpf/beaver.hpp),
[F_Beaver](@ref beaver.hpp), [F_BeaverAuth](@ref beaver.hpp),
[constrained_cmp.hpp](@ref dpf/constrained_cmp.hpp) for `F_CCMP`.
## A column of shares {#tour_shuffle}
The secret position in this library is a key. One cell of a shared array is
a unit DPF, a public rotate, and a dot product
([Duoram](@ref app_duoram)).
A hidden shuffle is what you do when you already hold every row and you are
about to open the column. Three passes, from the pairwise seeds `k01`,
`k12`, and `k20`, leave one party out of each permutation. The opened order
is not the stored order, and no single party can recompute it.
`shuffle_hidden_pass` is one party's step. `shuffle_party` is the other
helper: one permutation from `k01`, which every holder of that seed can
recompute.
The shuffle does not look up a cell, and it does not sort. Those stay a DPF
and a DCF.
**Go deeper:** [An array of shares](@ref share_shuffle),
[shuffle.hpp](@ref dpf/shuffle.hpp).
## When the leaf is a circuit {#tour_yao}
\htmlonly
<div class="eli5"><b>ELI5.</b> Eval already gave you a share of the leaf. If the next step is a bit circuit the key does not contain, split that share into XOR bits, garble the circuit, and share the answer back as a leaf.</div>
\endhtmlonly
A point leaf is subtractive, so the split is `b2y` and the return is `y2b`.
A comparison leaf is additive, so the split is `a2y`. An `fss_share` opens
like a point leaf. Parties 0 and 1 garbling a replicated leaf use `rss2y`;
party 2 does not send. The bits are least-significant first. Both shares
are arguments to the conversion. The masked `x - r` that A2B opens is uniform.
The netlist is XOR, AND, XNOR, and NOT. Party 0 garbles with half-gates
([ePrint 2014/756](@ref bib_halfgates)): 32 bytes per AND, one message, XOR
shares out. The block this is for is `yao::aes_mmo`, the same zero-key
Matyas–Meyer–Oseas block `prg::aes128::eval` uses on a tree expand, 5120
ANDs. AES-128 under a shared key is the other packaged netlist, 6400 ANDs.
The Doerner–Shelat oblivious hash still evaluates that S-box as GMW layers
in [the dealer-free section](@ref tour_ds). `aes_mmo` is one of those blocks
in constant rounds, for a leaf or a seed you already hold as bits.
A comparison, an interval, a public-offset polynomial, and a product of
two leaves do not come here. Those are a key, Grotto, or one Beaver open.
`cost_pass` does not grow a Yao strategy for them.
A secret if/else or a menu of blocks on those bits is stacked garbling
([a secret branch](@ref yao_stack)). The transmitted rows follow the heavier
block. The leaf is still split in and shared back out. A comparison stays
on the key.
**Go deeper:** [A boolean function of a leaf](@ref yao_leaf),
[yao.hpp](@ref dpf/yao.hpp), [yao_share.hpp](@ref dpf/yao_share.hpp),
[F_Yao](@ref yao.hpp), [F_YaoShare](@ref yao_share.hpp).
## Three evaluators {#tour_dpf3}
\htmlonly
<div class="eli5"><b>ELI5.</b> make_dpf3 gives each of three parties a pair of two-party keys, and the payload is a Shamir share in fp61, so any two evaluation shares open the value and one share is independent of it. make_it_dpf3 is not that object: each party holds an additive share of a length-256 table, and the three shares sum to the point function.</div>
\endhtmlonly
`(2,3)` point keys follow Zyskind, Yanai, and Pentland, [ePrint 2024/1658](@ref bib_dpf3),
Figure 3: each evaluator key is a pair of `(2,2)`-VDPF+ keys.
Each key is a Shamir share in `fp61`.
@ -554,6 +661,10 @@ three `eval_it_dpf3` values is the point function. See
## Grotto: math after a public offset {#tour_grotto}
\htmlonly
<div class="eli5"><b>ELI5.</b> The parties open the public distance eta = x − r. A polynomial, a binomial jet, a carry, or a table lookup is then a correction of shares they already have. That correction does not walk another DPF.</div>
\endhtmlonly
Open `eta = x - r`. Then cheap public corrections give rich functions of `x`
without another tree walk.
@ -567,7 +678,8 @@ without another tree walk.
| Twisted jets | `c^m \lambda^c`, including dyadic `1/2` |
| Carry | Truncate, arithmetic shift, extend on shared limbs |
| Prefix parity | XOR or signed prefix sums along a key |
| LUTs | Constant, easy, dyadic, range, window, principal |
| LUT union | Several piecewise tables on one comparison and one prefix walk |
| LUTs | Constant, easy, dyadic, range, window, principal, Haar and bior(5,3) |
| Closed form / exact steps | Compositions and digit or bit counts |
| `fixed_mul` | Fixed-point product into a chosen width |
@ -599,8 +711,15 @@ local fixed-width arithmetic (`fixed_mul` uses at most 8 limbs). Carry
keys are one comparison per live recipe flag, on the limb width, plus
one Beaver bit-opening when the recipe multiplies share MSBs. Prefix
parity on `m` endpoints is one resumed path walk, `O(m n)` expands in
the worst case; that walk follows [ePrint 2023/108](@ref bib_grotto). The degree-0 exact
LUTs follow Appendix D of the same paper. Other LUT calls are a knot
the worst case; that walk follows [ePrint 2023/108](@ref bib_grotto).
Several piecewise LUTs share one such comparison:
[make_lut_union_plan](@ref grotto/lut_union.hpp) unions their breakpoints,
and [schedule_lut_union](@ref grotto/lut_union.hpp) is one `fss_cmp` of
`n` rounds. The prefix walk over the union is local.
The degree-0 exact
LUTs follow Appendix D of the same paper. Haar and bior(5,3) tables
compress a uniform grid and evaluate in \f$\Theta(1)\f$ arithmetic
([ePrint 2025/013](@ref bib_wave)). Other LUT calls are a knot
search plus a constant-size Horner; exact steps loop over the word.
Detail is on the Grotto pages.
@ -618,17 +737,55 @@ Ship keys over ASIO peers with `dpf::asio::make_dpf`.
## Running protocols {#tour_party}
The overview of links, composition, leaf MPC, and measurement is
[Network, parties, and MPC](@ref network_and_mpc).
The `party/` programs run a three-role mesh (`p0`, `p1`, dealer `p2`).
Flows cover Beaver, geneval, DCF, DPF3, Grotto, and adversarial checks.
Use `--list` and `--tag` to filter.
Composed protocols record strands on a `dpf::protocol::composer`
([compose.hpp](@ref dpf/compose.hpp)). Values are tagged with a share
domain (`fss`, `a`, `b`, `rss`, `y`). An FSS leaf consumed by an ABY2.0
product gets a local `b2a` (or `fss2a`) inserted by `as`, and the leaf stays
on the beaver barrier's critical path so Duoram / SUBLEQ scale cannot float
before the walk. Like blinds and expansions are interned across sub-strands,
and independent opens share one RoundSink round. Party-count changes use an
explicit `reshare` (`rss_from_y` for y→rss). Composer-owned Beaver sessions use
`beavers::schedule_objective::rounds` so sign×polynomial stays one online
round; dealer benches that want Appendix-E peels keep the default `prep`
objective. Express/Sabre-style audits use `fss_point_fused` /
`level_walk_fused` so the sketch rides in the last CW flush.
BGI early-stop is `fss_point_early_stop`; Poplar checkpoints are
`level_walk_prefixes`; DCF `block_width` sizes are `level_walk_sized`.
Doerner–Shelat is `level_walk_ds` / `level_walk_ds_sized` (OH AND-layers
match `ds_oh_exchanges_per_level`); adaptive idpf_agg is staged
`step_adaptive_prefix` → drive tail (`from_exchange_wave`) →
`retain_adaptive_prefix`; keyword PIR buckets are `multipoint_fan`;
multi-lane ABY is `aby_lane`; prepaid SUBLEQ expands are `defer_expand`.
Party drivers use `util::drive_composed` / `util::drive_composed_trio` on
`composer::default_plan()` with `u64_beaver_host` or
`u64_auth_beaver_host`. A client/server query is `client_servers`
(one upload, one answer). Many instances with uneven stalls are
`dpf::app::run_fleet`: a worker parks instead of spinning and runs
whichever side can still submit. Cross-party reshare is `reshare_with_mask`;
pads are `dealer_deliver`; extra payloads share a round via
`exchange_fuse`; occupied cuckoo buckets are `schedule_cuckoo_probes`.
The full API and what stays outside compose are
[Protocol composition](@ref protocol_compose).
**Go deeper:** [trio.hpp](@ref dpf/net/trio.hpp),
[compose.hpp](@ref dpf/compose.hpp),
`party/registry.hpp` (in-tree).
## Suggested reading order {#tour_order}
1. [First program](@ref basics), then this tour if you want the map in prose.
2. [Capabilities](@ref capabilities): verifiability, programmability, comparisons, multipoint, three servers, dealer-free keygen, Beaver, Grotto.
2. [Capabilities](@ref capabilities) for key features; [Network, parties, and MPC](@ref network_and_mpc) for the runtime; [Logging, statistics, and experiments](@ref experiment_costs) for the run log and CSV costs.
3. [Domains](@ref input_types) and [Payloads](@ref output_types).
4. [Evaluation](@ref evaluation) and the [code examples](@ref listings).
5. [Application sketches](@ref applications).
\htmlonly
<div class="tldr"><b>TL;DR.</b> Point keys are make_dpf: subtract to open, one correction word per level, early-stop packing when the output is small. Comparisons and intervals add instead of subtract. verifiable, extractable, and cuckoo multipoint share one paper (ePrint 2021/580). Dealer-free keygen is the Doerner–Shelat opening, with IKNP pads when nobody deals them. Three parties are either Shamir spines (any two open) or an information-theoretic table (all three add). After a public offset, Grotto corrects polynomials, carries, and tables without another walk. Live runs use the party mesh, composer, Beaver/Yao/arith on leaves, start_logging, and experiment CSVs — see Network &amp; MPC and Logging &amp; statistics.</div>
\endhtmlonly

View file

@ -1,3 +1,5 @@
An *input type* is the domain of the secret index.
Shorter domains make shorter keys and faster walks.
Prefer `std::uint16_t` over `unsigned short` so the depth is obvious.
@ -52,7 +54,8 @@ Three names cover every integer width:
in `GF(2)^N`. See `xor_wrapper` below.
`dpf::bit`, `dpf::twobit`, and `dpf::nyble` are packed output lanes of width
1, 2, and 4. They are not domains. A domain of that width is `modint<1>`,
1, 2, and 4. `dpf::gf2`, `gf22`, and `gf24` are the same widths in GF(2^k).
They are not domains. A domain of that width is `modint<1>`,
`modint<2>`, or `modint<4>` (or the matching `xint`).
**See also**\n
@ -70,7 +73,6 @@ and [Output types](@ref output_types) for the packed lanes.
<details class="type-note">
<summary>Extended-precision integer scalar types</summary>
The extended-precision (`128`-bit) integer scalar types provided as
compiler extensions by most major C++ compilers (e.g., `__int128` and `unsigned __int128`), including `g++` and
`clang++` when compiling for `64`-bit targets. (As these types are not
@ -93,6 +95,9 @@ to declare such 128-bit integers in compiler-independent way.
<details class="type-note">
<summary>dpf::modint&lt;Nbits&gt;</summary>
\htmlonly
<div class="eli5"><b>ELI5.</b> modint&lt;N&gt; is an integer modulo 2^N. The DPF depth is N, not the width of the integer stored underneath. Arithmetic uses that underlying integer and masks on read, so it is as fast as the raw word and still wraps at N bits.</div>
\endhtmlonly
Arbitrary-, yet fixed-bitlength unsigned integer types. `dpf::modint` is a
lightweight class template that adapts one of the above-mentioned integer
@ -127,7 +132,6 @@ Here are some examples of arithmetic operations with `dpf::modint`:
<details class="type-note">
<summary>dpf::bitstring&lt;Nbits&gt;</summary>
Arbitrary-, yet fixed- bitlength binary strings types. `dpf::bitstring` is
a class template that represents a binary string of any given length.
Compared with `dpf::modint`, a `dpf::bitstring` is well suited to cases
@ -162,7 +166,6 @@ shortest length -- and with the fastest evaluations -- possible.
<details class="type-note">
<summary>dpf::keyword&lt;Alphabet, N&gt;</summary>
Fixed-length strings over restricted alphabets. `dpf::keyword` is an alias
for the class template `dpf::basic_fixed_length_string`, which represents
a string of length `N` consisting solely of letters from
@ -175,6 +178,7 @@ internal representation. This produces representations that are
when `alphabet` comprises few elements.
For example
\code{cpp}
const char cstr[] = "7fffae02";
std::cout << (sizeof(cstr) - 1) * CHAR_BIT << "\n"; // prints 64
@ -218,6 +222,9 @@ The `dpf::alphabets` namespace for a catalog of predefined alphabets.
<details class="type-note">
<summary>dpf::xor_wrapper&lt;T&gt;</summary>
\htmlonly
<div class="eli5"><b>ELI5.</b> The group operation is XOR of the underlying word, not addition. A domain of this type walks the bits in GF(2)^N. A leaf of this type opens by XORing the two shares.</div>
\endhtmlonly
An element of `GF(2)^N` for `N=8*sizeof(T)`. `xor_wrapper` is a
lightweight class template that adapts "integer-like" types so that
@ -246,7 +253,6 @@ when the index is an ordinary integer of the same width.
<details class="type-note">
<summary>dpf::keyword2&lt;Pattern&gt;</summary>
A ranked string whose language is a static pattern. `dpf::keyword2` is the
replacement for `dpf::keyword`: the pattern is the type, and each accepted
string has one rank in `0 .. |L|-1`. That rank is the DPF input. The type
@ -290,6 +296,9 @@ auto [k0, k1] = dpf::make_dpf(alpha, std::uint64_t{1});
<details class="type-note">
<summary>grotto::fixedpoint&lt;FractionalBits, IntegralType&gt;</summary>
\htmlonly
<div class="eli5"><b>ELI5.</b> The value is an integer with a public split between integer bits and fraction bits. Leaf addition is integer addition of the raw word. It does not shift the binary point.</div>
\endhtmlonly
A fixed-point value stored in an integer backend. `FractionalBits` is the
number of bits after the binary point. `IntegralType` defaults to
@ -323,6 +332,9 @@ auto [k0, k1] = dpf::make_dpf(half, std::uint64_t{1});
<details class="type-note">
<summary>dpf::wildcard_value&lt;T&gt; as a domain</summary>
\htmlonly
<div class="eli5"><b>ELI5.</b> Keygen plants a random mask instead of the index. The parties later open mask − alpha, not alpha. Until that offset is reconstructed, evaluation throws.</div>
\endhtmlonly
A wildcard input defers the secret index. `make_dpf(dpf::wildcard_value<Input>{}, beta)` plants a random mask in each key's `offset_x`. The public value the parties later open is `mask - alpha`, not `alpha`.
@ -352,7 +364,6 @@ auto [k0, k1] = dpf::make_dpf(dpf::wildcard_value<std::uint8_t>{}, std::uint32_t
<details class="type-note">
<summary>Custom input type requirements</summary>
A type can be a DPF input when the library can walk its bits and, for interval
or full-domain evaluation, order its values.
@ -392,3 +403,7 @@ template <> struct mod_pow_2<input_type> {
}
\endcode
</details>
\htmlonly
<div class="tldr"><b>TL;DR.</b> Shorter domains make shorter keys. Use a fixed-width integer when it fits, modint or xint for any other width up to 256, and bitstring or keyword when the index is not a number. bit, twobit, and nyble are payloads, not domains. A wildcard index is filled in after keygen.</div>
\endhtmlonly

View file

@ -1,5 +1,5 @@
\htmlonly
<p class="hero-lead"><code>libdpf++</code> is a header-only C++17 library of distributed point functions: short keys that hide one secret index, then answer at public points as secret shares.</p>
<p class="hero-lead"><code>libdpf++</code> is a header-only C++17 library of distributed point functions: short keys that hide one secret index, then answer at public points as secret shares. The same tree ships a TLS party mesh, round scheduling, MPC on those leaf shares, leveled run logs, and paper-cost statistics — so readers do not have to dig the API to find them.</p>
<h2 id="features">What it does</h2>
<div class="feature-grid">
<a class="feature-card" href="verifiability.html"><span class="feature-kicker">Proofs</span><strong>Verifiability &amp; authenticity</strong><p>Honest correction seeds, a weight-1 sketch, and MACs on the leaves, on the same walk.</p></a>
@ -8,10 +8,19 @@
<a class="feature-card" href="multipoint_keys.html"><span class="feature-kicker">Many secrets</span><strong>Multipoint keys</strong><p>Pack many secret points into one cuckoo key. One batched proof covers the set.</p></a>
<a class="feature-card" href="multiparty.html"><span class="feature-kicker">Three parties</span><strong>Multiparty &amp; 3-server</strong><p>Any two of three open a Shamir key. Or an information-theoretic three-server DPF.</p></a>
<a class="feature-card" href="dealer_free.html"><span class="feature-kicker">No dealer</span><strong>Dealer-free keygen</strong><p>The parties already share the index. Doerner–Shelat, geneval, and IKNP finish the key.</p></a>
<a class="feature-card" href="beaver_triples.html"><span class="feature-kicker">Multiplication</span><strong>Beaver triples</strong><p>Authenticated products on leaf shares, including the ABY2.0 MAC check.</p></a>
<a class="feature-card" href="jet_and_ring.html"><span class="feature-kicker">After the offset</span><strong>Grotto</strong><p>Jets, polynomials, carry, and exact ring changes once a public offset is open.</p></a>
<a class="feature-card" href="jet_and_ring.html"><span class="feature-kicker">After the offset</span><strong>Grotto</strong><p>Jets, polynomials, carry, and exact ring changes once a public offset is open. Several LUTs share one comparison.</p></a>
<a class="feature-card" href="ppvc_manual.html"><span class="feature-kicker">Commitments</span><strong>Programmable vectors</strong><p>Bind a vector, then open one hidden coordinate or the sum.</p></a>
<a class="feature-card" href="applications.html"><span class="feature-kicker">Protocols</span><strong>Application sketches</strong><p>The DPF step of Duoram, keyword PIR, PSI, Prio, LLAMA, and the rest.</p></a>
<a class="feature-card" href="applications.html"><span class="feature-kicker">Sketches</span><strong>Application sketches</strong><p>The DPF step of Duoram, keyword PIR, PSI, Prio, LLAMA, and the rest.</p></a>
</div>
<h2 id="around-the-keys">Around the keys</h2>
<p class="hero-lead" style="font-size:1.05rem;margin-bottom:0.6rem">Live parties, composition, MPC on leaf shares, logging, and statistics — first-class, not buried in the reference.</p>
<div class="feature-grid">
<a class="feature-card" href="network_and_mpc.html"><span class="feature-kicker">Links</span><strong>Network &amp; party mesh</strong><p>TLS 1.3 party sessions, trio mesh, RoundSink rounds, lanes, and reconnect.</p></a>
<a class="feature-card" href="protocol_compose.html"><span class="feature-kicker">Schedule</span><strong>Protocol composition</strong><p>FSS walks next to Beaver opens on one RoundSink plan, with explicit reshares.</p></a>
<a class="feature-card" href="beaver_triples.html"><span class="feature-kicker">Multiplication</span><strong>Beaver triples</strong><p>Authenticated products on leaf shares, including the ABY2.0 MAC check.</p></a>
<a class="feature-card" href="arith_runtime.html"><span class="feature-kicker">Shares</span><strong>Arithmetic share runtime</strong><p>edaBits, truncate, share compare, matmul, and hidden shuffle beside FSS.</p></a>
<a class="feature-card" href="yao_leaf.html"><span class="feature-kicker">Circuits</span><strong>Yao on a leaf</strong><p>Split a leaf into bits, garble a netlist, share the answer back as a leaf.</p></a>
<a class="feature-card" href="experiment_costs.html"><span class="feature-kicker">Observability</span><strong>Logging &amp; statistics</strong><p>Leveled run logs, provenance banners, replayable seeds, and CSV wire / PRG / timing breakdowns.</p></a>
</div>
\endhtmlonly

View file

@ -158,3 +158,7 @@ grotto::offset_iterable shifted(knots.begin(), knots.end(), 10);
**Defined in**\n
@ref grotto/offset_iterable.hpp
\htmlonly
<div class="tldr"><b>TL;DR.</b> eval_interval and eval_full walk a subinterval; eval_sequence walks the listed points. indices_set_in, advice_bits_of, batch_of, tuple_as_zip, and rotated_by only change the step. None of them copy the buffer.</div>
\endhtmlonly

View file

@ -12,6 +12,10 @@ The same offset also drives [offset Horner](@ref offset_horner),
## Binomial jet {#offset_jet}
\htmlonly
<div class="eli5"><b>ELI5.</b> The dealer keys the binomial coefficients of (center + eta) up to a chosen degree. After eta is public, a dot with those shares is the monomial or the polynomial, with no further tree walk.</div>
\endhtmlonly
`make_offset_jet_keys(center, degree)` keys one incremental `gt` whose
payload is the vector of \f$\binom{\mathrm{center}}{k}\f$ in
\f$\mathbb{Z}/2^{64}\f$. After `eta` opens,
@ -61,6 +65,10 @@ one reciprocal after the shares are opened. Those are not separate APIs.
## Exact ring switch {#ring_switch}
\htmlonly
<div class="eli5"><b>ELI5.</b> An n-bit limb is rewritten into another modulus, a field, or a P-256 scalar by an exact map on the opened residue. The value does not go through floating point, and the map does not expand another key.</div>
\endhtmlonly
For an unsigned \f$n\f$-bit limb (\f$n\le 64\f$) with representatives in
\f$[0,2^n)\f$,
@ -117,6 +125,10 @@ See also [representation shift and twisted jets](@ref repr_and_twist).
## Offset Horner {#offset_horner}
\htmlonly
<div class="eli5"><b>ELI5.</b> Powers of a public center are already shared. Shifting them by the opened eta, the binomial way, evaluates the polynomial at the secret. The degree here is fixed in the template.</div>
\endhtmlonly
`make_offset_horner_keys<Input, Degree>(center)` keys one `gt` whose
payload is `center^m` for `m = 0 .. Degree`. `Degree` is at most 3
(`offset_horner_max_degree`). Pass `dpf::verifiable{}` for proof tokens.
@ -147,6 +159,10 @@ reusable key.
## Offset polynomial {#offset_poly}
\htmlonly
<div class="eli5"><b>ELI5.</b> Same shift as offset Horner, but the degree is an argument, so the number of powered shares is chosen when the keys are built.</div>
\endhtmlonly
`make_offset_poly_keys(center, degree)` is offset Horner at a runtime
degree, at most 16 (`offset_poly_max_degree`). One incremental `gt`
whose payload is the vector of powers. `offset_poly_eval<Party>` dots the shifted powers.
@ -176,6 +192,10 @@ auto opened = s0 + grotto::offset_poly_eval<1>(mat, knots, coeff, eta);
## Carry {#carry}
\htmlonly
<div class="eli5"><b>ELI5.</b> A carry across a shift is a short list of comparisons and bit corrections, not a generic circuit. The request names the source width, the shift, and the width of what comes out. Truncate, arithmetic shift, and sign-extend are the same plan with different output widths.</div>
\endhtmlonly
A `carry_request` names the source width `n`, the shift `s`, the output
width `out_n`, a `carry_mode` (`truncate_reduce`, `same_ring`, `extend`,
`window`), and a `sign_knowledge` (`unknown`, `nonnegative`, `negative`).
@ -210,6 +230,10 @@ auto clear = grotto::eval_carry_clear(keys.recipe, x0, x1);
## Prefix parity {#prefix_parity}
\htmlonly
<div class="eli5"><b>ELI5.</b> One walk of an existing key stops at the public endpoints and folds XOR or addition along that prefix. The fold is O(number of endpoints), not a new key per prefix.</div>
\endhtmlonly
`prefix_parities(key, endpoints)` walks a key to the sorted endpoints and
returns XOR shares of the prefix parities, plus the index of the first
endpoint on the wrap. `segment_parities` turns those into one share per
@ -241,6 +265,67 @@ auto signs = grotto::signed_prefix_parities(cmp0, ends);
**Defined in**\n
@ref grotto/prefix_parity.hpp
## Several LUTs, one comparison {#lut_union}
\htmlonly
<div class="eli5"><b>ELI5.</b> Stack the breakpoints of every table into one sorted list. One comparison and one prefix walk label the pieces of that list. Each table then sums the labels that fall inside its own intervals.</div>
\endhtmlonly
`make_lut_union_plan(luts, eta)` shifts every piecewise LUT by the public
`eta`, inserts the same domain-minimum and carry cuts as
[offset polynomial](@ref offset_poly), and sorts the union. A piece of one
LUT is a span of those union knots: `[begin, end)`, or
`[begin, end-of-union) ∪ [0, end)` when the piece wraps. The span stores
that piece's binomial shift by its public `kappa`. Spans of one LUT
partition the union.
The interactive plan is one comparison, whatever the number of LUTs and
whatever the number of union knots:
- `plan.comparisons` and `plan.prefix_walks` are 1.
- `plan.depth` and `plan.geneval_rounds()` are the bitlength of the input.
- `plan.degree` is the widest polynomial. The payload is
`1, center, …, center^degree`.
- `schedule_lut_union` records one `fss_cmp` of that depth. The slot is
`lut_union_slot_bytes`: one AES block, or `lanes * 8` when the power
vector is wider. Prefix parity of the union is local after that
comparison.
`lut_union_eval<Party>` reads one `make_offset_poly_keys` key of degree
at least `plan.degree` and returns a share per LUT. `geneval_lut_union`
opens that comparison from XOR shares of the center, as in
`geneval_offset_horner`. Pass `dpf::arith_input` when the shares add to
the center in the input group.
`piecewise_from_easy` and `piecewise_from_constant` adapt the cleartext
tables. An `easy_lut` denominator other than 1 is a rounding division, so
`piecewise_from_easy` rejects it. Powers that are zero on every piece are
dropped, and the shared payload stays only as wide as the widest remaining
degree.
\code{cpp}
grotto::piecewise_lut<std::uint8_t> low{{0, 10}, {{1, 0}, {0, 2}}};
grotto::piecewise_lut<std::uint8_t> high{{0, 4, 12}, {{3, 0}, {1, 1}, {9, 4}}};
const std::uint8_t center = 12;
const std::uint8_t eta = 3;
auto plan = grotto::make_lut_union_plan({low, high}, eta);
auto mat = grotto::make_offset_poly_keys<std::uint8_t>(center, plan.degree);
auto s0 = grotto::lut_union_eval<0>(mat, plan);
auto s1 = grotto::lut_union_eval<1>(mat, plan);
dpf::protocol::composer composer(0);
grotto::schedule_lut_union(composer, plan); // rounds == plan.depth
\endcode
**Code samples**\n
<div class="tabbed">
- <b class="tab-title">lut_union.cpp</b> \include{cpp} grotto/lut_union.cpp
</div>
**Defined in**\n
@ref grotto/lut_union.hpp
## Cleartext maps {#grotto_luts}
These functions take a raw fixed-point word (`n << fractional_bits`) and
@ -250,6 +335,10 @@ return a raw word. They do not build a DPF. The type
## Fixed-point product {#fixedpoint_mul}
\htmlonly
<div class="eli5"><b>ELI5.</b> The product lives in a ring wide enough for both fixed-point operands. One Beaver triple in that ring is the product; the binary point is placed by a public shift afterward.</div>
\endhtmlonly
`fixed_mul<IntegerBits, FractionalBits>(lhs, rhs)` multiplies two
`fixedpoint` values and keeps that many integer bits (including the sign)
and fraction bits. Bits below the fraction are floored. The product type
@ -267,6 +356,10 @@ auto prod = grotto::fixed_mul<16, 16>(q16{1.5}, q16{2.0});
## Lookup tables {#lookup_tables}
\htmlonly
<div class="eli5"><b>ELI5.</b> A cleartext approximation is replaced by a table addressed with the secret. Constant, easy, range, window, and principal tables differ in how many bits of the input they consume and how the correction is added.</div>
\endhtmlonly
Constant, easy, principal, range, and window tables are included from
`grotto.hpp`. The dyadic table comes in through `exact_steps.hpp`, which
`grotto.hpp` also includes.
@ -303,6 +396,8 @@ Constant, easy, principal, range, and window tables are included from
(`principal_precision`). Names: `ln`, `exp`, `sin`, `tanf`, `tang`,
`sinh`, `cosh`, `sqrt`, `coth`, `sec`, `gsec`, `csch`, `inv`, `rsqrt`,
`invsq`.
- **Wavelet.** Haar and bior(5,3) compressed tables:
[Wavelet lookup tables](@ref dwt_luts).
\code{cpp}
auto sign = grotto::make_exact_constant_lut<std::int32_t>(
@ -368,6 +463,69 @@ picks the piece and calls that Horner step.
@ref grotto/window_lut.hpp, @ref grotto/principal_lut.hpp,
@ref grotto/piecewise.hpp
## Wavelet lookup tables {#dwt_luts}
\htmlonly
<div class="eli5"><b>ELI5.</b> A wavelet step is a fixed linear combination. The LUT stores that combination so the signal stays in shares and never enters a floating-point routine. Haar and biorthogonal 5/3 are the two filters.</div>
\endhtmlonly
`make_haar_dwt_lut` and `make_bior53_dwt_lut` compress a real signal of
length \f$2^n\f$ and evaluate it as a fixed-point word. The construction
is the cleartext Haar and bior(5,3) lookup of Reis, Ugurbil, Wagh, Henry,
and de Vega, [ePrint 2025/013](@ref bib_wave), Equations (7) and (8).
`sample_dwt_signal(domain_bits, fractional_bits, f)` writes the grid
\f$i \cdot 2^{-f}\f$ for \f$i \in [0, 2^n)\f$.
Both builders run the depth-\f$j\f$ low-pass with the smooth edge
extension used for that paper's accuracy tables. PyWavelets calls these
filters `haar` and `bior2.2`; bior(5,3) is the same pair, named there by
vanishing moments. Building either table is \f$\Theta(N)\f$ arithmetic
and extra memory, \f$N = 2^n\f$.
Haar then multiplies the approximation coefficients by \f$2^{-j/2}\f$
and rounds down to \f$f\f$ fraction bits. On this grid that coefficient
is the mean of each block of \f$2^j\f$ samples. Evaluation reads
`coeff[raw >> j]`, one indexing step, \f$\Theta(1)\f$.
bior(5,3) multiplies by \f$2^{j/2}\f$ and rounds down the same way.
Smooth extension prepends two coefficients, so the bin `msb = raw >> j`
lives at index `msb + 2`, and the next tap at `msb + 3`, wrapping in the
stored vector. With `lsb = raw mod 2^j`,
\f[
y = \bigl\lfloor\bigl(c_{\mathrm{msb}+2}\,(2^j - \mathrm{lsb})
+ c_{\mathrm{msb}+3}\,\mathrm{lsb}\bigr) / 2^{2j}\bigr\rfloor.
\f]
That is Equation (8): the Lemma 6 weights \f$(2^j - \mathrm{lsb}_j)\f$
and \f$\mathrm{lsb}_j\f$, in integer arithmetic. Two multiplications and
a shift, \f$\Theta(1)\f$.
The paper's online protocols look these tables up under a DPF. Haar is
paired there with a deterministic Pika truncation; bior(5,3) is paired
with segment parity. Those protocols are not a key type here. The value
they open is `table(raw)`.
\code{cpp}
auto samples = grotto::sample_dwt_signal(6, 4, [](double x) {
return 1.0 / (1.0 + std::exp(-(x - 2.0)));
});
auto haar = grotto::make_haar_dwt_lut(samples, 4, 2);
auto bior = grotto::make_bior53_dwt_lut(samples, 4, 2);
auto h = haar(std::uint64_t{32});
auto b = bior(std::uint64_t{33});
\endcode
**Code samples**\n
<div class="tabbed">
- <b class="tab-title">dwt_lut.cpp</b> \include{cpp} grotto/dwt_lut.cpp
</div>
**Defined in**\n
@ref grotto/dwt_lut.hpp
## Closed form {#closed_form}
`eval_closed(closed::atanh, fractional_bits, raw)` composes
@ -423,3 +581,7 @@ for a functor type.
**Defined in**\n
@ref grotto/gadgets.hpp, @ref grotto/gadget_hints.hpp
\htmlonly
<div class="tldr"><b>TL;DR.</b> Open eta = x − r once. Jets and Horner turn that public distance into polynomial powers. Ring switch and carry move the integer. Prefix parity folds a key you already hold. Lookup tables, including the wavelet tables, replace cleartext math. A fixed-point product is one Beaver triple.</div>
\endhtmlonly

View file

@ -4,6 +4,7 @@
- \subpage output_type_examples
- \subpage evaluation_examples
- \subpage grotto_examples
- \subpage protocol_examples
- \subpage iteratable_examples
\page input_type_examples Domain samples
@ -17,24 +18,31 @@
- \subpage input_types_2custom_8cpp
\page "input_types_2integral_types_8cpp" input_types/integral_types.cpp
\include{cpp} input_types/integral_types.cpp
\page "input_types_2extended_types_8cpp" input_types/extended_types.cpp
\include{cpp} input_types/extended_types.cpp
\page "input_types_2modint_8cpp" input_types/modint.cpp
\include{cpp} input_types/modint.cpp
\page "input_types_2bitstring_8cpp" input_types/bitstring.cpp
\include{cpp} input_types/bitstring.cpp
\page "input_types_2keyword_8cpp" input_types/keyword.cpp
\include{cpp} input_types/keyword.cpp
\page "input_types_2xor_wrapper_8cpp" input_types/xor_wrapper.cpp
\include{cpp} input_types/xor_wrapper.cpp
\page "input_types_2custom_8cpp" input_types/custom.cpp
\include{cpp} input_types/custom.cpp
\page output_type_examples Payload samples
@ -42,30 +50,42 @@
- \subpage output_types_2integral_types_8cpp
- \subpage output_types_2extended_types_8cpp
- \subpage output_types_2bit_8cpp
- \subpage output_types_2gf2_8cpp
- \subpage output_types_2bitstring_8cpp
- \subpage output_types_2wildcard_8cpp
- \subpage output_types_2xor_wrapper_8cpp
- \subpage output_types_2custom_8cpp
\page "output_types_2integral_types_8cpp" output_types/integral_types.cpp
\include{cpp} output_types/integral_types.cpp
\page "output_types_2extended_types_8cpp" output_types/extended_types.cpp
\include{cpp} output_types/extended_types.cpp
\page "output_types_2bit_8cpp" output_types/bit.cpp
\include{cpp} output_types/bit.cpp
\page "output_types_2gf2_8cpp" output_types/gf2.cpp
\include{cpp} output_types/gf2.cpp
\page "output_types_2bitstring_8cpp" output_types/bitstring.cpp
\include{cpp} output_types/bitstring.cpp
\page "output_types_2wildcard_8cpp" output_types/wildcard.cpp
\include{cpp} output_types/wildcard.cpp
\page "output_types_2xor_wrapper_8cpp" output_types/xor_wrapper.cpp
\include{cpp} output_types/xor_wrapper.cpp
\page "output_types_2custom_8cpp" output_types/custom.cpp
\include{cpp} output_types/custom.cpp
\page evaluation_examples Evaluation samples
@ -84,52 +104,84 @@
- \subpage evaluation_2eval_dpf3_cmp_ic_8cpp
\page "evaluation_2eval_point_8cpp" evaluation/eval_point.cpp
\include{cpp} evaluation/eval_point.cpp
\page "evaluation_2eval_interval_8cpp" evaluation/eval_interval.cpp
\include{cpp} evaluation/eval_interval.cpp
\page "evaluation_2eval_full_8cpp" evaluation/eval_full.cpp
\include{cpp} evaluation/eval_full.cpp
\page "evaluation_2defer_eval_8cpp" evaluation/defer_eval.cpp
\include{cpp} evaluation/defer_eval.cpp
\page "evaluation_2eval_sequence_8cpp" evaluation/eval_sequence.cpp
\include{cpp} evaluation/eval_sequence.cpp
\page "evaluation_2memoizers_8cpp" evaluation/memoizers.cpp
\include{cpp} evaluation/memoizers.cpp
\page "evaluation_2output_buffers_8cpp" evaluation/output_buffers.cpp
\include{cpp} evaluation/output_buffers.cpp
\page "evaluation_2eval_inner_product_8cpp" evaluation/eval_inner_product.cpp
\include{cpp} evaluation/eval_inner_product.cpp
\page "evaluation_2buffered_prg_8cpp" evaluation/buffered_prg.cpp
\include{cpp} evaluation/buffered_prg.cpp
\page "evaluation_2eval_dpf3_point_8cpp" evaluation/eval_dpf3_point.cpp
\include{cpp} evaluation/eval_dpf3_point.cpp
\page "evaluation_2eval_dpf3_doerner_shelat_8cpp" evaluation/eval_dpf3_doerner_shelat.cpp
\include{cpp} evaluation/eval_dpf3_doerner_shelat.cpp
\page "evaluation_2eval_dpf3_cmp_ic_8cpp" evaluation/eval_dpf3_cmp_ic.cpp
\include{cpp} evaluation/eval_dpf3_cmp_ic.cpp
\page grotto_examples Grotto samples
- \subpage grotto_2jet_and_ring_8cpp
- \subpage grotto_2repr_and_twist_8cpp
- \subpage grotto_2lut_union_8cpp
- \subpage grotto_2dwt_lut_8cpp
\page "grotto_2jet_and_ring_8cpp" grotto/jet_and_ring.cpp
\include{cpp} grotto/jet_and_ring.cpp
\page "grotto_2repr_and_twist_8cpp" grotto/repr_and_twist.cpp
\include{cpp} grotto/repr_and_twist.cpp
\page "grotto_2lut_union_8cpp" grotto/lut_union.cpp
\include{cpp} grotto/lut_union.cpp
\page "grotto_2dwt_lut_8cpp" grotto/dwt_lut.cpp
\include{cpp} grotto/dwt_lut.cpp
\page protocol_examples Protocol composition samples
- \subpage protocol_2compose_schedule_8cpp
\page "protocol_2compose_schedule_8cpp" protocol/compose_schedule.cpp
\include{cpp} protocol/compose_schedule.cpp
\page iteratable_examples Iterable samples
- \subpage iterables_2setbit_index_iterable_8cpp
@ -140,19 +192,25 @@
- \subpage iterables_2zip_iterable_8cpp
\page "iterables_2setbit_index_iterable_8cpp" iterables/setbit_index_iterable.cpp
\include{cpp} iterables/setbit_index_iterable.cpp
\page "iterables_2advice_bit_iterable_8cpp" iterables/advice_bit_iterable.cpp
\include{cpp} iterables/advice_bit_iterable.cpp
\page "iterables_2parallel_bit_iterable_8cpp" iterables/parallel_bit_iterable.cpp
\include{cpp} iterables/parallel_bit_iterable.cpp
\page "iterables_2subinterval_iterable_8cpp" iterables/subinterval_iterable.cpp
\include{cpp} iterables/subinterval_iterable.cpp
\page "iterables_2subsequence_iterable_8cpp" iterables/subsequence_iterable.cpp
\include{cpp} iterables/subsequence_iterable.cpp
\page "iterables_2zip_iterable_8cpp" iterables/zip_iterable.cpp
\include{cpp} iterables/zip_iterable.cpp

View file

@ -1,5 +1,9 @@
# Multiparty & 3-server {#multiparty}
\htmlonly
<div class="eli5"><b>ELI5.</b> make_dpf3 is two ordinary spines per party and a Shamir payload, so any two parties open and the third key is independent of the value. make_it_dpf3 shares the whole truth table instead, and all three shares are required. Doerner–Shelat and geneval are the two-party, no-dealer alternatives.</div>
\endhtmlonly
Two-party keys are the default. The library also builds three-evaluator
`(2,3)` keys, information-theoretic three-server DPFs, and dealer-free
two-party keygen when the index is already shared.
@ -12,6 +16,13 @@ two-party keygen when the index is already shared.
| [dpf::make_it_dpf3](@ref dpf/it_dpf3.hpp) | Information-theoretic 3-server DPF ([ePrint 2023/028](@ref bib_itdpf)) | `it_dpf3.hpp` |
| [dpf::make_dpf_doerner_shelat](@ref dpf/doerner_shelat.hpp) | Two parties, shared index, no dealer for the point | `doerner_shelat.hpp` |
| [dpf::geneval_*](@ref dpf/geneval.hpp) | Answer shares for one query, no reusable key | `geneval.hpp` |
| [dpf::shamir::deal](@ref dpf/shamir.hpp) | (K,N) Shamir. `(2,3)` is the `make_dpf3` payload | `shamir.hpp` |
The payload inside `make_dpf3` is degree-1 Shamir on the points `1`, `2`,
and `3`. That access structure is `shamir::two_of_three`, the type
`shamir_share`. The same split takes other thresholds:
`shamir::deal<T, K, N>` and `shamir::share_secret`. `make_dpf3` stays
`(2,3)`. See [secret shares](@ref secret_shares) and `examples/mwe/shamir.cpp`.
```cpp
auto [k1, k2, k3] = dpf::make_dpf3(std::uint8_t{42}, dpf::fp61{7});

View file

@ -1,5 +1,9 @@
# Multipoint keys {#multipoint_keys}
\htmlonly
<div class="eli5"><b>ELI5.</b> Cuckoo hashing puts each secret point in one of a few buckets, with three hash functions and one point key per bucket. A query probes its three candidate buckets. Key size is linear in the number of points. The S&amp;P 2025 PCG packing, which would derive those buckets from one seed, is not implemented.</div>
\endhtmlonly
`make_multipoint(alphas, betas)` packs many secret points into cuckoo
buckets (de Castro–Polychroniadou, EUROCRYPT 2022, §4 / [ePrint 2021/580](@ref bib_vdpf)):
`κ = 3` hashes, one point key per bucket. The bucket count is linear in

View file

@ -0,0 +1,92 @@
# Network, parties, and MPC around the keys {#network_and_mpc}
\htmlonly
<div class="eli5"><b>ELI5.</b> The library is about DPFs first. The same headers also wire the parties together, schedule rounds, multiply and garble leaf shares, write leveled run logs, and emit paper-cost statistics — so a PIR or Duoram sketch does not start from bare sockets.</div>
\endhtmlonly
Keys and evaluation are the core story. This page is the map of everything
that sits *around* those keys when you run a real protocol: links, party
roles, composition, arithmetic and boolean MPC on leaf shares, logging, and
cost instrumentation. Open a linked page for the details; the [call index](@ref api_reference)
lists the headers.
## Network and party mesh
| Piece | Role |
| --- | --- |
| [party_session](@ref dpf/net/party_session.hpp) | One role joins a mesh: dial/accept, lanes, reconnect, dealer link |
| [trio](@ref dpf/net/trio.hpp) | Convenience `(2+1)` mesh (`p0`, `p1`, dealer `p2`) |
| [TLS 1.3](@ref dpf/net/tls.hpp) / [security](@ref dpf/net/security.hpp) | Default on peer links; `--encryption=off` for plaintext benches |
| [RoundSink](@ref dpf/net/round_sink.hpp) / [edge_mesh](@ref dpf/net/edge_mesh.hpp) | Batched exchange rounds on star / clique / dealer topologies |
| [stream arrays](@ref dpf/net/stream_array.hpp) | Sync or async TCP (and optional SCTP) byte lanes |
Processes may start in any order: lower ids accept, higher ids connect and
retry. After connect (or TLS handshake) both ends exchange a fixed hello so
mismatched party id, epoch, transport, or lane count fail at join time.
```cpp
// Conceptual shape — see party_session / trio headers for the full API.
dpf::net::session_options opt; // TLS on by default
dpf::net::party_session session(/*role*/, opt);
session.join(/*host:port table*/); // or in-process rendezvous
auto & sink = session.round_sink(/*peer*/);
```
**Go deeper:** [trio.hpp](@ref dpf/net/trio.hpp),
[party_session.hpp](@ref dpf/net/party_session.hpp),
[secure_channel.hpp](@ref dpf/net/secure_channel.hpp),
examples under `examples/protocol/`.
## Protocol composition
[dpf::protocol::composer](@ref protocol_compose) records FSS walks, Beaver
opens, and reshares as one RoundSink plan. Share domains are tagged
(`fss`, `a`, `b`, `rss`, `y`); party-count changes are explicit `reshare`.
Independent opens share a wave; dependency chains become successive waves.
Drive with `drive` / `drive_via_schedule`, or party helpers
`util::drive_composed` / `util::drive_composed_trio`. Named application
skeletons (PIR upload/answer, SUBLEQ, hushmap, …) live in
[app_plans.hpp](@ref dpf/app_plans.hpp).
**Go deeper:** [Protocol composition](@ref protocol_compose).
## MPC on leaf shares
The DPF answers the secret index. What you do *with* the opened (or still
shared) leaf is ordinary MPC in the same library:
| Tool | When you reach for it |
| --- | --- |
| [Beaver triples](@ref beaver_triples) | Products, dots, polynomials, optional MACs (ABY2.0) |
| [Arithmetic share runtime](@ref arith_runtime) | edaBits, truncate, share compare, matmul, hidden shuffle |
| [Yao on a leaf](@ref yao_leaf) | A boolean circuit the key does not contain; half-gates |
| [Dealer-free keygen](@ref dealer_free) | Doerner–Shelat / IKNP when nobody deals the pads |
| [Multiparty & 3-server](@ref multiparty) | `(2,3)` Shamir spines or IT three-server tables |
**Go deeper:** those capability pages, then the [guided tour](@ref guided_tour)
sections on Beaver, Yao, and running protocols.
## Logging and statistics
Observability is first-class, not an afterthought:
| Facility | Role |
| --- | --- |
| [Run log](@ref run_log) (`log.hpp` / `run_log.hpp`) | Leveled `key=value` lines to stderr, syslog, or a file; provenance banner (build, host, argv, env, config); link and trial events |
| [Statistics & CSVs](@ref statistics) (`experiment.hpp`) | Replayable master seeds, per-party streams, wire bytes, rounds, wall/CPU, symmetric-key blocks by purpose×primitive |
| [`prg::count`](@ref dpf/prg_count.hpp) | Thread-local expand / hash / harness counters (AES, ChaCha, LowMC) |
| [`thread_work`](@ref dpf/thread_work.hpp) | Charge compute-pool kernels back to the owning party |
In-tree `party/` drivers and `run_parties` / `app::run` / `app::run_measured`
call `app::start_logging` and drive the mesh end-to-end. Before any party
starts its clock, all parties wait at a start gate.
**Go deeper:** [Logging, statistics, and experiments](@ref experiment_costs),
[compose.hpp](@ref dpf/compose.hpp),
[experiment.hpp](@ref dpf/experiment.hpp),
[log.hpp](@ref dpf/log.hpp).
\htmlonly
<div class="tldr"><b>TL;DR.</b> Start with make_dpf and eval. When you need live parties, party_session / trio give TLS links and RoundSink rounds; the composer schedules FSS next to Beaver and Yao; start_logging and experiment stamp the run log and paper CSVs. The keys stay the product — the net, MPC, logging, and statistics stack is how you run and measure them.</div>
\endhtmlonly

View file

@ -1,3 +1,5 @@
An output type is the group element at the secret index.
Every output on one key has the same width, and the type is trivially copyable.
Leaf addition is the group operation.
@ -11,8 +13,9 @@ Leaf addition is the group operation.
| `vec<T, N>` | `N` lanes, no carry between them | [vec.hpp](@ref dpf/vec.hpp) |
| `wildcard_value<T>` | Filled in later | [wildcard.hpp](@ref dpf/wildcard.hpp) |
| `field64` / `field128` / `fp61` | Prime field | [fp61.hpp](@ref dpf/fp61.hpp) |
| `gf2` / `gf22` / `gf24` / `gf28` / `gf216` / `gf232` / `gf264` | GF(2^k) | [gf2.hpp](@ref dpf/gf2.hpp) |
| `p256` / `p256_scalar` | Curve point or scalar | [p256.hpp](@ref dpf/p256.hpp) |
| Shares | (2,2), (3,3), or (2,3) replicated | [secret shares](@ref secret_shares) |
| Shares | (2,2), (3,3), (2,3) replicated, (K,N) Shamir | [secret shares](@ref secret_shares) |
| `fixedpoint` | Fixed-point word | [fixedpoint.hpp](@ref grotto/fixedpoint.hpp) |
`bool` is an 8-bit integer. A one-bit payload is `dpf::bit`.
@ -27,7 +30,6 @@ The notes below are closed. Open one when you need the rules.
<details class="type-note">
<summary>Integer scalar types</summary>
Fixed-width integers (`uint32_t`, `int64_t`, and the other `psnip` widths)
are an additive group. Leaves use SIMD add, subtract, and multiply, including
`char`, `long long`, `char16_t`, `char32_t`, and `wchar_t` at 8, 16, 32, and
@ -46,6 +48,9 @@ add the underlying word; values wider than one AES block add with
<details class="type-note">
<summary>Secret shares</summary>
\htmlonly
<div class="eli5"><b>ELI5.</b> The wrapper is the same bits as the payload, plus a rule for opening. Subtractive shares open as share0 − share1, which is what a point leaf returns. Additive shares open as a sum, which is what a comparison returns. A replicated share gives two of the three additive pieces to each party, so any two parties suffice. A Shamir share is one point on a polynomial; any K of the N points open the constant term.</div>
\endhtmlonly
`dpf::additive_share<T, Party>` and `dpf::subtractive_share<T, Party>` are
(2,2) shares, layout-identical to `T` (`Party` is `0` or `1`).
@ -62,6 +67,38 @@ so the secret is `x_0 + x_1 + x_2`. Any two parties reconstruct.
`as_additive3()` is that party's (3,3) component. `add_replicated` folds a
(3,3) sharing in by updating both holders of each component.
`dpf::shamir::share<T, Party, K, N>` is a Shamir share over a field.
The secret is the constant term of a polynomial of degree `K - 1`.
Party `i` holds that polynomial at `x = i + 1`. Any `K` shares open it.
`make_shamir_shares<K, N>` and `shamir::deal` are the deterministic split.
`shamir::share_secret` draws the higher coefficients with `uniform_sample`.
`shamir::reconstruct` opens typed shares, or runtime `shamir::point_share`s.
The first `K` shares are the interpolating set and are not checked.
Each further share must lie on that polynomial. That catches a share that
does not belong. It does not name the bad share, and it does not correct it.
Exactly `K` shares accept any values. When `K = N` that is every share, so
a tampered full set is not detected. A complete set of shares of a different
secret is consistent and opens that secret. A zero leading coefficient is a
lower threshold than `K`.
`(2,3)` is `shamir::two_of_three`. `shamir::share<T, Party, 2, 3>` is
`dpf::shamir_share<T, Party>` (`sharing::shamir`).
`make_shamir_shares(secret, slope)` is that case with one coefficient.
`shamir3` is the same polynomial on `fp61`, with party indices `1`, `2`,
and `3`. `make_dpf3` uses that case. The field inverse is
`detail::shamir_field`. `fp61` and `gf2n` specialize it. For `gf2n`
the integers `1 .. N` are bit patterns, so `N` must be less than `2^k`
or a party lands on `0` or on another party's point. Addition in that
field is XOR, so a share added to itself opens `0`.
A complete program is `examples/mwe/shamir.cpp`.
\code{cpp}
const std::array<dpf::fp61, 2> coeff{{dpf::fp61{2}, dpf::fp61{3}}};
auto shares = dpf::make_shamir_shares<3, 5>(dpf::fp61{10}, coeff);
auto opened = dpf::shamir::reconstruct(
std::get<0>(shares), std::get<2>(shares), std::get<4>(shares));
\endcode
Creating from a plaintext puts the value on party 0. For a replicated share
that value is component `x_0`, which party 2 also stores as `next`.
@ -84,7 +121,8 @@ Conversions that keep the secret on one party are `a2b`, `b2a`, `a2fss`,
`fss2a`, `b2fss`, and `fss2b` (`a` additive, `b` subtractive, `fss` the leaf
share), plus `rss2y` and `y2rss` between a replicated share and its (3,3)
components (`y`). `s2y`, `y2s`, `s2rss`, and `rss2s` open a reconstructing
set and split again (`s` is (2,3) Shamir, `shamir_share`, points 1, 2, 3).
set and split again (`s` is the (2,3) case of (K,N) Shamir, `shamir_share`,
points 1, 2, 3).
A cast that changes how many parties hold the secret, including anything
that would need a garbled circuit, is not a conversion. `rss_mul` is the
replicated product: each party forms `x_i y_i + x_i y_{i+1} + x_{i+1} y_i`,
@ -94,7 +132,6 @@ and the three terms are an additive sharing of the product.
<details class="type-note">
<summary>Extended-precision integer scalar types</summary>
`simde_int128`, `simde_uint128`, `uint128_t`, and `uint256_t` are additive.
Their leaf arithmetic is ordinary addition of those integers.
</details>
@ -102,7 +139,6 @@ Their leaf arithmetic is ordinary addition of those integers.
<details class="type-note">
<summary>dpf::bit</summary>
A one-bit output. The group is XOR: `operator+` and `operator-` are both XOR,
and a leaf multiply is AND with an all-zero or all-one mask. Many `dpf::bit`
outputs are packed into each leaf, low bit first. `dpf::bit::zero` and
@ -123,7 +159,6 @@ auto [k0, k1] = dpf::make_dpf(std::uint8_t{3}, dpf::bit::one);
<details class="type-note">
<summary>dpf::twobit</summary>
A 2-bit output in the ring Z/4Z. Values are `0` through `3`
(`twobit::zero` .. `twobit::three`). Scalar `+` and `-` wrap modulo 4.
A leaf packs one lane every two bits, low lane in the low bits of the first
@ -144,7 +179,6 @@ auto [k0, k1] = dpf::make_dpf(std::uint8_t{3}, dpf::twobit::two);
<details class="type-note">
<summary>dpf::nyble</summary>
A 4-bit output in the ring Z/16Z. Values are `0` through `15`. Scalar `+`
and `-` wrap modulo 16. A leaf packs one lane every four bits, low nibble
first. Leaf addition is not XOR and is not a byte add: a carry must not
@ -165,7 +199,6 @@ auto [k0, k1] = dpf::make_dpf(std::uint8_t{3}, dpf::to_nyble(0xau));
<details class="type-note">
<summary>dpf::bitstring&lt;Nbits&gt;</summary>
A fixed string of bits in the XOR group. `operator+`, `operator-`, and leaf
addition are XOR. The leftmost character of a literal or of `to_string` is
the most significant bit, matching `0b` notation. Bits above `Nbits` are not
@ -176,7 +209,6 @@ part of the value. `dpf::bitN_t` is `bitstring<N>` for `N` from 1 through
<details class="type-note">
<summary>dpf::vec&lt;T, N&gt;</summary>
`N` lanes of an ordinary output `T`, stored with lane 0 in the least-significant
place. `T` is an integer, `modint`, `xint` / `xor_wrapper`, `fixedpoint`,
`twobit`, or `nyble`. `+`, `-`, and `*` on a `vec` run per lane and do not
@ -200,6 +232,9 @@ auto [k0, k1] = dpf::make_dpf(std::uint8_t{3}, beta);
<details class="type-note">
<summary>dpf::wildcard_value&lt;T&gt;</summary>
\htmlonly
<div class="eli5"><b>ELI5.</b> The leaf group is the group of T, but the correction word is withheld. Evaluation throws until assign writes it. The assign itself is a Beaver correction, not a new tree.</div>
\endhtmlonly
A placeholder for an output of type `T`. The leaf group is the group of `T`.
`make_dpf` accepts an empty `wildcard_value<T>{}` (or `dpf::wildcard<T>`, or
@ -245,6 +280,9 @@ A wildcard *input* is separate: it masks the secret index. See
<details class="type-note">
<summary>dpf::xor_wrapper&lt;T&gt;</summary>
\htmlonly
<div class="eli5"><b>ELI5.</b> The group operation is XOR of the underlying word, not addition. A domain of this type walks the bits in GF(2)^N. A leaf of this type opens by XORing the two shares.</div>
\endhtmlonly
An element of `GF(2)^n` for `n = 8 * sizeof(T)`, or `n = N` for
`dpf::xint<N>`. `operator+` and `operator-` are XOR, `operator*` is AND, and
@ -258,8 +296,42 @@ leaf scaling is AND.
</details>
<details class="type-note">
<summary>Prime fields and P-256</summary>
<summary>GF(2^k)</summary>
`dpf::gf2`, `gf22`, `gf24`, `gf28`, `gf216`, `gf232`, and `gf264` are
GF(2^k) for k = 1, 2, 4, 8, 16, 32, and 64. Leaf addition is XOR. Leaf
scaling is field multiplication. Widths below 8 bits pack one field
element per lane, low lane first, and addition XORs that lane. A negative
integer constructs the same element as its magnitude.
`shamir::deal` and `shamir::reconstruct` run over these fields. The
shareholder count `N` must be less than `2^k`: the point integers are
stored as bit patterns, and `2^k` itself is `0`. `gf2` can hold one
shareholder. `gf22` can hold three, which is a `(2,3)` sharing. A share
added to itself is `0`.
| Type | Modulus |
| --- | --- |
| `gf2` | `x + 1` |
| `gf22` | `x^2 + x + 1` |
| `gf24` | `x^4 + x + 1` |
| `gf28` | `x^8 + x^4 + x^3 + x + 1` (AES, `0x11B`) |
| `gf216` | `x^16 + x^5 + x^3 + x^2 + 1` |
| `gf232` | `x^32 + x^31 + x^28 + x^21 + 1` |
| `gf264` | `x^64 + x^63 + x^62 + x^53 + 1` |
```cpp
auto [k0, k1] = dpf::make_dpf(std::uint8_t{3}, dpf::gf28{0x1b});
auto [s0, s1, s2] = dpf::make_shamir_shares(dpf::gf28{0x1b}, dpf::gf28{0x5a});
auto opened = dpf::reconstruct(s0, s2);
```
**Defined in**\n
@ref dpf/gf2.hpp
</details>
<details class="type-note">
<summary>Prime fields and P-256</summary>
`dpf::field64` is GF(2^64 − 2^32 + 1), the same prime as libprio `Field64`.
`dpf::field128` is GF(340282366920938462946865773367900766209), the same
@ -294,6 +366,9 @@ recipes still use the integer limb channel.
<details class="type-note">
<summary>grotto::fixedpoint&lt;FractionalBits, IntegralType&gt;</summary>
\htmlonly
<div class="eli5"><b>ELI5.</b> The value is an integer with a public split between integer bits and fraction bits. Leaf addition is integer addition of the raw word. It does not shift the binary point.</div>
\endhtmlonly
A fixed-point output stored in an integer backend. `FractionalBits` is the
number of bits after the binary point. `IntegralType` defaults to
@ -318,7 +393,6 @@ auto [k0, k1] = dpf::make_dpf(std::uint8_t{3}, fp{1});
<details class="type-note">
<summary>Custom output type requirements</summary>
Specialize `dpf::leaf_arithmetic::add_t`, `subtract_t`, and `multiply_t` for
the exterior node type (`simde__m128i` for the default AES PRG, and
`simde__m256i` when that node is used). Each functor's call operator receives
@ -335,3 +409,7 @@ Outputs must be trivially copyable and standard layout. `dpf::utils::make_from_i
should build `T` from the integer `1` when tests or `make_default` need a
nonzero payload. See `test/tests/helpers/custom_output_type_small.hpp`.
</details>
\htmlonly
<div class="tldr"><b>TL;DR.</b> One key, one leaf width. Integers and modint add; xint, bit, bitstring, and GF(2^k) XOR; twobit and nyble wrap inside their lane. Typed shares remember whether opening adds or subtracts, and whether any two parties suffice. A wildcard leaf throws until it is assigned.</div>
\endhtmlonly

View file

@ -1,5 +1,9 @@
# Point-programmable vector commitments {#ppvc_manual}
\htmlonly
<div class="eli5"><b>ELI5.</b> The vector is bound before the coordinate is chosen. Each coordinate is a pair of 1-bit DPF roots, and both roots are committed with a Naor string. Opening one side of each pair writes the hidden coordinate, or the sum, onto a public index.</div>
\endhtmlonly
A point-programmable vector commitment binds a vector
`x` in `(Z/2^s Z)^n` and still lets one hidden coordinate be chosen
after the commitment is published.
@ -53,6 +57,10 @@ storage, so two expansions on one thread must not overlap.
## What an opening proves {#ppvc_verify}
\htmlonly
<div class="eli5"><b>ELI5.</b> verify recomputes the Naor string on each opened root and compares it to the commitment. A root that was not the committed one fails. The check does not reveal the other coordinates.</div>
\endhtmlonly
`verify` checks each opened root against its Naor string.
`accept` also checks the programmed statement: the rotated coordinate
when `mu` is 0, the column sum when `mu` is 1.
@ -89,3 +97,7 @@ The generator is `dpf::prg::aes128` unless another 128-bit PRG is named.
Naor's string commitment is Moni Naor, [Bit Commitment Using Pseudorandomness](@ref bib_naor), Journal of Cryptology 1991.
The point keys are the Boyle–Gilboa–Ishai construction named in
[DPF basics](@ref point_functions).
\htmlonly
<div class="tldr"><b>TL;DR.</b> Commit binds the vector before the coordinate is chosen, by committing both DPF roots. Open writes one coordinate or the sum onto a public index. verify checks those roots against the Naor strings.</div>
\endhtmlonly

View file

@ -1,5 +1,9 @@
# Programmability {#programmability}
\htmlonly
<div class="eli5"><b>ELI5.</b> A wildcard is a placeholder with the correction word not yet fixed. Assigning an input or a leaf is a Beaver multiplication against that placeholder: one blinded share is exchanged, then both parties write the same correction. An updatable leaf can be patched again the same way.</div>
\endhtmlonly
Fill in a secret index or payload after the key exists, rewrite an
updatable leaf, or commit to a vector before choosing which coordinate
to open.

View file

@ -7,6 +7,10 @@ with a public Pascal shift plus \f$\lambda^{\kappa}\f$.
## Representation shift {#offset_repr}
\htmlonly
<div class="eli5"><b>ELI5.</b> Offset Horner is the Pascal-matrix case of a shift-invariant recurrence. Representation shift is the same idea for a general linear recurrence: a public step count kappa advances the shared state without rebuilding the key.</div>
\endhtmlonly
Offset Horner is the unipotent (Pascal) case of a shift-invariant module.
Here the dealer keys an arbitrary state
\f$S_c\in(\mathbb{Z}/2^{64})^d\f$ at the hidden center. After `eta` opens, each
@ -55,6 +59,10 @@ squaring, at most 63 squarings) and applies it on every refined piece,
## Twisted jets {#offset_twist}
\htmlonly
<div class="eli5"><b>ELI5.</b> The dealer keys one comparison whose payload is the vector of twisted powers c^m λ^c, including the dyadic case c = 1/2. After eta opens, the parties scale that vector. They do not re-expand the tree.</div>
\endhtmlonly
The dealer keys one comparison whose payload is the vector of twisted powers
\f$c^{m}\lambda^{c}\f$ in \f$\mathbb{Z}/2^{64}\f$. After `eta` opens, the segment
walk returns those shares on the hot piece. A public binomial shift of the
@ -112,5 +120,10 @@ auto half_keys = grotto::make_offset_twist_keys<std::uint8_t>(
center, 2, grotto::twist_half);
\endcode
Offset Horner, offset polynomials, carry, prefix parity, and the cleartext
Offset Horner, offset polynomials, a union of several LUTs on one
comparison, carry, prefix parity, and the cleartext
LUTs are on [jet and ring](@ref jet_and_ring).
\htmlonly
<div class="tldr"><b>TL;DR.</b> Both start from the opened offset eta = x − r. Representation shift advances a linear recurrence by a public step count. Twisted jets scale a vector of powers that already includes the constant factor.</div>
\endhtmlonly

View file

@ -1,5 +1,9 @@
# Verifiability & authenticity {#verifiability}
\htmlonly
<div class="eli5"><b>ELI5.</b> The proof rides on the same walk as the payload. verifiable checks that the correction seeds were the honest ones. extractable checks that the path is weight 1, so a second programmed point fails. A MAC checks the leaf share after it is opened.</div>
\endhtmlonly
Prove that a DPF walk used honest correction seeds, or that a weight-1
sketch over the path is consistent. The tags ride along as extra
`make_dpf` arguments; eval still returns the usual leaf share.

View file

@ -17,11 +17,13 @@ Types in full are on [Input types](@ref input_types) and [Output types](@ref out
- The parties already share `alpha` and want a reusable key: [dpf::make_dpf_doerner_shelat](@ref dpf/doerner_shelat.hpp).
- The parties want the answer and no key: [dpf::geneval_point](@ref dpf/geneval.hpp), [dpf::geneval_interval](@ref dpf/geneval.hpp), [dpf::geneval_cmp](@ref dpf/geneval.hpp).
- Three evaluators: [dpf::make_dpf3](@ref dpf/dpf3.hpp) and [dpf::make_dpf3_doerner_shelat](@ref dpf/dpf3_ds.hpp).
- A Shamir secret for any threshold, not only (2,3): [shamir::deal](@ref dpf/shamir.hpp) and `examples/mwe/shamir.cpp`.
- One point, a comparison, or a public interval: [dpf::gt](@ref dpf/dcf.hpp) and the other predicates, and [dpf::ic](@ref dpf/interval.hpp).
- Many secret points: [dpf::make_multipoint](@ref dpf/multipoint.hpp).
- A vector whose hidden coordinate is chosen after the commitment: [point-programmable vector commitments](@ref ppvc_manual).
- Several lanes at one leaf: [dpf::vec](@ref dpf/vec.hpp).
- A payload filled in later: [dpf::wildcard_value](@ref dpf/wildcard.hpp).
- A proof the key is well formed: [dpf::verifiable](@ref dpf/verifiable.hpp).
- A leaf share must enter a bit circuit the key does not compute: [b2y](@ref yao_leaf) and the netlist on that page. A comparison, an interval, or a public-offset polynomial stays a key.
The call index, with one line each, is the [API reference](@ref api_reference).

124
doc/pages/yao.md Normal file
View file

@ -0,0 +1,124 @@
# A boolean function of a leaf {#yao_leaf}
\htmlonly
<div class="eli5"><b>ELI5.</b> The key still answers the point, the comparison, and the interval. When the leaf you already hold has to go through a bit circuit the key does not contain, split that leaf into XOR bits, garble the circuit, and share the result back in the leaf's own type.</div>
\endhtmlonly
Party 0 garbles. Party 1 evaluates. Free-XOR, half-gates
([ePrint 2014/756](@ref bib_halfgates)). One table row is 32 bytes.
Outputs are XOR shares of the output bits: the garbler's share is the
permute bit, the evaluator's share is the color of the label it holds.
The circuit this library is built around is the zero-key AES-MMO block
`prg::aes128::eval` already uses on every tree expand. A seed or a leaf
can enter that block without being opened. The same netlist is a hand-built
straight line of XOR, AND, XNOR, and NOT. Inputs are declared first.
| Leaf you hold | Into the netlist | Back out |
| --- | --- | --- |
| Point leaf (subtractive) | `b2y` | `y2b` |
| Comparison leaf (additive) | `a2y` | `y2a` |
| `fss_share` | `fss2y` | `y2fss` |
| Replicated, parties 0 and 1 garble | `rss2y` | `y2rss` |
Bits are least-significant first, one byte each, 0 or 1. `width` 0 means
the whole ring on the way in, and the vector length on the way out.
Both parties' shares are arguments. The edaBit open of `x - r` and the
daBit open of `b ⊕ r` are resolved inside the call, the same way
`edabit::a2b_gmw_pair` does. Those masked values are uniform. The integer
stays shared.
`rss2y` does not wake party 2. Party 0 already holds `x0` and `x1`. Party 1
already holds `x2`. `r0.next` must equal `r1.own`. Top-level `dpf::rss2y`
and `dpf::y2rss` are the local (3,3) casts and are different functions.
```cpp
auto [k0, k1] = dpf::make_dpf(std::uint8_t{42}, std::uint32_t{0x6b});
auto s0 = *dpf::eval_point(k0, std::uint8_t{42});
auto s1 = *dpf::eval_point(k1, std::uint8_t{42});
auto [y0, y1] = dpf::yao::b2y(s0, s1, 8);
dpf::yao::netlist n;
dpf::yao::bit in[8];
for (int i = 0; i < 8; ++i)
in[i] = n.shared_in();
n.out(n.and_(in[0], in[1]));
dpf::yao::session garbler;
auto out0 = garbler.eval(0, n, y0.data(), link); // party 1 passes y1
auto [z0, z1] = dpf::yao::y2b<std::uint32_t>(out0, out1, 1);
// reconstruct(z0, z1) == (0x6b & 1) & ((0x6b >> 1) & 1)
```
`session` keeps the IKNP base OT. The first `eval` that needs a choice
label runs Chou–Orlandi. Later evals on that session only extend. Do not
interleave `iknp::sample` on the same channel. Party 0 is the garbler for
the life of the session. Tables are one-time. The model is semi-honest.
## What stays a key {#yao_not}
A public query against a secret point is a comparison key. An interval is
an interval key. A polynomial in a public offset is Grotto, one comparison
and a local dot. A product of two leaf shares is a Beaver triple. A word
mux is one bit×ring inject. `cost_pass` still chooses among a DCF mask, an
edaBit MSB, and a full adder for those.
Garble when the AND depth is the cost and the circuit is this shape: the
MMO block (5120 ANDs, one message, about 160 KiB of tables), the AES-128
block under a shared key (6400 ANDs, key schedule included), or a netlist
you built because the leaf bits are the input. The Boyar–Peralta S-box
inside those blocks is 32 ANDs. `correction_level` is that hash as one garble: eight MMO blocks, lanes
`0..3` on each seed, 40960 ANDs, XOR shares of the four-block digest.
The level and the prefix share are packed into the inputs
(`pack_correction_level`); the netlist itself does not change per level.
`party/oblivious_hash.hpp` still walks the same S-box as GMW AND layers,
80 opens per level.
## A secret branch {#yao_stack}
The netlist above is a straight line. A secret if/else, or a secret choice
among a few blocks, still belongs on that leaf circuit: the bits came from
`b2y` / `a2y` / `rss2y`, and the answer goes back with `y2b` / `y2a` /
`y2rss`. It is not a reason to open the leaf.
`yao::eval_if` stacks the two branches
([Heath and Kolesnikov, CRYPTO 2020](@ref bib_stacked)). Each branch is
garbled from the hash of the control label for that semantic bit. The
generator XORs the materials. The evaluator rebuilds the inactive branch
from a seed under the control label and XORs it out. Transmitted AND rows
follow the heavier branch, plus four translation rows per output bit so the
garbler's share does not depend on the branch.
`yao::eval_one_hot` is the same stack for `k` netlists, `k` from 2 to 8
([Heath and Kolesnikov, CCS 2021](@ref bib_onehot)). The index is
`index_p0 XOR index_p1`. The demux row for the selector color carries the
inactive seeds.
Both parties pass the same netlists. Branch inputs use the same layout as
`eval_pair`: a shared entry is that party's XOR share, `priv0` is read on
party 0, `priv1` on party 1. The reconstructed bit is `share0[i] XOR share1[i]`.
`stack_blocks` is the stacked material. `naive_blocks` is the sum of the
branches. A comparison or a public-offset polynomial still stays on the key.
## Calls {#yao_calls}
| Call | What it does |
| --- | --- |
| `yao::netlist` | XOR, AND, XNOR, NOT, XOR with a public bit. `n_and()` is the row count |
| `yao::eval_plain` | The same wires in the clear |
| `yao::eval_local` / `eval_pair` | Garble and evaluate in one process |
| `yao::session::eval` | Garble on the peer channel. Party 0 sends the tables |
| `yao::a2y` `b2y` `fss2y` `rss2y` | Ring share to LSB-first XOR bits |
| `yao::y2a` `y2b` `y2fss` `y2rss` | Those bits back to a ring share |
| `yao::aes_mmo` | One zero-key MMO block on a session. `pos` is public |
| `yao::correction_level` | Eight of those blocks: the correction-seed hash for one tree level |
| `yao::aes128` | AES-128 of a shared block under a shared key |
| `yao::eval_if` | Stacked if/else. Rows follow the heavier branch ([CRYPTO 2020](@ref bib_stacked)) |
| `yao::eval_one_hot` | One stack over `k` branches ([CCS 2021](@ref bib_onehot)) |
**Go deeper:** [a secret branch](@ref yao_stack), [yao.hpp](@ref dpf/yao.hpp),
[yao_stack.hpp](@ref dpf/yao_stack.hpp), [yao_share.hpp](@ref dpf/yao_share.hpp),
[yao_aes.hpp](@ref dpf/yao_aes.hpp), [F_Yao](@ref yao.hpp),
[F_YaoShare](@ref yao_share.hpp), [half-gates](@ref bib_halfgates),
[edaBits](@ref bib_edabits). The walk that still uses GMW for the hash is
[the dealer-free tour](@ref tour_ds).