libdpf/doc/pages/evaluation.md
Ryan Henry 0d22946a0e Checkpoint the party/runtime stack before share-program and malicious-mode work.
Ship the TLS mesh, composer, Beaver/Yao/leaf MPC, prep/online paths, apps, and docs so the tree is pushable before elevating share_expr, security_mode, and prep resume.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-28 05:59:19 -06:00

613 lines
29 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

<!-- # Evaluating DPFs {#evaluation} -->
On this page:
[wildcard assign](@ref wildcard_assign),
[memoizers](@ref memoizers),
[buffers](@ref output_buffers),
[point](@ref eval_point),
[interval](@ref eval_interval),
[full domain](@ref eval_full),
[deferred input](@ref defer_eval),
[inner product](@ref eval_inner_product),
[sequence](@ref eval_sequence),
[three-party](@ref dpf3),
[IT 3-server](@ref it_dpf3).
`make_dpf(x, y)` returns one key per party. Evaluation of a key yields that
party's share. Leaf outputs are subtractive shares: open them with
`dpf::reconstruct`, which computes `share0 - share1`. Comparison outputs are
additive shares: `reconstruct` computes `share0 + share1`. A single-output
`eval_point` returns a small handle; `*handle` is the share. An unassigned
`dpf::wildcard` output throws `std::runtime_error`.
`[from, to]` is inclusive. The points passed to `eval_sequence` are a
nondecreasing range; an unsorted range throws `std::runtime_error`.
Let `n` be the input bit length and `λ` the PRG seed width (128 bits for
`prg::aes128`). A `make_dpf` key stores one `λ`-bit root and, on each of
the `n` levels, one `λ`-bit correction word and an advice bit, plus one
leaf per output slot. Key size is `Θ(n λ)` bits plus those payloads.
That layout is Boyle, Gilboa, and Ishai, CCS 2016 (full version
[ePrint 2018/707](@ref bib_fss2018)), whose stated length is `n(λ+2)+λ+⌈log2|G|⌉` bits, not
their [EUROCRYPT 2015](@ref bib_fss2015) key of `4n(λ+1)` bits. For a small output group `G`, Remark 3.4 of the full version stops
`ν = log2(λ / log2|G|)` levels early and shortens the key by `ν(λ+2)`
bits. This generator does that: `depth` is `n - lg(outputs_per_leaf)`,
where `outputs_per_leaf` is how many copies of `G` fit in one `λ`-bit
block. Those low input bits select the lane. A comparison still writes
a value word on each remaining level. The default expand is that CCS 2016
tree. Guo, Yang, Wang, Zhang, Xie, Zhang, and Liu ([ePrint 2022/1431](@ref bib_halftree)) keep
the same key length and the same `n`-hash point evaluation, and state
about `2n+2` random-permutation calls for key generation versus about
`4n`, and `1.5N` calls for a full-domain evaluation versus `2N`.
`prg::aes128_ccr` selects that Half-Tree expand.
Keygen expands once per level. `at<N>` and `idpf` add leaves; the spine
is still `n` levels. A comparison adds one payload word per level.
`block_width<B>` stores one word every `B` levels instead, and the seed
spine stays `Θ(n λ)`. That is ahead of Elette Boyle, Nishanth Chandran,
Niv Gilboa, Divya Gupta, Yuval Ishai, Nishant Kumar, and Mayank Rathee
(EUROCRYPT 2021, [ePrint 2020/1392](@ref bib_dcf)), who publish a value word on every
level: about `B` times fewer payload words, paid for by expanding the
siblings between checkpoints at evaluation. See the
[bibliography](@ref bibliography).
`eval_point` expands those `n` levels. A basic path memoizer keeps the
`n` nodes (`Θ(n λ)` bits) and the next call expands only the suffix after
the common prefix. The nonmemoizing memoizer keeps one node. `reconstruct`
of two shares, or of two or three Shamir shares, is a constant amount of
group arithmetic.
`eval_point<I>` and `eval_point(dpf::out<I>, key, x)` read output slot `I`.
`dpf::out<I, W>` is that slot with prefix length `W` checked against the
key. A bare leaf's prefix is the input bit length. `dpf::at<N>` uses `N`.
`dpf::prefix_deduce` (the default of `out<I>`) reads the prefix from the key.
`eval_point(dpf::cmp, key, x)` reads the comparison channel.
`eval_point(dpf::cmp_prefix<L>, key, x)` reads the first `L` bits of an
`idcf` key. The [eval_point](@ref eval_point) section has the calls.
On an `idpf` / multilevel key, `eval_prefixes(out<I,N>, key)` walks from
the root and materializes all `2^N` prefix shares (Poplar-style).
`idpf_eval_ctx` / `eval_until` ([ePrint 2021/017](@ref bib_poplar)'s EvaluateUntil) resume
under a live prefix list so the saved node count stays
`O(|prefixes|)`, not `O(2^N)`. The I-DPF max / k-th walk in
[dpf/idpf_agg.hpp](@ref dpf/idpf_agg.hpp) ([ePrint 2024/1190](@ref bib_idpfagg)) is that
caller. See [eval_until](@ref dpf/eval_until.hpp) and the application
mockup [I-DPF max and k-th](@ref app_idpf_agg).
## Assigning a wildcard leaf {#wildcard_assign}
\htmlonly
<div class="eli5"><b>ELI5.</b> The blank leaf is a Beaver slot from keygen. Filling it in rewrites the correction word: the parties exchange one blinded share of the new payload and both apply the same patch. assign_cmp rewrites the n comparison words locally and sends nothing. An updatable leaf is the same patch later, O(λ) and independent of the depth.</div>
\endhtmlonly
An output wildcard is a placeholder for a payload filled after keygen.
The type and the `dpf::wildcards` names are on
[Output types](@ref output_types). Evaluation of an unassigned slot throws
`std::runtime_error`.
In one process, each party holds a share of the payload and the two keys.
Slot `I` is `std::get<I>(key.leaf_nodes)` (`party_key` is the key). The
exchange is three calls:
\code{cpp}
auto & w0 = std::get<0>(k0.leaf_nodes);
auto & w1 = std::get<0>(k1.leaf_nodes);
auto b0 = w0.compute_and_get_blinded_output_share(share0);
auto b1 = w1.compute_and_get_blinded_output_share(share1);
auto l0 = w0.compute_and_get_leaf_share(b1);
auto l1 = w1.compute_and_get_leaf_share(b0);
w0.reconstruct_correction_word(l1);
w1.reconstruct_correction_word(l0);
\endcode
`share0` and `share1` are the two parties' shares of the payload.
`compute_and_get_blinded_output_share` also accepts a `secret_share`.
A wrapper that is already ready takes `begin_update()` before the same
three calls; that assign adds a difference onto the payload already
installed.
Over a socket, `key.async_assign_leaf(peer, share, token)` (and
`async_assign_leaf<I>`) runs that exchange. The method is present when
the library is built with ASIO. That is two rounds: blinded output
shares, then leaf shares. Each message is one share of the output (the
second round sends the packed leaf). The Beaver triple on the wildcard
slot is the preprocessing; keygen stored it. Local time is linear in
the leaf width.
`assign_cmp` on a pair of keys rewrites the `n` value-correction words
in place and sends nothing. `assign_cmp_local` is that patch on one key
once `δ` and the addend share are already known. An input wildcard is
one round: the parties exchange one `n`-bit offset share and reconstruct
locally. See [Input types](@ref input_types).
A wildcard *comparison* payload is different. Generate with
`dpf::lt(dpf::wildcard_value<Beta>{})` (or `leq` / `gt` / `geq`), then
`dpf::assign_cmp(k0, k1, if_true, if_false)` on the pair. Eval before
that throws. `dpf::assign_cmp_local(key, delta, addend_share)` is the
one-key form: both parties pass the same public `delta` and additive
shares of the absorb target.
\code{cpp}
auto [k0, k1] = dpf::make_dpf(std::uint8_t{40},
dpf::lt(dpf::wildcard_value<std::uint64_t>{}));
dpf::assign_cmp(k0, k1, std::uint64_t{7});
auto y0 = dpf::eval_point(dpf::cmp, k0, std::uint8_t{10});
\endcode
An input wildcard masks the index. The calls are
`offset_x.compute_and_get_share` and `offset_x.reconstruct` on
[Input types](@ref input_types). Eager `eval_interval` / `eval_full` require
that offset to be ready. To expand **before** assign (full-domain identity,
then rotate once `δ` is known), use
[Deferred input evaluation](@ref defer_eval).
**Defined in**\n
@ref dpf/leaf_wrapper.hpp, @ref dpf/dpf_key.hpp, @ref dpf/incremental.hpp
Memoizers hold interior nodes between calls. Output buffers hold the shares
a multi-point evaluation writes. Pass both as mutable named objects when a
later call should reuse them. The factories
`make_basic_path_memoizer`, `make_basic_interval_memoizer`, and
`make_*_sequence_memoizer` unwrap `party_key`, so a workspace built from
either party's type accepts both parties. Name that type with
`dpf::unwrap_party_key_t<std::decay_t<decltype(key)>>`.
## Memoizers {#memoizers}
\htmlonly
<div class="eli5"><b>ELI5.</b> The first walk stores interior nodes. The next query starts from the deepest stored node that still lies on its path, instead of from the root. An interval memoizer stores the nodes that cover a range; a sequence memoizer stores the nodes along a sorted list.</div>
\endhtmlonly
## Path memoizers {#path_memoizers}
`eval_point` walks one root-to-leaf path. `make_basic_path_memoizer<Key>()`
keeps every node of the previous point, and the next point recomputes only
the suffix after the common prefix. `make_nonmemoizing_path_memoizer<Key>()`
keeps one node and starts from the root on every call. A one-off
`eval_point(key, x)` uses the nonmemoizing memoizer.
Pass the memoizer as a mutable lvalue. The default argument is a new
temporary, so it has no previous point to resume from. Keep a separate
memoizer for each key you are in the middle of evaluating. A different root
restarts the path. One point evaluation is `Θ(n)` expands. The basic
memoizer's extra memory is the `n` saved nodes.
**Code samples**\n
<div class="tabbed">
- <b class="tab-title">memoizers.cpp</b> \include{cpp} evaluation/memoizers.cpp
</div>
## Interval memoizers {#interval_memoizers}
`eval_interval` and `eval_full` expand every leaf in a range.
`make_basic_interval_memoizer<Key>(from, to)` stores two levels of that
range. That is the workspace the convenience overloads allocate.
`make_full_tree_interval_memoizer<Key>(from, to)` keeps every level.
`make_basic_full_memoizer<Key>()` and `make_full_tree_full_memoizer<Key>()`
are the same workspaces sized for the whole domain.
Size the memoizer for the widest interval you will pass to it. A wider
interval throws `std::length_error`. The same key and the same endpoints
leave the final interior level in place. A different key or a different
interval rebuilds into the same allocation.
Passing only the memoizer still allocates a fresh output buffer and returns
`std::pair(buffer, iterable)`.
Let `L` be the number of inputs in the inclusive interval. The node count
at a level is the width of that interval shifted up to the level, and the
walk expands each of those nodes once, `Θ(n + L)` expands in total. The
output buffer is `L` slots per selected output. A basic memoizer keeps two
levels, `O(L)` nodes. A full-tree memoizer keeps every level, still
`Θ(n + L)` nodes. `eval_full` is the same walk with `L = 2^n`: time and
the output buffer are `Θ(2^n)` slots, and the basic full memoizer holds
two levels of that tree.
## Sequence memoizers {#sequence_memoizers}
`make_sequence_recipe<Key>(begin, end)` compiles a sorted point list into a
traversal. The recipe depends on the input type, and one recipe serves every
key of that type.
A sequence memoizer stores a reference to the recipe object it was built
from and checks later calls by address. Pass that same object, and keep the
recipe alive for as long as the memoizer is used. A copy of the recipe
throws `std::logic_error`.
`make_double_space_sequence_memoizer<Key>(recipe)` keeps two levels. It is
what `eval_sequence(key, recipe, buffer)` allocates when you omit the
memoizer. `make_inplace_reversing_sequence_memoizer<Key>(recipe)` keeps one
level and reverses direction as it descends.
`make_full_tree_sequence_memoizer<Key>(recipe)` retains every level. A key
whose depth differs from the recipe throws `std::logic_error`.
Let `m` be the number of listed points. `make_sequence_recipe` makes one
pass per level and binary-searches each live block, so the build is
`O(n m log m)` comparisons in the worst case. The recipe stores one step
per block per level, `O(n m)` records when every level splits. Evaluating
it expands one node per step, at most `O(n m)` expands, and writes `m`
shares. The double-space memoizer keeps two levels of that traversal
(`O(m)` nodes at the widest level). The inplace memoizer keeps one level.
The full-tree memoizer keeps every level.
## Output buffers {#output_buffers}
`output_buffer<T>` is move-only storage with `size`, iterators, `data`, and
`operator[]`. Build it with the factory that matches the evaluation:
- `make_output_buffer_for_interval(key, from, to)`
- `make_output_buffer_for_full(key)`
- `make_output_buffer_for_subsequence(key, begin, end, tag)`
- `make_output_buffer_for_recipe_subsequence(key, recipe, tag)`
On a `party_key`, leaf slots are `subtractive_share`s and comparison slots
are `additive_share`s. `dpf::bit`, `dpf::twobit`, and `dpf::nyble` slots are
packed. Trivially default-constructible slot types are left uninitialized;
the evaluation overwrites every slot it is responsible for.
`eval_interval` and recipe `eval_sequence` take the buffer as a non-const
reference, so the argument is a named object. The returned iterable refers
into that buffer. Read it while the buffer is alive, and only over the
points the iterable covers. The next evaluation overwrites those slots.
The buffer holds one slot per input in the interval, or one slot per
listed point. Its space is that many slots times the slot width. Packed
`bit`, `twobit`, and `nyble` slots share words.
For one output, the convenience overload returns the buffer itself as the
first element of the pair. For several output indices it returns a tuple of
buffers.
`make_output_buffer(dpf::out<I>, key, from, to)` and
`make_output_buffer(dpf::cmp, key, n)` size a buffer for one channel of a
multi-output or comparison key. The slot types follow the same party-share
rule.
**Code samples**\n
<div class="tabbed">
- <b class="tab-title">output_buffers.cpp</b> \include{cpp} evaluation/output_buffers.cpp
</div>
## dpf::eval_point {#eval_point}
`eval_point(key, x)` evaluates output 0 at one input.
`eval_point<I>(key, x)` selects another output. Two or more indices,
`eval_point<0, 1>(key, x)`, return a tuple of shares rather than handles.
`eval_point(key, x, path)` continues a path memoizer.
`eval_point(dpf::out<I>, key, x, path)` and
`eval_point(dpf::out<I, W>, key, x, path)` are the same walk with an
explicit slot. `W` must equal that slot's prefix.
`eval_point(dpf::cmp, key, x, path)` reads the comparison channel.
`eval_point(dpf::cmp_prefix<L>, key, x, path)` reads an `idcf` prefix of
`L` bits. Comparison results are additive shares. An unassigned wildcard
comparison throws until `assign_cmp`; see
[Assigning a wildcard leaf](@ref wildcard_assign).
\code{cpp}
auto [k0, k1] = dpf::make_dpf(
std::uint8_t{0x2a},
dpf::at<4>(std::uint8_t{5}),
std::uint8_t{9});
auto hi = *dpf::eval_point(dpf::out<0, 4>, k0, std::uint8_t{0x2a});
auto leaf = *dpf::eval_point(dpf::out<1, 8>, k0, std::uint8_t{0x2a});
auto [c0, c1] = dpf::make_dpf(std::uint8_t{40},
dpf::idcf(dpf::gt(std::uint64_t{1})));
auto full = dpf::eval_point(dpf::cmp, c0, std::uint8_t{50});
auto pref = dpf::eval_point(dpf::cmp_prefix<4>, c0, std::uint8_t{50});
\endcode
**Code samples**\n
<div class="tabbed">
- <b class="tab-title">eval_point.cpp</b> \include{cpp} evaluation/eval_point.cpp
</div>
## dpf::eval_interval {#eval_interval}
`eval_interval(key, from, to)` evaluates every input from `from` through
`to`. The iterable yields one share per input, in that order. Optional
arguments are an output buffer and then an interval memoizer. An output
index pack, `eval_interval<0, 1>(key, from, to, buffers, memo)`, writes each
selected output.
**Code samples**\n
<div class="tabbed">
- <b class="tab-title">eval_interval.cpp</b> \include{cpp} evaluation/eval_interval.cpp
</div>
## dpf::eval_full {#eval_full}
`eval_full(key)` is the closed interval from
`std::numeric_limits<Input>::min()` through `max()`. The buffer and
full-domain memoizer overloads match `eval_interval`.
`make_output_buffer_for_full(key)` and `make_basic_full_memoizer<Key>()`
size both for that domain.
**Code samples**\n
<div class="tabbed">
- <b class="tab-title">eval_full.cpp</b> \include{cpp} evaluation/eval_full.cpp
</div>
## Deferred input evaluation {#defer_eval}
\htmlonly
<div class="eli5"><b>ELI5.</b> The PRG expand runs once, into a full-domain buffer, before the index is known. When the offset opens, get() rotates that buffer. The tree is not expanded again.</div>
\endhtmlonly
Additive input blinding evaluates in tree coordinates `x ↦ x + δ`, where
`δ` is reconstructed by `assign_wildcard_input` into `offset_x`. Eager
`eval_interval` folds that map into the traversed range and throws if the
offset is not ready.
When `δ` is still unknown, call `defer_eval_interval` or `defer_eval_full`:
1. Expand the **full** input domain at identity into a full-sized buffer
(`make_output_buffer_for_full`).
2. After assign, `.get()` on the returned `deferred_rotated_subinterval`
rotates by `offset_x(0)` and yields the logical `[from, to]` (including
wrap-around when `from > to` in domain-walk order).
Interior-only prep (assigned input, leaf still a wildcard) is
`defer_traverse_interval`. Full-domain interior with both still unset is
`defer_traverse_full`.
Keep the deferred object alive for the lifetime of any `.get()` range.
Buffers must outlive both.
**Code samples**\n
<div class="tabbed">
- <b class="tab-title">defer_eval.cpp</b> \include{cpp} evaluation/defer_eval.cpp
</div>
**Defined in**\n
@ref dpf/deferred_rotated_subinterval.hpp, @ref dpf/eval_interval.hpp,
@ref dpf/eval_full.hpp
## dpf::eval_inner_product {#eval_inner_product}
\htmlonly
<div class="eli5"><b>ELI5.</b> The walk is the same as an interval or full-domain eval, but each leaf is multiplied by a public weight and added into one accumulator. The expanded vector is not stored. A sequence inner product does that only at the listed points.</div>
\endhtmlonly
`eval_inner_product` multiply-accumulates DPF shares against another vector
during the walk. It does not write the output vector.
Three local forms:
- **Batched leaf walk** (no tag): `eval_inner_product(key, from, to, weights, memo)`.
One output, weights in `eval_interval` layout, same batched exterior AES as
that walk, O(1) accumulator.
- **`dpf::paired`** (row-wise): one weight row per input. A scalar pairs with
one output; a `tuple` / `array` zips several outputs (leaf slots or an
ancestor prefix plus the leaf) off one path. Products are summed.
- **`dpf::columns`** (transposed): one output, several weight streams, one
accumulator per stream. Products stay apart. A stream is `w[i]` or `w(i)`.
`dpf::project` maps the share first; `dpf::also` sees it unmapped.
A single-stream `columns` result matches `paired` on that stream. A
single-output batched leaf walk matches `paired` when the interval is
leaf-aligned so covering-leaf weights equal the clipped domain points.
Unaligned intervals still weight every lane of the covering leaves (same
layout as one-key `eval_interval` buffers). Cohort evaluation uses a separate
*interleaved* leaf layout (`cohort_index`); `interleave_leaves` builds that
order from per-key buffers (see `dpf/interleave_leaves.hpp` and
`dpf/cohort.hpp`).
Pass `dpf::paired`. One element of the other vector is consumed per input, in
the same order as `eval_interval` or `eval_sequence`.
A scalar element pairs with one output:
```cpp
eval_inner_product(dpf::paired, key, from, to, weights);
eval_full_inner_product(dpf::paired, key, weights);
```
A `std::tuple` or `std::array` element pairs componentwise with the output
indices in the template pack. Those outputs may be several slots on the same
leaf, or an ancestor prefix slot and the leaf. Both are read from one path:
```cpp
eval_inner_product<0, 1>(dpf::paired, key, from, to, rows);
eval_sequence_inner_product<0, 1>(key, begin, end, rows);
eval_sequence_inner_product<0, 1>(key, recipe, begin, end, rows);
```
`share * component` must be defined. The product type must support `+`.
The walk costs the same as the interval or sequence it follows. Each
input adds a constant amount of arithmetic per paired output or per
column. The accumulator is the only result; there is no output vector
of shares.
The point list for a sequence or recipe is sorted nondecreasing, same as
`eval_sequence`. The recipe overload also checks that the list length matches
the recipe.
`dpf::columns` is the same walk with the products kept apart. One output,
several weight streams, one accumulator per stream:
```cpp
auto [sum, dot, sq] = eval_sequence_inner_product<0>(
dpf::columns, key, begin, end, std::tie(ones, r, r2));
```
A stream is anything with `w[i]` or `w(i)`, so a challenge can be a function
of the list index instead of a stored vector. `i` is the position in the
point list, or in the interval in `eval_interval` order. `dpf::project(fn)`
maps the share before the multiply. `dpf::also(fn)` is called as
`fn(i, x, share)` on the share before that map, which is how a mailbox slot
is updated in the same walk. Pass either tag, both, or neither.
The older single-output range walk (no `dpf::paired`) is still the batched
leaf inner product: `eval_inner_product<I>(key, from, to, weights, memo)`.
**Code samples**\n
<div class="tabbed">
- <b class="tab-title">eval_inner_product.cpp</b> \include{cpp} evaluation/eval_inner_product.cpp
</div>
## dpf::eval_sequence {#eval_sequence}
`eval_sequence(key, begin, end, tag)` evaluates a sorted list.
`dpf::return_output_only_tag_` stores one share per listed point.
`dpf::return_entire_node_tag_` stores whole leaves; it is the default.
The iterable still yields one share per listed point, in list order.
`eval_sequence(key, recipe, buffer, memo, tag)` repeats that list.
`memo` is a sequence memoizer bound to `recipe`. Omit `memo` to allocate a
`double_space` workspace for that call.
`eval_sequence_breadth_first(key, begin, end, buffer)` writes one share
per listed point, expanding a level at a time. The list is still sorted.
`eval_sequence_breadth_first(dpf::out<I>, key, begin, end)` allocates the
buffer and returns it. The `out<I, W>` form checks the prefix the same way
`eval_point` does. The cost is the same order as recipe `eval_sequence`
on that list: at most `O(n m)` expands and `m` output slots.
**Code samples**\n
<div class="tabbed">
- <b class="tab-title">eval_sequence.cpp</b> \include{cpp} evaluation/eval_sequence.cpp
</div>
## Buffered PRG {#buffered_prg}
\htmlonly
<div class="eli5"><b>ELI5.</b> One AES expand produces more blocks than a single tree node needs. The buffered PRG keeps the leftover blocks and serves the next nodes from them, so a wide walk makes fewer expands. The keys do not change.</div>
\endhtmlonly
`dpf::randomness::buffered_prg<PRG, Ts...>` (alias
`dpf::randomness::aes_buffered_prg<Ts...>`) is a forward cursor with one
PRG stream per value type. `get<I>()` and `fill<I>(out, n)` consume the
cursor. `at<I>(index)` reads an absolute index and leaves the cursor where
it is. `sampled<I>()` is how far `get` and `fill` have advanced.
`per_stream_buffer_elems` is at least 1.
`dpf::randomness::lane_table<T>` is the seekable form for a runtime set of
roles. `value_at(role, index)` and `mask_at(role, index)` are independent
streams, and a repeated index returns the same element.
`get` and `fill` of `q` elements do `Θ(q)` PRG work. The cursor keeps
`per_stream_buffer_elems` elements per stream. `at` reads one absolute
index and does not move the cursor. A `lane_table` keeps one cache
window per role (the constructor's `window`, default 256) for the value
stream and one for the mask stream. A repeated index returns the same
element. `fill_values` / `fill_masks` of `q` elements are `Θ(q)`.
**Code samples**\n
<div class="tabbed">
- <b class="tab-title">buffered_prg.cpp</b> \include{cpp} evaluation/buffered_prg.cpp
</div>
## Three-party (2,3) DPF {#dpf3}
\htmlonly
<div class="eli5"><b>ELI5.</b> Each evaluator key is two VDPF+ spines. eval_point walks both, Θ(n) expands, then scales into fp61. Opening two or three as_share values is a constant amount of field arithmetic. An updatable rewrite patches four leaves and refreshes an offset, independent of n.</div>
\endhtmlonly
`make_dpf3(α, β)` builds three evaluator keys after Guy Zyskind, Avishay Yanai, and Alex "Sandy" Pentland, [ePrint 2024/1658](@ref bib_dpf3), Figure 3:
two VDPF+ spines plus Shamir embedding in `fp61`. `eval_point` returns a
field share; open with `dpf::reconstruct` on any two (or all three)
`dpf::as_share` values. Tags: `verifiable`, `extractable`, `updatable`
(Fig. 10 in-place payload update).
`make_dpf3_doerner_shelat(x0, x1, β)` is the dual-spine Doerner–Shelat path:
XOR shares of `α`, same clear `β`. Socket orchestration lives in
`party/dist_dpf3.hpp` (`dist_with_dpf3_key`): after keygen, role **p0** holds
party 1, **p2** holds party 2, **p1** holds party 3. Path bits open to p0/p1;
p2 does not learn `α`. Dist keys are verifiable; payload updates use dealer
`updatable` keys or `remake_dpf3`.
Comparison / interval: `make_dpf3_cmp`, `make_dpf3_cmp_blocked`, `make_dpf3_ic`.
Multipoint: `make_multipoint3` (cuckoo buckets of point keys).
`make_dpf3` is local dealer work. Each evaluator key is two VDPF+ spines,
`Θ(n λ)` bits with a constant factor over one two-party key. `eval_point`
is two walks, `Θ(n)` expands, then a field scale. Opening two or three
`as_share` values is a constant number of `fp61` operations. An
`updatable` rewrite is four leaf patches and an offset refresh, `O(λ)`
and independent of `n`. `make_dpf3_doerner_shelat` is two
Doerner–Shelat spines: twice the rounds and the pad tape of one
two-party opening (see [Doerner–Shelat](@ref tour_ds)). Comparison,
interval, and multipoint forms add the same extra material as the
two-party comparison, interval, or cuckoo packing, on top of those spines.
## Information-theoretic 3-server DPF {#it_dpf3}
\htmlonly
<div class="eli5"><b>ELI5.</b> There is no PRG tree. For this domain each key is an additive share of a 256-word table. The three shares sum to beta at alpha and to zero elsewhere. A PIR answer is three inner products with the database; those three dots sum to the record.</div>
\endhtmlonly
`make_it_dpf3(α, β)` ([ePrint 2023/028](@ref bib_itdpf)) is a different object from
`make_dpf3`. Each of three parties holds an additive share of the
characteristic vector on `{0..255}`; the **sum** of all three
`eval_it_dpf3` values is the point function. Any single key is
independent of `(α, β)`. For this domain size the key is a full
truth-table share (`N = 256` words), not the paper's matching-vector
packing. PIR is `eval_it_dpf3_inner_product` on each server; the three
dots sum to the record. See [it_dpf3.hpp](@ref dpf/it_dpf3.hpp) and
[Three-server PIR](@ref app_pir3).
Socket (2+1) keygen lives in `party/dist_ds.hpp`. The two-party peer that
replaces the pad dealer is IKNP (`party/iknp_deal.hpp`,
`dist_with_*_iknp`): same Doerner–Shelat walk after Jack Doerner and abhi
shelat, CCS 2017 ([ePrint 2017/827](@ref bib_ds)), with pads from `dpf::iknp::sample`
(Ishai, Kilian, Nissim, and Petrank, CRYPTO 2003; Chou–Orlandi base OT,
[ePrint 2015/267](@ref bib_chou)). There is no two-party `(2,3)` Shamir path —
`dist_dpf3` needs three key holders. Fig-10 `updatable` payload rewrites
also stay on `dist_dpf3`.
**IKNP vs dealer vs Half-Tree (costs).** With input bit length `n` and
seed width `λ = 128`, a reveal point tape has length `T = Θ(n)`
(`ncw = n`, `nblock = 2n + n_leaf`). An oblivious (non-reveal) tape adds
`n · 40960` bit×block pads because each level runs the Boyar–Peralta
32-AND S-box hash ([ePrint 2011/332](@ref bib_boyar); see `hash_level_and_count()`).
- Dealer `make_dpf`: 0 rounds, 0 bytes, `Θ(n)` AES expands.
- DS + p2 (`dist_with_*`): dealer sends `Θ(n λ)` bits of pads offline;
online is `n` rounds and `Θ(n λ)` bits of peer opens (guided tour
[Doerner–Shelat](@ref tour_ds)).
- Half-Tree §5.2 ([ePrint 2022/1431](@ref bib_halftree)): `n+3` rounds in the COT/OLE hybrid
with no beaver-pad dealer. This library's IKNP path does **not** use
that hybrid; it keeps the per-level DS open after OT-sampled pads.
- IKNP + DS: two Chou–Orlandi sessions of `κ = 128` base OTs, then two
OT-extension directions whose U-matrix and correction traffic is
`Θ(κ T)` bits, then the same `n` DS opens as the dealer walk (no p2
frames). Local work adds `Θ(κ)` P-256 scalar muls and the AES column
expands of the extension.
One localhost measurement (`party_bench --case iknp_geneval_point
--repeat 1 --warmup 0`, `uint8` so `n = 8`, reveal): ≈ 1.08 s wall;
about 10–16 KiB per direction on the p0–p1 link; harness `rounds = 47`
and `prg_evals = 2431`. That is a measurement of this harness, not a
claimed speedup over dealer keygen or over Half-Tree §5.2.
**Code samples**\n
<div class="tabbed">
- <b class="tab-title">eval_dpf3_point.cpp</b> \include{cpp} evaluation/eval_dpf3_point.cpp
- <b class="tab-title">eval_dpf3_doerner_shelat.cpp</b> \include{cpp} evaluation/eval_dpf3_doerner_shelat.cpp
- <b class="tab-title">eval_dpf3_cmp_ic.cpp</b> \include{cpp} evaluation/eval_dpf3_cmp_ic.cpp
</div>
\htmlonly
<div class="tldr"><b>TL;DR.</b> eval_point is one path. An interval costs the path plus the length of the range. A full domain costs the size of the domain; an inner product does that walk and keeps only the accumulator. Wildcard assign and an updatable rewrite patch the leaf and do not depend on the depth. Three-party eval is two walks and a short field open. The information-theoretic key is a 256-word share, not a walk.</div>
\endhtmlonly