libdpf/doc/pages/evaluation.md
Ryan Henry 0d22946a0e Checkpoint the party/runtime stack before share-program and malicious-mode work.
Ship the TLS mesh, composer, Beaver/Yao/leaf MPC, prep/online paths, apps, and docs so the tree is pushable before elevating share_expr, security_mode, and prep resume.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-28 05:59:19 -06:00

29 KiB
Raw Blame History

On this page: [wildcard assign](@ref wildcard_assign), [memoizers](@ref memoizers), [buffers](@ref output_buffers), [point](@ref eval_point), [interval](@ref eval_interval), [full domain](@ref eval_full), [deferred input](@ref defer_eval), [inner product](@ref eval_inner_product), [sequence](@ref eval_sequence), [three-party](@ref dpf3), [IT 3-server](@ref it_dpf3).

make_dpf(x, y) returns one key per party. Evaluation of a key yields that party's share. Leaf outputs are subtractive shares: open them with dpf::reconstruct, which computes share0 - share1. Comparison outputs are additive shares: reconstruct computes share0 + share1. A single-output eval_point returns a small handle; *handle is the share. An unassigned dpf::wildcard output throws std::runtime_error.

[from, to] is inclusive. The points passed to eval_sequence are a nondecreasing range; an unsorted range throws std::runtime_error.

Let n be the input bit length and λ the PRG seed width (128 bits for prg::aes128). A make_dpf key stores one λ-bit root and, on each of the n levels, one λ-bit correction word and an advice bit, plus one leaf per output slot. Key size is Θ(n λ) bits plus those payloads. That layout is Boyle, Gilboa, and Ishai, CCS 2016 (full version [ePrint 2018/707](@ref bib_fss2018)), whose stated length is n(λ+2)+λ+⌈log2|G|⌉ bits, not their [EUROCRYPT 2015](@ref bib_fss2015) key of 4n(λ+1) bits. For a small output group G, Remark 3.4 of the full version stops ν = log2(λ / log2|G|) levels early and shortens the key by ν(λ+2) bits. This generator does that: depth is n - lg(outputs_per_leaf), where outputs_per_leaf is how many copies of G fit in one λ-bit block. Those low input bits select the lane. A comparison still writes a value word on each remaining level. The default expand is that CCS 2016 tree. Guo, Yang, Wang, Zhang, Xie, Zhang, and Liu ([ePrint 2022/1431](@ref bib_halftree)) keep the same key length and the same n-hash point evaluation, and state about 2n+2 random-permutation calls for key generation versus about 4n, and 1.5N calls for a full-domain evaluation versus 2N. prg::aes128_ccr selects that Half-Tree expand. Keygen expands once per level. at<N> and idpf add leaves; the spine is still n levels. A comparison adds one payload word per level. block_width<B> stores one word every B levels instead, and the seed spine stays Θ(n λ). That is ahead of Elette Boyle, Nishanth Chandran, Niv Gilboa, Divya Gupta, Yuval Ishai, Nishant Kumar, and Mayank Rathee (EUROCRYPT 2021, [ePrint 2020/1392](@ref bib_dcf)), who publish a value word on every level: about B times fewer payload words, paid for by expanding the siblings between checkpoints at evaluation. See the [bibliography](@ref bibliography).

eval_point expands those n levels. A basic path memoizer keeps the n nodes (Θ(n λ) bits) and the next call expands only the suffix after the common prefix. The nonmemoizing memoizer keeps one node. reconstruct of two shares, or of two or three Shamir shares, is a constant amount of group arithmetic.

eval_point<I> and eval_point(dpf::out<I>, key, x) read output slot I. dpf::out<I, W> is that slot with prefix length W checked against the key. A bare leaf's prefix is the input bit length. dpf::at<N> uses N. dpf::prefix_deduce (the default of out<I>) reads the prefix from the key. eval_point(dpf::cmp, key, x) reads the comparison channel. eval_point(dpf::cmp_prefix<L>, key, x) reads the first L bits of an idcf key. The [eval_point](@ref eval_point) section has the calls.

On an idpf / multilevel key, eval_prefixes(out<I,N>, key) walks from the root and materializes all 2^N prefix shares (Poplar-style). idpf_eval_ctx / eval_until ([ePrint 2021/017](@ref bib_poplar)'s EvaluateUntil) resume under a live prefix list so the saved node count stays O(|prefixes|), not O(2^N). The I-DPF max / k-th walk in [dpf/idpf_agg.hpp](@ref dpf/idpf_agg.hpp) ([ePrint 2024/1190](@ref bib_idpfagg)) is that caller. See [eval_until](@ref dpf/eval_until.hpp) and the application mockup [I-DPF max and k-th](@ref app_idpf_agg).

Assigning a wildcard leaf

\htmlonly

ELI5. The blank leaf is a Beaver slot from keygen. Filling it in rewrites the correction word: the parties exchange one blinded share of the new payload and both apply the same patch. assign_cmp rewrites the n comparison words locally and sends nothing. An updatable leaf is the same patch later, O(λ) and independent of the depth.
\endhtmlonly

An output wildcard is a placeholder for a payload filled after keygen. The type and the dpf::wildcards names are on [Output types](@ref output_types). Evaluation of an unassigned slot throws std::runtime_error.

In one process, each party holds a share of the payload and the two keys. Slot I is std::get<I>(key.leaf_nodes) (party_key is the key). The exchange is three calls:

\code{cpp} auto & w0 = std::get<0>(k0.leaf_nodes); auto & w1 = std::get<0>(k1.leaf_nodes); auto b0 = w0.compute_and_get_blinded_output_share(share0); auto b1 = w1.compute_and_get_blinded_output_share(share1); auto l0 = w0.compute_and_get_leaf_share(b1); auto l1 = w1.compute_and_get_leaf_share(b0); w0.reconstruct_correction_word(l1); w1.reconstruct_correction_word(l0); \endcode

share0 and share1 are the two parties' shares of the payload. compute_and_get_blinded_output_share also accepts a secret_share. A wrapper that is already ready takes begin_update() before the same three calls; that assign adds a difference onto the payload already installed.

Over a socket, key.async_assign_leaf(peer, share, token) (and async_assign_leaf<I>) runs that exchange. The method is present when the library is built with ASIO. That is two rounds: blinded output shares, then leaf shares. Each message is one share of the output (the second round sends the packed leaf). The Beaver triple on the wildcard slot is the preprocessing; keygen stored it. Local time is linear in the leaf width.

assign_cmp on a pair of keys rewrites the n value-correction words in place and sends nothing. assign_cmp_local is that patch on one key once δ and the addend share are already known. An input wildcard is one round: the parties exchange one n-bit offset share and reconstruct locally. See [Input types](@ref input_types).

A wildcard comparison payload is different. Generate with dpf::lt(dpf::wildcard_value<Beta>{}) (or leq / gt / geq), then dpf::assign_cmp(k0, k1, if_true, if_false) on the pair. Eval before that throws. dpf::assign_cmp_local(key, delta, addend_share) is the one-key form: both parties pass the same public delta and additive shares of the absorb target.

\code{cpp} auto [k0, k1] = dpf::make_dpf(std::uint8_t{40}, dpf::lt(dpf::wildcard_valuestd::uint64_t{})); dpf::assign_cmp(k0, k1, std::uint64_t{7}); auto y0 = dpf::eval_point(dpf::cmp, k0, std::uint8_t{10}); \endcode

An input wildcard masks the index. The calls are offset_x.compute_and_get_share and offset_x.reconstruct on [Input types](@ref input_types). Eager eval_interval / eval_full require that offset to be ready. To expand before assign (full-domain identity, then rotate once δ is known), use [Deferred input evaluation](@ref defer_eval).

Defined in\n @ref dpf/leaf_wrapper.hpp, @ref dpf/dpf_key.hpp, @ref dpf/incremental.hpp

Memoizers hold interior nodes between calls. Output buffers hold the shares a multi-point evaluation writes. Pass both as mutable named objects when a later call should reuse them. The factories make_basic_path_memoizer, make_basic_interval_memoizer, and make_*_sequence_memoizer unwrap party_key, so a workspace built from either party's type accepts both parties. Name that type with dpf::unwrap_party_key_t<std::decay_t<decltype(key)>>.

Memoizers

\htmlonly

ELI5. The first walk stores interior nodes. The next query starts from the deepest stored node that still lies on its path, instead of from the root. An interval memoizer stores the nodes that cover a range; a sequence memoizer stores the nodes along a sorted list.
\endhtmlonly

Path memoizers

eval_point walks one root-to-leaf path. make_basic_path_memoizer<Key>() keeps every node of the previous point, and the next point recomputes only the suffix after the common prefix. make_nonmemoizing_path_memoizer<Key>() keeps one node and starts from the root on every call. A one-off eval_point(key, x) uses the nonmemoizing memoizer.

Pass the memoizer as a mutable lvalue. The default argument is a new temporary, so it has no previous point to resume from. Keep a separate memoizer for each key you are in the middle of evaluating. A different root restarts the path. One point evaluation is Θ(n) expands. The basic memoizer's extra memory is the n saved nodes.

Code samples\n

  • memoizers.cpp \include{cpp} evaluation/memoizers.cpp

Interval memoizers

eval_interval and eval_full expand every leaf in a range. make_basic_interval_memoizer<Key>(from, to) stores two levels of that range. That is the workspace the convenience overloads allocate. make_full_tree_interval_memoizer<Key>(from, to) keeps every level. make_basic_full_memoizer<Key>() and make_full_tree_full_memoizer<Key>() are the same workspaces sized for the whole domain.

Size the memoizer for the widest interval you will pass to it. A wider interval throws std::length_error. The same key and the same endpoints leave the final interior level in place. A different key or a different interval rebuilds into the same allocation.

Passing only the memoizer still allocates a fresh output buffer and returns std::pair(buffer, iterable).

Let L be the number of inputs in the inclusive interval. The node count at a level is the width of that interval shifted up to the level, and the walk expands each of those nodes once, Θ(n + L) expands in total. The output buffer is L slots per selected output. A basic memoizer keeps two levels, O(L) nodes. A full-tree memoizer keeps every level, still Θ(n + L) nodes. eval_full is the same walk with L = 2^n: time and the output buffer are Θ(2^n) slots, and the basic full memoizer holds two levels of that tree.

Sequence memoizers

make_sequence_recipe<Key>(begin, end) compiles a sorted point list into a traversal. The recipe depends on the input type, and one recipe serves every key of that type.

A sequence memoizer stores a reference to the recipe object it was built from and checks later calls by address. Pass that same object, and keep the recipe alive for as long as the memoizer is used. A copy of the recipe throws std::logic_error.

make_double_space_sequence_memoizer<Key>(recipe) keeps two levels. It is what eval_sequence(key, recipe, buffer) allocates when you omit the memoizer. make_inplace_reversing_sequence_memoizer<Key>(recipe) keeps one level and reverses direction as it descends. make_full_tree_sequence_memoizer<Key>(recipe) retains every level. A key whose depth differs from the recipe throws std::logic_error.

Let m be the number of listed points. make_sequence_recipe makes one pass per level and binary-searches each live block, so the build is O(n m log m) comparisons in the worst case. The recipe stores one step per block per level, O(n m) records when every level splits. Evaluating it expands one node per step, at most O(n m) expands, and writes m shares. The double-space memoizer keeps two levels of that traversal (O(m) nodes at the widest level). The inplace memoizer keeps one level. The full-tree memoizer keeps every level.

Output buffers

output_buffer<T> is move-only storage with size, iterators, data, and operator[]. Build it with the factory that matches the evaluation:

  • make_output_buffer_for_interval(key, from, to)
  • make_output_buffer_for_full(key)
  • make_output_buffer_for_subsequence(key, begin, end, tag)
  • make_output_buffer_for_recipe_subsequence(key, recipe, tag)

On a party_key, leaf slots are subtractive_shares and comparison slots are additive_shares. dpf::bit, dpf::twobit, and dpf::nyble slots are packed. Trivially default-constructible slot types are left uninitialized; the evaluation overwrites every slot it is responsible for.

eval_interval and recipe eval_sequence take the buffer as a non-const reference, so the argument is a named object. The returned iterable refers into that buffer. Read it while the buffer is alive, and only over the points the iterable covers. The next evaluation overwrites those slots. The buffer holds one slot per input in the interval, or one slot per listed point. Its space is that many slots times the slot width. Packed bit, twobit, and nyble slots share words.

For one output, the convenience overload returns the buffer itself as the first element of the pair. For several output indices it returns a tuple of buffers.

make_output_buffer(dpf::out<I>, key, from, to) and make_output_buffer(dpf::cmp, key, n) size a buffer for one channel of a multi-output or comparison key. The slot types follow the same party-share rule.

Code samples\n

  • output_buffers.cpp \include{cpp} evaluation/output_buffers.cpp

dpf::eval_point

eval_point(key, x) evaluates output 0 at one input. eval_point<I>(key, x) selects another output. Two or more indices, eval_point<0, 1>(key, x), return a tuple of shares rather than handles. eval_point(key, x, path) continues a path memoizer.

eval_point(dpf::out<I>, key, x, path) and eval_point(dpf::out<I, W>, key, x, path) are the same walk with an explicit slot. W must equal that slot's prefix. eval_point(dpf::cmp, key, x, path) reads the comparison channel. eval_point(dpf::cmp_prefix<L>, key, x, path) reads an idcf prefix of L bits. Comparison results are additive shares. An unassigned wildcard comparison throws until assign_cmp; see [Assigning a wildcard leaf](@ref wildcard_assign).

\code{cpp} auto [k0, k1] = dpf::make_dpf( std::uint8_t{0x2a}, dpf::at<4>(std::uint8_t{5}), std::uint8_t{9}); auto hi = *dpf::eval_point(dpf::out<0, 4>, k0, std::uint8_t{0x2a}); auto leaf = *dpf::eval_point(dpf::out<1, 8>, k0, std::uint8_t{0x2a});

auto [c0, c1] = dpf::make_dpf(std::uint8_t{40}, dpf::idcf(dpf::gt(std::uint64_t{1}))); auto full = dpf::eval_point(dpf::cmp, c0, std::uint8_t{50}); auto pref = dpf::eval_point(dpf::cmp_prefix<4>, c0, std::uint8_t{50}); \endcode

Code samples\n

  • eval_point.cpp \include{cpp} evaluation/eval_point.cpp

dpf::eval_interval

eval_interval(key, from, to) evaluates every input from from through to. The iterable yields one share per input, in that order. Optional arguments are an output buffer and then an interval memoizer. An output index pack, eval_interval<0, 1>(key, from, to, buffers, memo), writes each selected output.

Code samples\n

  • eval_interval.cpp \include{cpp} evaluation/eval_interval.cpp

dpf::eval_full

eval_full(key) is the closed interval from std::numeric_limits<Input>::min() through max(). The buffer and full-domain memoizer overloads match eval_interval. make_output_buffer_for_full(key) and make_basic_full_memoizer<Key>() size both for that domain.

Code samples\n

  • eval_full.cpp \include{cpp} evaluation/eval_full.cpp

Deferred input evaluation

\htmlonly

ELI5. The PRG expand runs once, into a full-domain buffer, before the index is known. When the offset opens, get() rotates that buffer. The tree is not expanded again.
\endhtmlonly

Additive input blinding evaluates in tree coordinates x ↦ x + δ, where δ is reconstructed by assign_wildcard_input into offset_x. Eager eval_interval folds that map into the traversed range and throws if the offset is not ready.

When δ is still unknown, call defer_eval_interval or defer_eval_full:

  1. Expand the full input domain at identity into a full-sized buffer (make_output_buffer_for_full).
  2. After assign, .get() on the returned deferred_rotated_subinterval rotates by offset_x(0) and yields the logical [from, to] (including wrap-around when from > to in domain-walk order).

Interior-only prep (assigned input, leaf still a wildcard) is defer_traverse_interval. Full-domain interior with both still unset is defer_traverse_full.

Keep the deferred object alive for the lifetime of any .get() range. Buffers must outlive both.

Code samples\n

  • defer_eval.cpp \include{cpp} evaluation/defer_eval.cpp

Defined in\n @ref dpf/deferred_rotated_subinterval.hpp, @ref dpf/eval_interval.hpp, @ref dpf/eval_full.hpp

dpf::eval_inner_product

\htmlonly

ELI5. The walk is the same as an interval or full-domain eval, but each leaf is multiplied by a public weight and added into one accumulator. The expanded vector is not stored. A sequence inner product does that only at the listed points.
\endhtmlonly

eval_inner_product multiply-accumulates DPF shares against another vector during the walk. It does not write the output vector.

Three local forms:

  • Batched leaf walk (no tag): eval_inner_product(key, from, to, weights, memo). One output, weights in eval_interval layout, same batched exterior AES as that walk, O(1) accumulator.
  • dpf::paired (row-wise): one weight row per input. A scalar pairs with one output; a tuple / array zips several outputs (leaf slots or an ancestor prefix plus the leaf) off one path. Products are summed.
  • dpf::columns (transposed): one output, several weight streams, one accumulator per stream. Products stay apart. A stream is w[i] or w(i). dpf::project maps the share first; dpf::also sees it unmapped.

A single-stream columns result matches paired on that stream. A single-output batched leaf walk matches paired when the interval is leaf-aligned so covering-leaf weights equal the clipped domain points. Unaligned intervals still weight every lane of the covering leaves (same layout as one-key eval_interval buffers). Cohort evaluation uses a separate interleaved leaf layout (cohort_index); interleave_leaves builds that order from per-key buffers (see dpf/interleave_leaves.hpp and dpf/cohort.hpp).

Pass dpf::paired. One element of the other vector is consumed per input, in the same order as eval_interval or eval_sequence.

A scalar element pairs with one output:

eval_inner_product(dpf::paired, key, from, to, weights);
eval_full_inner_product(dpf::paired, key, weights);

A std::tuple or std::array element pairs componentwise with the output indices in the template pack. Those outputs may be several slots on the same leaf, or an ancestor prefix slot and the leaf. Both are read from one path:

eval_inner_product<0, 1>(dpf::paired, key, from, to, rows);
eval_sequence_inner_product<0, 1>(key, begin, end, rows);
eval_sequence_inner_product<0, 1>(key, recipe, begin, end, rows);

share * component must be defined. The product type must support +. The walk costs the same as the interval or sequence it follows. Each input adds a constant amount of arithmetic per paired output or per column. The accumulator is the only result; there is no output vector of shares. The point list for a sequence or recipe is sorted nondecreasing, same as eval_sequence. The recipe overload also checks that the list length matches the recipe.

dpf::columns is the same walk with the products kept apart. One output, several weight streams, one accumulator per stream:

auto [sum, dot, sq] = eval_sequence_inner_product<0>(
    dpf::columns, key, begin, end, std::tie(ones, r, r2));

A stream is anything with w[i] or w(i), so a challenge can be a function of the list index instead of a stored vector. i is the position in the point list, or in the interval in eval_interval order. dpf::project(fn) maps the share before the multiply. dpf::also(fn) is called as fn(i, x, share) on the share before that map, which is how a mailbox slot is updated in the same walk. Pass either tag, both, or neither.

The older single-output range walk (no dpf::paired) is still the batched leaf inner product: eval_inner_product<I>(key, from, to, weights, memo).

Code samples\n

  • eval_inner_product.cpp \include{cpp} evaluation/eval_inner_product.cpp

dpf::eval_sequence

eval_sequence(key, begin, end, tag) evaluates a sorted list. dpf::return_output_only_tag_ stores one share per listed point. dpf::return_entire_node_tag_ stores whole leaves; it is the default. The iterable still yields one share per listed point, in list order.

eval_sequence(key, recipe, buffer, memo, tag) repeats that list. memo is a sequence memoizer bound to recipe. Omit memo to allocate a double_space workspace for that call.

eval_sequence_breadth_first(key, begin, end, buffer) writes one share per listed point, expanding a level at a time. The list is still sorted. eval_sequence_breadth_first(dpf::out<I>, key, begin, end) allocates the buffer and returns it. The out<I, W> form checks the prefix the same way eval_point does. The cost is the same order as recipe eval_sequence on that list: at most O(n m) expands and m output slots.

Code samples\n

  • eval_sequence.cpp \include{cpp} evaluation/eval_sequence.cpp

Buffered PRG

\htmlonly

ELI5. One AES expand produces more blocks than a single tree node needs. The buffered PRG keeps the leftover blocks and serves the next nodes from them, so a wide walk makes fewer expands. The keys do not change.
\endhtmlonly

dpf::randomness::buffered_prg<PRG, Ts...> (alias dpf::randomness::aes_buffered_prg<Ts...>) is a forward cursor with one PRG stream per value type. get<I>() and fill<I>(out, n) consume the cursor. at<I>(index) reads an absolute index and leaves the cursor where it is. sampled<I>() is how far get and fill have advanced. per_stream_buffer_elems is at least 1.

dpf::randomness::lane_table<T> is the seekable form for a runtime set of roles. value_at(role, index) and mask_at(role, index) are independent streams, and a repeated index returns the same element.

get and fill of q elements do Θ(q) PRG work. The cursor keeps per_stream_buffer_elems elements per stream. at reads one absolute index and does not move the cursor. A lane_table keeps one cache window per role (the constructor's window, default 256) for the value stream and one for the mask stream. A repeated index returns the same element. fill_values / fill_masks of q elements are Θ(q).

Code samples\n

  • buffered_prg.cpp \include{cpp} evaluation/buffered_prg.cpp

Three-party (2,3) DPF

\htmlonly

ELI5. Each evaluator key is two VDPF+ spines. eval_point walks both, Θ(n) expands, then scales into fp61. Opening two or three as_share values is a constant amount of field arithmetic. An updatable rewrite patches four leaves and refreshes an offset, independent of n.
\endhtmlonly

make_dpf3(α, β) builds three evaluator keys after Guy Zyskind, Avishay Yanai, and Alex "Sandy" Pentland, [ePrint 2024/1658](@ref bib_dpf3), Figure 3: two VDPF+ spines plus Shamir embedding in fp61. eval_point returns a field share; open with dpf::reconstruct on any two (or all three) dpf::as_share values. Tags: verifiable, extractable, updatable (Fig. 10 in-place payload update).

make_dpf3_doerner_shelat(x0, x1, β) is the dual-spine Doerner–Shelat path: XOR shares of α, same clear β. Socket orchestration lives in party/dist_dpf3.hpp (dist_with_dpf3_key): after keygen, role p0 holds party 1, p2 holds party 2, p1 holds party 3. Path bits open to p0/p1; p2 does not learn α. Dist keys are verifiable; payload updates use dealer updatable keys or remake_dpf3.

Comparison / interval: make_dpf3_cmp, make_dpf3_cmp_blocked, make_dpf3_ic. Multipoint: make_multipoint3 (cuckoo buckets of point keys).

make_dpf3 is local dealer work. Each evaluator key is two VDPF+ spines, Θ(n λ) bits with a constant factor over one two-party key. eval_point is two walks, Θ(n) expands, then a field scale. Opening two or three as_share values is a constant number of fp61 operations. An updatable rewrite is four leaf patches and an offset refresh, O(λ) and independent of n. make_dpf3_doerner_shelat is two Doerner–Shelat spines: twice the rounds and the pad tape of one two-party opening (see [Doerner–Shelat](@ref tour_ds)). Comparison, interval, and multipoint forms add the same extra material as the two-party comparison, interval, or cuckoo packing, on top of those spines.

Information-theoretic 3-server DPF

\htmlonly

ELI5. There is no PRG tree. For this domain each key is an additive share of a 256-word table. The three shares sum to beta at alpha and to zero elsewhere. A PIR answer is three inner products with the database; those three dots sum to the record.
\endhtmlonly

make_it_dpf3(α, β) ([ePrint 2023/028](@ref bib_itdpf)) is a different object from make_dpf3. Each of three parties holds an additive share of the characteristic vector on {0..255}; the sum of all three eval_it_dpf3 values is the point function. Any single key is independent of (α, β). For this domain size the key is a full truth-table share (N = 256 words), not the paper's matching-vector packing. PIR is eval_it_dpf3_inner_product on each server; the three dots sum to the record. See [it_dpf3.hpp](@ref dpf/it_dpf3.hpp) and [Three-server PIR](@ref app_pir3).

Socket (2+1) keygen lives in party/dist_ds.hpp. The two-party peer that replaces the pad dealer is IKNP (party/iknp_deal.hpp, dist_with_*_iknp): same Doerner–Shelat walk after Jack Doerner and abhi shelat, CCS 2017 ([ePrint 2017/827](@ref bib_ds)), with pads from dpf::iknp::sample (Ishai, Kilian, Nissim, and Petrank, CRYPTO 2003; Chou–Orlandi base OT, [ePrint 2015/267](@ref bib_chou)). There is no two-party (2,3) Shamir path — dist_dpf3 needs three key holders. Fig-10 updatable payload rewrites also stay on dist_dpf3.

IKNP vs dealer vs Half-Tree (costs). With input bit length n and seed width λ = 128, a reveal point tape has length T = Θ(n) (ncw = n, nblock = 2n + n_leaf). An oblivious (non-reveal) tape adds n · 40960 bit×block pads because each level runs the Boyar–Peralta 32-AND S-box hash ([ePrint 2011/332](@ref bib_boyar); see hash_level_and_count()).

  • Dealer make_dpf: 0 rounds, 0 bytes, Θ(n) AES expands.
  • DS + p2 (dist_with_*): dealer sends Θ(n λ) bits of pads offline; online is n rounds and Θ(n λ) bits of peer opens (guided tour [Doerner–Shelat](@ref tour_ds)).
  • Half-Tree §5.2 ([ePrint 2022/1431](@ref bib_halftree)): n+3 rounds in the COT/OLE hybrid with no beaver-pad dealer. This library's IKNP path does not use that hybrid; it keeps the per-level DS open after OT-sampled pads.
  • IKNP + DS: two Chou–Orlandi sessions of κ = 128 base OTs, then two OT-extension directions whose U-matrix and correction traffic is Θ(κ T) bits, then the same n DS opens as the dealer walk (no p2 frames). Local work adds Θ(κ) P-256 scalar muls and the AES column expands of the extension.

One localhost measurement (party_bench --case iknp_geneval_point --repeat 1 --warmup 0, uint8 so n = 8, reveal): ≈ 1.08 s wall; about 10–16 KiB per direction on the p0–p1 link; harness rounds = 47 and prg_evals = 2431. That is a measurement of this harness, not a claimed speedup over dealer keygen or over Half-Tree §5.2.

Code samples\n

  • eval_dpf3_point.cpp \include{cpp} evaluation/eval_dpf3_point.cpp

  • eval_dpf3_doerner_shelat.cpp \include{cpp} evaluation/eval_dpf3_doerner_shelat.cpp

  • eval_dpf3_cmp_ic.cpp \include{cpp} evaluation/eval_dpf3_cmp_ic.cpp

\htmlonly

TL;DR. eval_point is one path. An interval costs the path plus the length of the range. A full domain costs the size of the domain; an inner product does that walk and keeps only the accumulator. Wildcard assign and an updatable rewrite patch the leaf and do not depend on the depth. Three-party eval is two walks and a short field open. The information-theoretic key is a 256-word share, not a walk.
\endhtmlonly