libdpf/doc/pages/applications.md
Ryan Henry 0d22946a0e Checkpoint the party/runtime stack before share-program and malicious-mode work.
Ship the TLS mesh, composer, Beaver/Yao/leaf MPC, prep/online paths, apps, and docs so the tree is pushable before elevating share_expr, security_mode, and prep resume.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-28 05:59:19 -06:00

32 KiB
Raw Blame History

Application mockups

These programs are the DPF step of a larger protocol. Each one is a single process. A dealer stands in where the paper generates keys from shares. They compile from the repository root:

c++ -std=c++17 -march=native -pthread -I include -I thirdparty examples/applications/duoram3.cpp

The online compose sketch at the end of each file uses [dpf::app::run_measured](@ref dpf/app_flow.hpp): it prints rounds, live bytes, wall/CPU, PRG evals, random bytes, and the master seed hex. Set DPF_EXPERIMENT_DIR=/tmp/run to also write the CSV tables described in [Experiments](@ref experiment_costs). The same compile line with the other filenames builds the rest. Timing for the party protocols, including word garbling, stacked branches, FLUTE, and the hidden column reorder, is party_bench / profile_party (see [The battery](@ref experiment_costs)).

Optional Python bindings configure with -DLIBDPF_PYTHON=ON in the test build directory, then make pydpf (and make pydpf_pytest). pydpf exposes point / interval / full / sequence / recipe eval on uint8→uint64, two-leaf multileaf, wildcard assign, 16-bit eval_until, and it_dpf3. Not in this module yet: geneval, Doerner–Shelat, VDPF prove/sketch, DCF, or grotto. What those programs had to do by hand is [the library surface underneath](@ref application_gaps).

Sketch What the DPF step is
[3-party Duoram](@ref app_duoram) Unit key, rotate, inner product
[SUBLEQ](@ref app_subleq) Instruction fetch under a secret PC
[BitMore](@ref app_bitmore) 2^L servers, one hot slot
[Keyword PIR](@ref app_keyword) Cuckoo buckets over keywords
[Prio](@ref app_prio) Heavy-hitter prefix walk
[I-DPF max / k-th](@ref app_idpf_agg) eval_until order statistics
[LLAMA](@ref app_llama) Lookup and truncation
[Pika](@ref app_pika) Comparison tree
[Express](@ref app_express) Authenticated memory
[PRAC](@ref app_prac) Range and prefix counts
[Splinter](@ref app_splinter) Function secret-sharing queries
[Mastic](@ref app_mastic) Aggregation
[Waldo](@ref app_waldo) Private search
[Sabre](@ref app_sabre) Robust aggregation
[(2,3) ledger](@ref app_ledger23) Three-party update
[PSI](@ref app_psi) Private set intersection
[Range count](@ref app_range_count) Interval payload
[Floram](@ref app_floram) ORAM read
[Three-server PIR](@ref app_pir3) Information-theoretic DPF
[Protocol composition](@ref protocol_compose) Fused walks, early-stop, RSS refresh, ABY scale
PIRsona fetch BitMore star upload+answer (pirsona_fetch.cpp)
Hushmap ADD Dealer tape + two opens (hushmap_add.cpp)
[What the walk folds in](@ref application_gaps) Library surface those programs used to do by hand

3-party Duoram

\htmlonly

ELI5. Preprocessing plants a unit DPF at a random index r. Online the parties open i* − r, rotate the expanded unit vector by that public shift, and dot with the memory. The update rotates a payload vector the same way and adds it in.
\endhtmlonly

Vadapalli, Henry, and Goldberg ([USENIX Security 2023](@ref bib_duoram)) keep a memory in shares and read or add at a secret index. Preprocessing builds unit DPFs at a random index r. Online, the parties open i* - r and cyclic-shift the expanded vector. The read is the dot product of that vector with the memory. The update adds a payload vector, shifted the same way. Opening the whole memory, rather than one index, is a hidden shuffle of that column ([an array of shares](@ref share_shuffle)), not another DPF.

The program uses one dealer unit key for the read and one payload key for the update. eval_full_inner_product(dpf::paired, key, memory, dpf::rotate{s}) dots the unit key with the memory read at (i + s) mod n, so the caller keeps one unrotated vector. The update expands straight into the subtractive memory shares with eval_full_add_into(mem, key, dpf::rotate{s}) — no separate expansion, shift, and add. A 1-bit leaf at the same index lifts to +1 or -1; its sign share is recorded at keygen with dpf::unit_sign. A word payload of 1 opens to +1.

This is a (2,2) key. The auxiliary party holds some of those keys. [dpf::make_dpf3](@ref dpf/dpf3.hpp) shares one point Shamir-style among three evaluators.

\include{cpp} applications/duoram3.cpp

MPC SUBLEQ

\htmlonly

ELI5. The address is not known when the keys are built, so the unit vectors are expanded early into a deferred buffer. Online, opening the address rotates that buffer. The read is a dot with memory; the write adds the scaled unit vector back.
\endhtmlonly

Jiang and Henry ([MSc thesis, University of Calgary](@ref bib_subleq)) emulate the subtract-and-branch-if-less-than-or-equal-to-zero (SUBLEQ) OISC for private function evaluation. One instruction is

D[B] -= D[A]
pc    = (D[B] ≤ 0) ? C : pc + 3

The DPF work is prepaid. The dealer ships wildcard unit keys [[* | 1]] for the addresses that will be read (and for writing B). Each party expands those keys over the full address domain with defer_eval_full before the addresses are known — that is all of the PRG. Online, the parties open each address into offset_x, and .get() on the deferred view rotates the prepaid unit vector. A read is the Du-Atallah / local dot of that view with the memory shares. The write reuses the same e_B view: add (-D[A]) · e_B into the memory shares (a Beaver multiplication in the protocol; the listing opens the scale). The branch is a path evaluation only — make_dpf(0, dpf::leq(1)) evaluated at x = D'[B] — never a full-domain expand of the word domain. The full protocol can instead assign a wildcard comparison key to x and evaluate at public 0 with geq (same bit).

Instruction fetch is the same prepaid unit dotted against three sliding windows of D (the three addresses of the instruction). The listing starts after (A, B, C) are already shares.

An equivalent prepaid form plants the unit at a random r and opens addr - r, then folds the shift with dpf::rotate{s} on eval_full_inner_product / eval_full_add_into (the Duoram / Pika surface). Same online AES; only the blinding convention differs.

Still by hand: Du-Atallah blinds, the ABY2 mux of C against pc+3, and out-of-bounds prefix-parity checks from the thesis.

\include{cpp} applications/subleq.cpp

BitMore, 2^L servers

\htmlonly

ELI5. The label is L bits. Each bit is its own 1-bit DPF, expanded over the whole domain. Stacking the L bit-vectors and reading a column produces the 2^L-server answer for that label.
\endhtmlonly

Hafiz and Henry ([PoPETs 2019](@ref bib_bitmore), §5.2) query ell = 2^L servers with L independent 1-bit DPFs, all at the same row. Server j receives key number j_e from DPF e. dpf::pack_bit_columns(keys...) runs the full-domain bit walk once per key and writes lane e = key e into one integer per row, so the server reads symbol[row] instead of unpacking one int per bit. Off the secret row every server computes the same digit. On the secret row the digits are symbol(0) XOR j, a permutation of 0 .. ell-1. When the server count ℓ is not a power of two, dpf::mod_bit_columns<ℓ>(keys...) is that same integer modulo ℓ. The first key is still the low bit. The information-theoretic virtual-bucket response starts from those digits.

L = 1 is the two-server member of the same family. Each party XORs the records its bit selects. The two answers XOR to the record.

\include{cpp} applications/bitmore.cpp

Keyword PIR

\htmlonly

ELI5. The keyword is hashed into cuckoo buckets. Each bucket is a point key, and the record is the inner product of the probed buckets with the dictionary. The S&P 2025 seed-packing of those buckets is not what this program does.
\endhtmlonly

Gilboa and Ishai ([EUROCRYPT 2014](@ref bib_dpf2014)) retrieve one record by a keyword. [dpf::keyword](@ref dpf/keyword.hpp) is the domain, so the DPF point is the keyword itself. Each server walks the dictionary with eval_sequence and XORs the record where its bit share is set. Off the keyword the two bits match, so the record cancels. A keyword absent from the dictionary opens to 0. eval_sequence takes the dictionary in nondecreasing order.

A batch of keywords that should come back as separate records is one key per keyword. [dpf::make_multipoint](@ref dpf/multipoint.hpp) adds several secret points into one output. That is a histogram of hits, or a payload the client chose. That packing is de Castro–Polychroniadou (EUROCRYPT 2022): t points into m ≈ O(t) cuckoo buckets, each an ordinary GGM key. S&P 2025 (Boyle, Gilboa, Hamilis, Ishai, Tu) shrinks the dealer message with a silent OLE / PCG that expands a short seed into many correlated bucket keys. That is a different primitive (the same family as silent VOLE), not a new multipoint_params field, and this library does not ship a PCG stack.

\include{cpp} applications/keyword_pir.cpp

Prio and the heavy-hitter prefix walk

\htmlonly

ELI5. A one-hot vote is a unit DPF in the Prio field. Each server adds the expanded vector into a running histogram. The heavy-hitter pass is the same key read at successive prefixes: the servers add the opened prefix shares and keep the heavy nodes.
\endhtmlonly

Corrigan-Gibbs and Boneh ([NSDI 2017](@ref bib_prio)) aggregate client encodings. A frequency count is a one-hot vector. A unit DPF is that vector, compressed. Each client keys a point at its bin with payload 1 in [dpf::field64](@ref dpf/field64.hpp), the same prime as libprio Field64. Each server folds a full-domain expansion into its running histogram share with eval_full_add_into(hist, key). The opened bin is the count. Prio's validity proof for a general encoding is a SNIP. This program is the DPF encoding only.

The prefix walk is Poplar (Boneh, Boyle, Corrigan-Gibbs, Gilboa, and Ishai, [ePrint 2021/017](@ref bib_poplar)). [dpf::idpf](@ref dpf/placement.hpp) plants a 1 on prefix lengths 1, 2, and 3, counting from the high bit. Two clients hold 0xA0 and 0xB0. They share the length-3 prefix 101 and split at the next bit. The servers evaluate out<i, length> at one representative of each node and add the opened values.

\include{cpp} applications/prio.cpp

I-DPF max and k-th

\htmlonly

ELI5. Each secret integer is one incremental DPF with a unit payload on every prefix. At each bit the servers open the two children. Max keeps the child that holds mass. The k-th keeps the 1-child when its count covers k, and otherwise descends the 0-child with k reduced.
\endhtmlonly

Cheng, Mitrokotsa, Zhang, and Hartmann ([ePrint 2024/1190](@ref bib_idpfagg)) aggregate secret values with an incremental DPF. Communication tracks the bit length of the domain, not how many secret inputs were summed. Each secret uint16_t is one [dpf::idpf](@ref dpf/placement.hpp) with a unit payload on every prefix length. Servers resume only the live prefixes with [dpf::eval_until](@ref dpf/eval_until.hpp). At each depth they open the two children: max keeps the nonempty 1-child; the k-th largest takes the 1-child when its count is at least k. idpf_agg_max / idpf_agg_kth in [dpf/idpf_agg.hpp](@ref dpf/idpf_agg.hpp) share that walk with the gtest.

\include{cpp} applications/idpf_agg.cpp

LLAMA

\htmlonly

ELI5. The comparison is the gate. eval_point on a gt or lt key opens to the payload when the public query is on the true side of the secret, including across the sign bit, because signed inputs flip the high bit before the walk.
\endhtmlonly

Gupta, Kumaraswamy, Chandran, and Gupta ([ePrint 2022/793](@ref bib_llama)) evaluate a nonlinear gate from a dealer key and one opened masked input x_hat = x + r. The sign test is a comparison: make_dpf(knot + r, dpf::gt(1)), evaluated at x_hat. A degree-0 spline is one [dpf::ic](@ref dpf/interval.hpp) key per piece. The piece value is the payload. The same comparison on int8_t follows numeric order: 100 > -3 opens to 1.

\include{cpp} applications/llama.cpp

Pika

\htmlonly

ELI5. The dealer keys a unit DPF at a fresh r and the parties open x = r − a. Rotating the public table by x and dotting with the DPF reads the entry at a. The sign of an early-stop bit leaf is recorded at keygen, so the evaluators never open r.
\endhtmlonly

Wagh ([PoPETs 2022](@ref bib_pika), Fig. 1) looks up Func(a) in a table of a bounded domain. The dealer keys a unit DPF at a fresh index r and the parties open x = r - a. The inner product of the DPF with the table read at (i - x) mod n is the table entry at a; dpf::rotate{s} folds that offset into the walk. A word payload of 1 opens to +1. The paper's early-stop bit leaf opens to +1 or -1; the dealer records that sign at keygen with dpf::unit_sign (the final control bit Gen sees), so the evaluators never open r. A networked early-stop walk (BGI Remark 3.4) drops ν CW rounds with fss_point_early_stop on a [composer](@ref protocol_compose).

\include{cpp} applications/pika.cpp

Express

\htmlonly

ELI5. A mailbox write is a full-domain add of one DPF into the shared array. Every box is touched by the expand; only the programmed box survives when the shares are opened.
\endhtmlonly

Eskandarian, Corrigan-Gibbs, Zaharia, and Boneh ([USENIX Security 2021](@ref bib_express), §3.1) write one mailbox. Two servers hold subtractive shares of the mailboxes. The client sends one key each. Each server adds its expansion into its share. The owner opens one address. An [extractable](@ref dpf/verifiable.hpp) sketch accepts the honest write and rejects a second hot mailbox.

The message is [dpf::fp61](@ref dpf/fp61.hpp) because that is the sketch field. Each server folds the expansion into its mailbox shares and the one-hot audit in one walk: eval_full_add_into(box, key, dpf::sketch(s, r)). The extractable full-domain leaf now matches point evaluation on every lane, so this one call replaces the earlier eval_point-per-address loop and the separate sketch_fold pass.

On a networked walk the audit share rides in the last correction-word flush — fss_point_fused / level_walk_fused on a [composer](@ref protocol_compose) — so the sketch does not add a round. The same shape is Sabre's proof token.

\include{cpp} applications/express.cpp

PRAC

\htmlonly

ELI5. Binary search needs a unit vector on a stride that grows by one bit per comparison. One incremental key holds all of those prefixes, and a prefix inner product dots a stride without a separate point key per slot. A heap update is a three-lane vector at the parent and its two children.
\endhtmlonly

Sasy, Vadapalli, and Goldberg ([ePrint 2023/1897](@ref bib_prac)) run dynamic data structures on a 3-party Duoram. The new DPF shapes are an incremental key and a wide leaf.

Binary search on a sorted array reads lg n locations. The accessible set at each depth is a public stride, and the index in the next stride is the previous index with one comparison bit appended. One [idpf](@ref dpf/placement.hpp) supplies every stride's unit vector. The program's array is [1, 3, 5, 7, 9, 11, 13, 15] and the needle is 10. The public midpoint is index 3. eval_prefix_inner_product(out<i, length>, key, stride) walks to the prefix depth once and dots the 2^length prefix shares with the stride, replacing one eval_point per stride slot. The prefix of length 1 selects 11 from the stride {1, 5}. The prefix of length 2 selects 9 from {0, 2, 4, 6}. The bits 101 are the lower bound, index 5.

Heapify reads a parent and its two children, three strides at one index, with one unit key. The update at that index is a [dpf::vec](@ref dpf/vec.hpp) of three lanes. An incremental wide key is idpf of those vectors. That call already evaluates.

The dealer in this program knows the path and keys it up front. The protocol appends each comparison bit after the key exists.

\include{cpp} applications/prac.cpp

Splinter

\htmlonly

ELI5. The secret WHERE value is a unit DPF. The server dots it with a column that was already summed by attribute, which is the grouped SUM, and with an all-ones column, which is the COUNT. The queried attribute is not revealed.
\endhtmlonly

Wang, Yun, Goldwasser, Vaikuntanathan, and Zaharia ([NSDI 2017](@ref bib_splinter)) answer private queries on public data with two-server FSS. The client's private WHERE value is a unit DPF at that attribute. Each server dots the selector with a public column pre-aggregated by attribute, so eval_full_inner_product(dpf::paired, key, group_sum) is SELECT SUM(value) WHERE attribute = secret, and the all-ones column is COUNT(*). Neither server learns the queried key.

Still by hand: MAX / TOP-k and multi-predicate conjunctions are not one selector DPF; Splinter composes several FSS instances for those.

\include{cpp} applications/splinter.cpp

Mastic

\htmlonly

ELI5. This is the Poplar prefix walk with a weight instead of 1 on every prefix. Servers sum the prefix shares across clients and drop prefixes under the threshold. A path sketch can check that each client programmed a single path.
\endhtmlonly

Mastic (private weighted heavy-hitters and attribute-based metrics) is Poplar's prefix walk with a weight payload. Each client keys an [idpf](@ref dpf/placement.hpp) whose β on every prefix length is its weight, not 1. eval_prefixes(out<i, length>, key) returns the prefix shares; the servers sum them across clients and threshold to keep the heavy prefix. Two clients on 0xA0 and 0xB0 with weights 5 and 3 make prefix 101 heavy with total weight 8. verify_idpf_path<Depth> is the one-time VIDPF path check ([path_sketch.hpp](@ref dpf/path_sketch.hpp)).

\include{cpp} applications/mastic.cpp

Waldo

\htmlonly

ELI5. Each append is a fresh unit DPF added into the value shares; old events are not rewritten. A threshold query is a comparison inner product of those shares with a public magnitude column, so the sum past a secret threshold does not reveal the threshold or the matches.
\endhtmlonly

Dauterman, Rathee, Popa, and Stoica ([S&P 2022](@ref bib_waldo)) build a private time-series database from FSS. The store is append-only: each event is a fresh unit DPF folded into the servers' value shares with eval_full_add_into, never an update. A range/threshold aggregate uses the comparison channel: eval_full_inner_product(dpf::cmp, key, magnitude) dots the per-timestamp gt shares with a public magnitude column over the whole domain, so the SUM over timestamps past a secret threshold reveals neither the threshold nor the matches. The two halves open with reconstruct_cmp_halves.

Still by hand: Waldo's aggregation trees over several predicates, and range endpoints that are themselves secret-shared, compose more than one comparison key.

\include{cpp} applications/waldo.cpp

Sabre

\htmlonly

ELI5. The write is the same full-domain add as Express. The audit folds a constant-size proof token along that key. verify accepts one honest point and rejects a key that was hot in more than one place.
\endhtmlonly

Vadapalli, Storrier, and Henry ([S&P 2022](@ref bib_sabre)) send anonymous messages with a fast audit. The write is Express's full-domain add (eval_full_add_into). The audit is a verifiable DPF proof rather than Express's fp61 sketch: prove_full(key, dpf::prove(π)) folds a constant-size token per party, and dpf::verify(π0, π1) accepts an honest single-point write and rejects the mismatched fold a multi-point key produces. Networked audits pack the proof share into the last CW with fss_point_fused ([protocol composition](@ref protocol_compose)).

Still by hand: Sabre's blame / accountability phase that identifies a cheating client is protocol logic above the DPF proof.

\include{cpp} applications/sabre.cpp

A (2,3) ledger

\htmlonly

ELI5. An append is one verifiable three-party point at the slot. The three proof tokens are checked before the point is added into the slot shares. Any two servers reconstruct a balance.
\endhtmlonly

A replicated ledger held as (2-of-3) shares by three servers, on this group's dpf3 VDPF+ construction. Each append is one make_dpf3(slot, amount, dpf::verifiable{}); a verify_dpf3 over the three prove_dpf3 tokens rejects an append that is not a single well-formed point before it is applied. Each server folds the point into its slot shares with eval_full_add_into(shares, key) (the (2,3) full-domain overload), and any two servers shamir3::reconstruct a balance. Two credits to slot 5 (100 then 7) open to 107.

Still by hand: the transaction / consensus layer around the append (ordering, replay protection) is protocol logic above the DPF step.

\include{cpp} applications/ledger23.cpp

Private set intersection

\htmlonly

ELI5. Each element on one side is a unit DPF. Dotting it with the other side's table is the membership test. The cuckoo layout and the OPRF that would sit around that test are not in the DPF call.
\endhtmlonly

Kolesnikov, Kumaresan, Rosulek, and Trieu ([CCS 2016](@ref bib_kkrt)) test membership with an oblivious PRF. On this domain the PRF is a table both servers hold. Each receiver element is a unit DPF. eval_full_inner_product(dpf::paired, key, table) is that server's share of PRF(y). The sender publishes PRF(x) for each element of their own set. y is in the intersection when the opened tag appears in that image.

Still by hand: cuckoo hashing in the PSI application layer, and the GGM puncture that keeps a large domain at O(n) communication instead of a table. The multipoint cuckoo VDMPF stays the GGM packing above; S&P 2025 DMPF seed packing needs a PCG this library does not provide.

\include{cpp} applications/psi.cpp

Range count

\htmlonly

ELI5. A value falls in [lo, hi) when the greater-than bit at hi and the greater-than bit at lo differ. Each secret value is one comparison key. The count is the sum of those opened bits.
\endhtmlonly

Each secret value is one comparison. The interval [lo, hi) is public. eval_point(dpf::cmp, key, q) opens to 1 when q is strictly above the value, so the two endpoints differ by 1 exactly when lo <= v < hi. The count is the sum of those bits over the values. A neighbor just below lo, and the open end hi, contribute 0.

Still by hand: a secret interval over a public histogram is two comparison inner products (the Waldo aggregate, subtracted). This program is the other direction, secret values and a public range.

\include{cpp} applications/range_count.cpp

Floram

\htmlonly

ELI5. The address is shared, not known to a dealer. The Doerner–Shelat opening builds the unit key level by level, and the read or write is the inner product of that key with the array.
\endhtmlonly

Doerner and shelat ([CCS 2017](@ref bib_ds)) read and write an array at a secret address. Both parties see the memory. The address is XOR-shared, so the key is [dpf::make_dpf_doerner_shelat](@ref dpf/doerner_shelat.hpp) rather than a dealer key. The read dots a unit payload with the array. The write adds a payload key from the same address shares into subtractive copies of the memory, with eval_full_add_into.

Still by hand: Floram's stash, position map, and refresh. This program is the FSS access.

\include{cpp} applications/floram.cpp

Three-server PIR

\htmlonly

ELI5. All three servers hold the database. A Shamir DPF3 key makes each server return an inner product; any two of those field elements open the record. The information-theoretic key instead adds all three inner products.
\endhtmlonly

The database is public and replicated on three servers.

The computational path is one (2,3) point key from [dpf::make_dpf3](@ref dpf/dpf3.hpp) ([ePrint 2024/1658](@ref bib_dpf3)). Each server dots its share with the database: eval_full_inner_product(key, database). Any two of those dots shamir3::reconstruct to the record.

\include{cpp} applications/pir3.cpp

The information-theoretic path is [dpf::make_it_dpf3](@ref dpf/it_dpf3.hpp) ([ePrint 2023/028](@ref bib_itdpf)). Each server holds an additive share of the characteristic vector; the three dots sum to the record. Distinct from make_dpf3.

\include{cpp} applications/it_pir3.cpp

What the walk now folds in

\htmlonly

ELI5. Rotate-then-dot, the sign of a bit leaf, bit columns, prefix dots, and path sketches used to be loops around eval. They are parameters of one walk now.
\endhtmlonly

The calls the eight programs used to build by hand are now the library surface. See [dpf/eval_walk.hpp](@ref dpf/eval_walk.hpp). Multi-protocol schedules (FSS + ABY + RSS on one RoundSink) are [protocol composition](@ref protocol_compose). Each listing under examples/applications/ records that paper's online flow on a composer and calls dpf::app::run, which drives both parties on an in-process sink and prints name rounds= bytes=. That line is the experiment: compare it with the round and bandwidth column of the paper. PIR listings are a client and two or three servers (one upload round, one answer round). The servers do not open shares with each other.

Shift, then add

Duoram and Pika rotate a vector by an opened offset, then dot. dpf::rotate{s} on eval_full_inner_product reads the weight at (i + s) mod 2^n inside the walk, so the caller keeps one unrotated vector. dpf::cyclic_shift(buf, s) rotates a materialized share buffer. Duoram's update, Prio's histogram, and Express's mailbox add the expansion into a buffer the caller already holds with eval_full_add_into(buf, key) (and a dpf::rotate or dpf::sketch overload). Express's audit folds in the same pass.

Still by hand: none for the deferred leaf. dpf::leaf_later expands without the leaf correction word and fills a parallel control-bit buffer. After cyclic_shift_pair (or eval_full_add_into(..., leaf_later, rotate{s})), apply_leaf_correction(buf, control, F) does buf[i] += F * control[i]. [Doerner–Shelat](@ref dpf/doerner_shelat.hpp) already builds a key from index shares the parties hold.

The bit leaf's sign

Duoram's flag vector and Pika's early-stop leaf are 1-bit DPFs that lift to a ring unit of +1 or -1. dpf::make_dpf(alpha, dpf::bit::one, dpf::unit_sign{w0, w1}) records that sign at keygen — w0 - w1 is the ±1 unit — from the final control bit Gen sees and one key hides. Absent the tag, keygen is unchanged.

Answers that are XORs or bit columns

BitMore stacks L full-domain bit vectors and reads them as L-bit digits. dpf::pack_bit_columns(keys...) runs the bit walk once per key and writes one integer per row (lane e is key e). dpf::mod_bit_columns<ℓ>(keys...) reduces that integer modulo a server count ℓ from 2 through 32768, which is the digit when ℓ is not a power of two. The virtual-bucket response that consumes the digits stays outside the DPF layer.

Keyword PIR XORs dictionary records selected by a bit share with eval_sequence_xor(key, begin, end, records), which folds that loop into the sequence walk.

Prefixes

Poplar and PRAC score every node at one depth. eval_prefixes(out<i, length>, key) returns the 2^length prefix shares in one walk, and eval_prefix_inner_product(out<i, length>, key, values) dots them with a public vector, each replacing one eval_point per node.

verify_idpf_path<Depth> / sketch_path_level / sketch_path_parent ([path_sketch.hpp](@ref dpf/path_sketch.hpp)) check that an incremental key is a single path: each depth is weight-1 (or one nonzero weight), and each parent equals the sum of its children. Leaf one-hot remains [sketch_fold](@ref dpf/verifiable.hpp).

A growing index

PRAC's binary search key grows: the next comparison bit is appended to the index. dpf::extend / dpf::add_output realize [F_Grow](@ref ideal_functionalities): a dealer (or joint) view turns an existing dpf_key pair into a richer one. Specs are the same objects make_dpf accepts. Memoizer overloads skip the O(d) rewalk when both path memoizers are already filled through the frontier. extend_ds / add_output_ds realize [F_GrowDS](@ref ideal_functionalities): one Doerner–Shelat correction-word round on extend, or leaf-open only on add_output, from warm frontiers.

Cost (dealer F_Grow). O(d) rewalk of the spine, or O(1) seed reads with warm memoizers; one make_cw / advance on extend; one exterior leaf plant per new packing group. Zero rounds and zero bytes on the wire.

Cost (F_GrowDS). One interactive level on extend (blinds, CW share, advice, AND — same shape as one point_party level) plus leaf pads when planting. local_cw_protocol opens in-process (no wire). add_output_ds has no interior CW round.

A wide leaf is a vec. idpf of several vec payloads is the incremental wide key heapify uses for every level at once. That call already evaluates.

Gates LLAMA still has to assemble

One comparison and one public interval are already gates, and an int8_t comparison follows numeric order. A spline with several pieces is several dpf::ic keys in the program. One key whose payload is the coefficient vector of the selected piece, with Horner on the shares and the public x_hat, is the Boyle gate LLAMA cites. dpf::gt(hi, lo) with both payloads a [dpf::vec](@ref dpf/vec.hpp) carries the vector. That is a wide comparison leaf. dpf::gt(hi) alone uses a zero of the same type as hi, so a dpf::vec false payload is the zero vector. Signed extension and truncate-reduce are grotto::sign_extend / grotto::truncate_reduce (wrappers over make_carry_keys + eval_carry_extend / eval_carry_in). Grotto's piecewise Horner evaluates a public polynomial from a point key. It is a different construction.

Sketches on the group the protocol writes

Express audits the same key that writes the mailbox. dpf::blob<N> is the XOR leaf for a row of N bytes. A caller fold eval_full_add_into(buf, key, fold) invokes fold(index, share) once per written output in the same walk; dpf::sketch remains the fp61 weight-1 fold. Pika's malicious check is the same shape over Z/2^k, using their bilinear Schwartz–Zippel lemma.

eval_full on an extractable fp61 key now matches eval_point on every lane, including both lanes of the packed leaf that holds the programmed point. Express folds its expansion and audit in one eval_full_add_into(box, key, dpf::sketch(...)) walk. extractable_full_test pins the share-for-share match.

Punctured PRF

KKRT's PSI is cuckoo hashing plus a set check in the application. The library object is dpf::pprf_master / dpf::puncture / dpf::pprf_eval on the AES PRG, including a 128-bit domain. The programmed value at the punctured point is stored beside the puncture when the master holder computes it. The same walk over a set H returns a dpf::pprf_copath: every node whose parent lies on a path to H and which itself does not, with shared prefixes stored once. Empty H publishes the root; a full-domain H publishes nothing. The walk is O(|H| · n) and never materializes the domain. An audit opening of a replica-seed pool is that copath with program_hidden = false, so the live seeds stay out. The one-point layout stays for PSI; {α} with programming agrees with it.

\htmlonly

TL;DR. Each program is only the DPF step, in one process, with a dealer standing in for shared keygen. Memory reads are a unit vector, a public shift or a prefix, and a dot. PIR and grouped sums are inner products. Heavy hitters are prefix walks. Mailbox writes and the ledger are full-domain adds, plus a proof when the write must be a single point.
\endhtmlonly