# Protocol composition {#protocol_compose} \htmlonly
ELI5. Record every expand, multiply, and open as a node with a share domain. Identical work is interned. Independent opens share one RoundSink round; a dependency chain becomes successive waves. The schedule is what a hand-tuned Express or Duoram party would have written by hand.
\endhtmlonly `dpf::protocol::composer` ([compose.hpp](@ref dpf/compose.hpp)) records multi-protocol strands on one sink. Values are tagged with a share domain matching [secret_share.hpp](@ref dpf/secret_share.hpp): | Domain | Meaning | | --- | --- | | `fss` | FSS / DPF leaf share | | `a` | (2,2) additive | | `b` | (2,2) subtractive | | `rss` | (2,3) replicated | | `y` | (3,3) additive (RSS product factor) | Party-count changes are never implied: use `reshare` (or `rss_from_y` for `y`→`rss`). Local casts use `as`. ## Building a schedule ```cpp dpf::protocol::composer c(/*party=*/0); auto seed = c.input(dpf::protocol::domain::fss, 16); auto leaf = c.fss_point(seed, /*depth=*/8, /*slot_bytes=*/16); auto p = c.default_plan(); // == schedule(); RoundSink uses this // p.rounds() == p.exchange_waves() == 8 (no empty compute-only sink rounds) ``` Drive with [drive](@ref dpf::protocol::drive) or party `util::drive_composed` / `util::drive_composed_trio`, optionally with `drive_options` (`from_exchange_wave`, `compact_sink`, `beavers`). Express/Sabre audits use `util::schedule_fused_audit` (or `fss_point_fused`). The same plan lowers onto the RoundSink batch loop with `plan_to_schedule` → [`schedule_session`](@ref dpf::protocol::schedule_session) → `finish_schedule`, or the one-shot `drive_via_schedule`. Each exchange wave becomes a [`schedule_round`](@ref dpf::protocol::schedule_round) whose `produce` runs on this thread when that instance's peer slot is ready (the `schedule_session::drive` scan). Rounds carry an [`edge_id`](@ref dpf/net/edge_mesh.hpp) on an [`edge_mesh`](@ref dpf/net/edge_mesh.hpp) (star / clique / dealer), a [`receive_rule`](@ref dpf::protocol::receive_rule) (`domain_open`, `copy_peer`, `field_sum`, `any_two`, `verify_*`, `eq_check`), optional `branch` (skip send) and `next` (jump after an open). [`session_host`](@ref dpf/session_host.hpp) queues micro-plans on a durable mesh. Pad graphs live in [pad_graphs.hpp](@ref dpf/pad_graphs.hpp) (`du_atallah_mul_graph`, `star_upload_answer_graph`); IKNP setup in [iknp_graphs.hpp](@ref dpf/iknp_graphs.hpp); PIRsona / hushmap builders in [mesh_apps.hpp](@ref dpf/mesh_apps.hpp). Named application plans and `drive_star` / `submit_and_drive` live in [app_plans.hpp](@ref dpf/app_plans.hpp). Replayable seeds and paper cost CSVs: see [Experiments](@ref experiment_costs). In short, `dpf::experiment` / `app::measure_plan` / `app::run_measured` attach a per-thread master seed (default random, or `replay`) and emit CSV breakdowns. `plan::rounds()` equals `exchange_waves()` and `slot_bytes_all().size()` — the sink and the scheduler share one count. Trailing compute-only DAG waves still appear in `waves()` / `wave(i)` but do not allocate RoundSink rounds. Composer-owned ABY sessions default to `beavers::schedule_objective::rounds` so online sign×polynomial stays one round. Dealer benches that want Appendix-E peels keep `prep`. Independent circuits use `aby_lane(i)`. Live δ openings use `drive_options::beavers` (`util::u64_beaver_host` wraps `party_batch_stepper`). ## Hand-schedule shapes the API covers | Shape | Call | What it saves | | --- | --- | --- | | L‖R PRG stretch | `expand_pair` / default `level_walk` | Half the AES vs separate child expands | | Shared-seed fan (Grotto LUT) | `fan` of `fss_cmp` on one seed | One expand per level, not N×depth | | Several piecewise LUTs | `schedule_lut_union` | One `fss_cmp`; the union's prefix walk is local | | Express / Sabre audit | `fss_point_fused` / `schedule_fused_audit` | Sketch in last CW — no +1 exchange wave | | BGI Remark 3.4 early-stop | `fss_point_early_stop` | Drop ν interactive CW rounds | | Poplar / idpf prefix checkpoints | `level_walk_prefixes` | Prefix share after each CW, no extra rounds | | Adaptive idpf_agg | `step` → drive tail → `retain` → `step`… | One packed L‖R open per depth; child bit never on the wire | | DCF `block_width` | `level_walk_sized` | Per-level CW slot bytes | | Doerner–Shelat keygen | `level_walk_ds` / `level_walk_ds_sized` | blind‖CW‖advice‖AND + OH AND-layers | | Keyword PIR / PSI buckets | `multipoint_fan` / `exchange_pack` | Bucket CWs pack; answers one open | | RSS mul + neighbor refresh | `rss_product_replicated` / `rss_from_y` | One y-exchange, not a reconstructing open | | FSS leaf → ABY scale | `aby_product` after `fss_point` | Leaf stays on the beaver barrier critical path | | Multi-lane ABY | `aby_lane(i)` / `aby_product(..., lane)` | Independent barriers, shared wave | | Round-aware Beaver | `composer::aby()` | `schedule_objective::rounds` | | Prepaid expand / rotate | `defer_expand` / `rotate_share` | Zero online FSS rounds (SUBLEQ offline) | | Duoram leaf_later | `leaf_later_walk` / `apply_leaf_correction` | Path CWs only; leaf apply is local | | RSS column, hidden order | `shuffle_hidden` | Three `shuffle_send` waves. `aux` is the left-out party: 2, then 0, then 1 | Adaptive prefixes: `step_adaptive_prefix` schedules only the next packed open; drive with `from_exchange_wave = exchanges_flushed` and a sink sized by `plan::slot_bytes_from` (`compact_sink = true`); then `retain_adaptive_prefix` adds one local tip. Never unroll a full-depth adaptive walk into one static schedule. 3PC RSS: `rss_from_y` schedules a `domain::y` exchange (neighbor receive). `util::drive_composed_trio` splits peer maps per exchange — `y` on the neighbor ring, `dealer_deliver` from p2, everything else on p0↔p1. Cross-party `reshare` refuses a reconstructing open; use `reshare_with_mask`. Opens use domain algebra (`a`/`fss` sum, `b` subtractive, `y` copy, optional `field_open::fp61`). Authenticated Beaver sessions size barriers as `auth_opening` and drive through `u64_auth_beaver_host`. `exchange_fuse` packs extra payloads into one open. `schedule_cuckoo_probes` turns occupied bucket ids into `multipoint_fan`. ## Client and servers {#compose_client} Keyword PIR and both three-server PIRs are a star, not a correction-word walk. `client_servers(servers, query_bytes, answer_bytes)` records two waves and nothing between the servers: 1. Upload. The slot is `servers * query_bytes` (every key in one round). 2. Answer. The slot is `servers * answer_bytes`. It depends on the upload, so it cannot share that wave. `drive` copies the peer's message into the exchange node. It does not add the shares. A two-party sink stands in for one client and one server; the slot width is still the full parallel payload, which is what `dpf::app::exercise` reports as `bytes`. ```cpp auto q = c.client_servers(/*servers=*/2, /*query_bytes=*/256, /*answer_bytes=*/4); auto p = c.default_plan(); // p.rounds() == 2 // p.slot_bytes(0) == 512, p.slot_bytes(1) == 8 ``` ## Parking and the fleet {#compose_fleet} `drive` used to spin until the peer flushed, then throw. A step that simply takes a long time looks like that failure, and a pool that always resumes the side already waiting on a receive never runs the peer who could unblock it. `drive_options::park_if_waiting` returns instead. `drive_cursor` remembers the wave, the exchange index, and whether the submit already happened. The next `drive` with that cursor receives if the peer has caught up, or parks again. `one_exchange` stops after a single completed round so a scheduler can interleave other instances. `dpf::app::run_fleet(composer, instances, chaos_seed)` is that scheduler. Each instance is a pair of parties on an in-process sink. Workers prefer a side that still has a submit to do over a side parked on a receive, and among those they prefer the one further behind. Delays are per side and per step. They do not line up, which is the case that makes round-robin and "always resume the waiter" stall. `dpf::app::run(name, composer, expect_rounds)` is the one-shot experiment used by `examples/applications/`. It drives both parties and prints ``` name rounds=R bytes=B ``` `bytes` is the sum of one lane's exchange slots. Opcodes from `beaver_delta` (300) through `beaver_delta + 1023` are δ barriers. `user_base` (1000) sits inside that window, so a hand-rolled kernel opcode in that range is skipped as a beaver. Use 5000 and up for kernels `drive` must run or reject. ## What eval_full reuses `eval_full(key, buffer)` keeps a thread-local workspace for that key type. The same key evaluated again does not rebuild the interior. A different key still does. Leaves of one block are stretched eight at a time (`eval_x8`). Pass your own memoizer when two evaluations on one thread must not share that cache. DS OH: `level_walk_ds(..., oh=true)` emits `net::ds_oh_exchanges_per_level` (80) AND-layer opens per level. Size sinks with `net::compose_ds_slot_bytes` (schedule-exact) or the conservative `net::ds_walk_slot_bytes` for hand `dist_ds` paths. \include{cpp} protocol/compose_schedule.cpp ## Gap analysis {#compose_gaps} | Item | Status | | --- | --- | | Point-key stretch | Builtin `fss_expand_pair` runs `prg::aes128`; `fss_step+1` XORs the opened correction onto the children | | Setup plus walk | `iknp_setup_graph` / `plan_with_pad_setup` — pad frames are `schedule_round`s | | Edge mesh | `edge_mesh` + `make_memory_star` / `make_memory_clique`; `drive_via_schedule(plan, mesh, …)` | | Receive rules | `field_sum`, `any_two`, `verify_sketch` / `verify_proof`, `eq_check` in `apply_peer_slot` | | Jump / host | `schedule_round::next` + `session_host` for SUBLEQ / idpf_agg / hushmap ops | | PIRsona / hushmap | `pirsona_bitmore_fetch`, `pirsona_gd_update_graph`, `hushmap_add_schedule` | | 2PC → RSS | `reshare` casts to a `y` share (p2 holds 0) and `rss_from_y`. `reshare_fresh` adds a dealer-sampled zero mask | | Auth openings | `auth_beaver_host` / `u64_auth_beaver_host`. `party_auth_batch_stepper::apply_peer` rejects a bad tag | | One DS sink | `ds_walk_slot_bytes` is the envelope of the hand walk and `compose_ds_slot_bytes` (5 peer opens per level + OH + mux) | | Wide opens | Every multiple of 8 bytes is a lane add/sub (`field_open::fp61` for the Mersenne field) | | Fuse offset | `exchange_fuse` stores the segment offset in `aux` (`uint32_t`) | Still outside this scheduler: the cuckoo PRP, OPRF evaluation, the full IKNP wire body inside each pad `produce` (frames and count are in the schedule; `iknp::sample` still owns the crypto), and key-specific control-bit / verifiable / Doerner–Shelat advice logic. Those last ones override the builtin kernels. Trio multi-edge `drive_via_schedule` still needs `edge_sinks` wired from the party helper. ### Closed earlier Incremental drive, domain-correct `a`/`b`/`y` opens, split y/2PC peer maps, OH AND-layers, `level_walk_ds_sized`, `default_plan`, staged adaptive retain, multipoint fan, multi-lane ABY, `defer_expand` / `leaf_later_walk`, Express trailer fuse. **Go deeper:** [compose.hpp](@ref dpf/compose.hpp), [beaver.hpp](@ref dpf/beaver.hpp), [sink_exchange.hpp](@ref dpf/net/sink_exchange.hpp), [Application sketches](@ref applications), [guided tour](@ref tour_party).