↓ Skip to main content
  1. Index/
  2. 📄 Publications/

Perseus: Faster FHE Transformer Inference via Complex-Packing and Sparse Bootstraps

Alessandro Zirilli , 

Davide Marincione

, 

Evgenios M. Kornaropoulos

, 

Giuseppe Ateniese

, 

Emanuele Rodolà

·1 min
Paper coming soon! In the meantime, the code is already public.

Under CKKS every multiplication consumes part of a ciphertext’s level budget, and once it runs out the value must be refreshed by a bootstrap, by far the most expensive operation in the model. Where bootstraps sit, and which kind is used, decides most of the inference time, yet existing tools place them from the multiplication count alone, without looking at the values being refreshed.

Perseus records one encrypted forward pass of the model (every ciphertext, its data layout and its magnitude) and plans bootstraps from it: a minimum-cut algorithm picks the positions, each bootstrap gets the scaling its values allow, and a cheaper sparse bootstrap is used wherever a ciphertext holds only a few distinct values. On encrypted GPT-2 generation, Perseus runs 416 bootstraps per token against 584–894 for DaCapo, Orion and Fhelipe, and takes 11.78 s per token against 16.09–21.40 s.

In the card below you can find our codebase, with the planner, GPT-2 on the Perseus runtime, and every plan behind the results.

Alessandro Zirilli
Author
Alessandro Zirilli