Google Open-Sources HEIR: One-Click Compilers for Encrypted AI Inference
Google's HEIR compiler converts pre-trained AI models to run on encrypted inputs — servers compute on ciphertext and never see your data. Here's how it works and why it matters.
For years, fully homomorphic encryption (FHE) has carried a strange reputation in applied cryptography: simultaneously “the holy grail” and “forever ten years away.” On August 14, 2026, Google took a meaningful step toward closing that gap. In a post on its Security blog, staff software engineer Jeremy Kun unveiled HEIR — the Homomorphic Encryption Intermediate Representation — as the latest addition to Google’s Private Computing Toolkit, an open-source compiler that converts pre-trained AI models to operate on encrypted inputs.
The pitch is simple to state and profound in implication: a server running HEIR-compiled models can process ciphertexts and return encrypted results without ever seeing the underlying data. A cloud service can recommend content without knowing your features. A fraud-detection system can score transactions without reading them. The model’s owner also keeps their intellectual property safe — shipping a proprietary model to a user’s device is no longer the only privacy-preserving option.
The privacy-utility trade-off, redefined
Google frames HEIR against a familiar dilemma. Standard protections like end-to-end encryption keep user data safe from breaches, but they blind the service provider: no spam filtering, no virus detection, no personalization. Sensitive sectors such as healthcare and finance are even more constrained, with strict regulations limiting data sharing across institutions. The alternative — local processing on the user’s device — is limited by device capability and risks leaking the model weights that constitute the provider’s IP.
Homomorphic encryption fundamentally alters this trade-off by allowing computation directly on encrypted data. It doesn’t eliminate cost — FHE still carries a nontrivial overhead compared to plaintext computation. But as Kun points out, it converts the capability/privacy trade-off into a question of cost, and that cost is rapidly decreasing. Unlike hardware-based approaches such as secure enclaves, FHE’s guarantees are purely cryptographic — no trusted manufacturing chain, no side-channel surface, just math.
The catch has always been usability. Manually converting an existing program to use homomorphic encryption efficiently, Google notes, “requires a team of cryptographers.” That is the problem HEIR exists to solve. The stated vision is a “one-click solution” that lets non-experts incorporate encrypted inference into production applications.
What HEIR actually is
HEIR is an MLIR-based compiler toolchain — built on the same Multi-Level Intermediate Representation infrastructure that underpins much of the modern LLVM ecosystem. A paper published on arXiv last year (“HEIR: A Universal Compiler for Homomorphic Encryption”) describes the ambition: support all mainstream homomorphic encryption techniques, integrate with the major FHE software libraries, and target hardware accelerators.
That last part is where the ecosystem story gets interesting. Google reports that since announcing the project’s intentions in 2023, the FHE community has embraced HEIR, and the company has partnered with hardware accelerator developers including Belfort, Niobium, Cornami, and Optalysys. Compilers and silicon co-evolving is how every major compute transition has actually happened; FHE appears to be following the same script.
HEIR has also become a productive research platform. Because cryptographers can build on the existing infrastructure for testing, benchmarking, and comparison rather than reinventing it, the project has drawn collaborations with Georgia Tech, Carnegie Mellon, UC Santa Barbara, Illinois Institute of Technology, Purdue, the University of Edinburgh, and Tsinghua University, among others. Four peer-reviewed publications have been built on HEIR to date, with more in preparation.
Four working demos
To demonstrate how far the technology has come, Google shared four private-inference applications, each compiled with HEIR, with latency numbers reported for a single-threaded CPU. The source code for all examples is available in the project’s GitHub repository:
- Deep Learning Recommendation Model (DLRM) — private content recommendation, joint work with Belfort Labs, LG, and New York University. This is the demo Kun highlights as literal proof of the core claim: a cloud service providing recommendations without seeing the user’s features.
- Credit card fraud detection — a fraud detector compiled together with Niobium and hardshell.ai, letting a payment provider score fraud risk without exposing raw transaction data to the inference server.
- Network intrusion detection — Google compiled the Kitsune anomaly-detection system with Niobium, enabling detection of threats in encrypted network traffic without revealing packet contents to the service provider.
- Hotword detection — compiled with Belfort Labs, allowing an audio-triggered AI agent to recognize wake words while protecting the privacy of the recordings. This one has obvious implications for always-listening assistant devices, one of the most scrutinized privacy surfaces in consumer tech.
The selection is pointed: recommendation, payments, network security, and voice are precisely the domains where the tension between useful AI and data privacy is most acute — and where regulators have been most active.
Why this lands now
Three currents converge to make HEIR’s graduation into Google’s Private Computing Toolkit well timed.
First, AI inference is moving into regulated and sensitive domains faster than privacy infrastructure can follow. Healthcare diagnostics, financial risk scoring, and enterprise data processing all want model intelligence without data exposure. FHE-compiled inference is a credible architectural answer.
Second, the cost curve is bending. Google explicitly promises to demonstrate the latency benefits of its hardware accelerator partnerships “in the near future,” and the single-threaded CPU latency figures in the demos suggest the overhead penalty — historically FHE’s killer objection — is now being engineered down rather than treated as a law of nature.
Third, the open-source posture matters. By shipping HEIR under an open license with working examples, Google is effectively inviting the industry to standardize on its compiler stack for encrypted computation — a land grab in the privacy-infrastructure layer analogous to what TensorFlow and then JAX did for ML frameworks. Rivals including AWS have been experimenting with FHE inference on SageMaker, but no vendor had previously put a serious, extensible compiler toolchain forward as community infrastructure.
The honest caveats
Enthusiasm should be calibrated. FHE remains substantially slower than plaintext inference for large models; the demos showcased are relatively small, narrow-purpose models — recommenders, fraud scorers, hotword detectors — not frontier LLMs. Latency numbers on single-threaded CPUs for such workloads are impressive, but nobody is running a transformer with billions of parameters homomorphically at interactive speeds today. The hardware accelerators from Belfort, Niobium, Cornami, and Optalysys are meant to change that equation, and the “one-click” vision is aspirational rather than fully realized.
Still, the trajectory is what matters. A compiler that lets ordinary engineers target encrypted execution without a cryptography team, backed by a company with the scale to push it into production, is how “forever ten years away” technologies eventually arrive. HEIR is worth watching — and, given that it’s on GitHub under an open license, worth trying.